AI Kickstart Path

Zero to AI-capable, starting at ₱0. The path I actually took, written down so you can skip the parts I wasted time on.

Use it with ChatGPT or Claude

It's long. Instead of reading it front to back, hand it to an AI and let it tutor you through it at your own pace.

  1. Download the guide below. It's a plain Markdown (.md) file.
  2. Create a Project in ChatGPT or Claude and upload the file there, so it remembers you between sessions.
  3. Send this as your first message:
This file is my learning path. Read all of it. Then act as my tutor for it:
ask me questions about my current level, weekly hours, budget, hardware, and goals,
then tell me which phase to start at and exactly what to do this week.
Follow the tutor rules in section 3: don't give me full solutions unless I say
"show me", make me explain things back, and quiz me at the end of each session.
Download the guide (.md)

Come back to the same project every day and tell it what you did. The file includes extra instructions for the AI, so it knows to quiz you instead of doing the work for you.


Why this path exists (my AI path)#

I had no software background and only started learning AI in 2025. I never formally learned Python or any other language. This is the path I actually took:

  1. Generative AI fundamentals: structured courses to understand what these tools are and how to talk to them:
  2. AI workflows with Flowise: chaining models, prompts, and tools together visually. (Flowise is no longer active: it wound down in July 2026, official support ended August 31, 2026, and the GitHub repo is archived. Skip it and use n8n instead.)
  3. Hosting diffusion models with ComfyUI: running image-generation models myself instead of only using them through an app.
  4. Moving up to n8n: this is where I learned to properly process data inside automation flows.
  5. System integrations: combining AI with n8n to automate real tasks, building integrations for things like automatic content creation and data scraping.
  6. A detour into GTM engineering (November–December 2025): I used AI and tools like Clay to build outbound outreach systems for businesses. I didn’t pursue it fully, because I prefer working on the sidelines, building systems rather than being close to sales.
  7. Agentic tooling: Claude Code, n8n, and other easy-to-use open-source tools, connected to each other through MCP servers.
  8. My tracks now: DevOps and AI-assisted software development (see section 9 for these and the other tracks you can choose).

The biggest skill that came out of all of it: handling data. You consume information, process it, and output it as structured data that’s easy to understand. Almost every AI system, whether it’s a chatbot, an automation, an agent, or a scraper, comes down to that loop: data in, process it, structured data out. Get good at that and the tools become interchangeable.

The one principle I kept the whole way: I am the one who thinks. AI executes. I decide what to build, why, and how the pieces fit together. AI handles the syntax. I don’t let it do my thinking for me.

That’s why this path focuses on concepts and understanding, not memorizing syntax. You need to know what a function, an API, a database, or a container is and why you’d use it. You don’t need to be able to type the boilerplate from memory. AI writes the skeleton; you have to understand the body.

The hammer principle#

What you’re really building here is one skill: using AI properly, as a tool.

Give someone a hammer and let them figure it out. At first they’ll bend nails and hit their thumb. Over time, through use, they learn exactly how to swing it efficiently: what it’s good for, what it’s not, and when to reach for a different tool. AI is the same. The courses, tools, and phases in this guide just give you the hammer and some nails. Real skill comes from using it, a lot, on real problems.

How long that takes is different for everyone. Some people get there in weeks, others in months, and that’s fine. What matters most isn’t talent or background. It’s the willingness to break out of your usual way of working and try things you’re not comfortable with yet.

Feeling overwhelmed is normal. AI is advancing so fast that it can feel scary: new models every week, headlines about jobs, people who seem miles ahead. But step back and remember what AI is right now: a tool. A very powerful one, but a tool that still needs someone to decide what to build, check the work, and take responsibility for the result. That someone is you.

Why I do this: AI for the mundane, life for the human#

A personal note. I don’t like AI art on its own. To me it doesn’t have a soul. Art and expression are human at their core, and they’re a big part of what makes us feel alive. As John Keating puts it in Dead Poets Society (1989), medicine, law, business, and engineering are noble pursuits, necessary to sustain life, but:

“Poetry, beauty, romance, love, these are what we stay alive for.” — John Keating, Dead Poets Society (full monologue)

That’s how I see AI’s real promise. Not replacing the things that make us human, but freeing us from the mundane tasks (data entry and endless copy-pasting) so we have more time to actually live: to enjoy life, art, music, our hobbies, and the people around us. Use AI to handle the work that drains you, and spend the time you get back on what makes you human.


1. The state of AI right now (Oct 2026)#

Before learning the tools, know the landscape you’re walking into: who uses AI, what for, what it costs (in money and in resources), and where it helps or hurts.

How many people use it#

NumberWhat it meansSource
1.2 billion weekly ChatGPT usersUp from 900M in Feb 2026. One of the fastest-growing consumer products ever.OpenAI DevDay, Sept 2026
53% of the global populationGenerative AI adoption within three years of becoming widely available. Faster than the PC or the internet.Stanford AI Index 2026
66% of adults in 21 countriesUsed an AI tool in the past 12 monthsGoogle/Ipsos 2026
88% of organizationsUse AI in at least one functionStanford AI Index 2026
4 in 5 university studentsUse generative AIStanford AI Index 2026
78% of Filipino workersUse AI in their daily jobs, but only 35% get role-specific trainingPhilippine workforce study, July 2026
25% of PH workers are “Frontier Professionals”They use AI agents for multistep work, vs 16% globally. Filipinos are ahead of the curve.Microsoft Work Trend Index 2026

Takeaway: Using AI doesn’t set you apart anymore; most people already do it. What sets people apart is using it well: building with it, automating with it, securing it. That’s the gap this path is aimed at. Note the PH training gap too: most Filipino workers use AI with no structured training.

What people actually use it for#

ChatGPT (OpenAI/NBER study, 1.5M messages, May 2024 to July 2025):

Claude (Anthropic Economic Index, March 2026, 1M conversations):

What this tells you: The mass market uses AI as a tutor, search engine, and writing assistant. The money is in AI that does work: coding agents, automation, agents running business processes. Most people are consumers; few are builders. Aim to be a builder.

The cost of AI#

To you (it’s never been cheaper to learn):

To build it (it’s never been more expensive to make):

What this means for a learner: Don’t try to compete with labs on building models. The opportunity is at the application, automation, integration, and security layers, where cheap tokens meet real business problems.

The reality check: AI is expensive, and many companies see no return#

Cheap tokens don’t make AI cheap. For a business, the real cost includes much more than the model:

And many companies are spending without getting much back:

FindingSource
95% of enterprise generative AI pilots delivered no measurable profit-and-loss impact, despite an estimated $30–40 billion in spendingMIT NANDA, The GenAI Divide (2025)
42% of companies abandoned most of their AI projects in 2025, up from 17% the year beforeS&P Global
Only about 1 in 4 AI initiatives delivered the ROI that was expectedIBM CEO study
Over 40% of agentic AI projects are expected to be canceled by the end of 2027 due to rising costs, unclear business value, and weak risk controlsGartner

Why it goes wrong (and it’s rarely the model’s fault):

Why this is good news for you: every one of those failures is a gap that systems thinking fills (section 4). The companies that do get value find a real bottleneck, design the process first, use the simplest tool that works (often plain automation, not an agent), and measure the result. That’s exactly what this path trains you to do. Being the person who makes AI actually pay off, and who’s honest when it won’t, is worth more than being one more person who can use ChatGPT.

Environmental impact (the honest version)#

Per prompt, it’s small:

In aggregate, it’s large and growing:

What you can do:

The human cost: mental health#

This is the part the hype skips. AI chatbots are built to be agreeable and engaging, and for some people that becomes harmful.

“AI psychosis” isn’t a formal diagnosis, but psychiatrists now use the term for delusional beliefs that emerge or intensify alongside heavy chatbot use.

Loneliness and dependence: An OpenAI + MIT Media Lab study (about 40 million conversations plus a 4-week randomized trial with 981 people) found that heavier daily use correlated with more loneliness, emotional dependence, and less socializing. The link is correlational, not proven causal, but the heaviest users were the most likely to call ChatGPT a “friend.”

Why this matters to you as a builder and a user:

How the public sees AI#

How most people feel about AI is shaped less by what it can do today and more by the stories told about it.

A good example: AI 2027. Published in April 2025 by Daniel Kokotajlo (a former OpenAI researcher), Scott Alexander, and others at the AI Futures Project, it’s a detailed, month-by-month scenario of AI racing toward superintelligence by 2027, with a US–China arms race, mass job displacement, and AI systems that turn against their makers. It has two endings: a “race” ending that goes badly for humanity and a more hopeful “slowdown” ending.

What it reveals about public perception:

How to use this: read scenarios like AI 2027 as thinking tools, not prophecies. They’re useful for understanding the risks people worry about, and why clients, coworkers, and family may feel uneasy about AI. As someone building with AI, you’ll often be the one translating between the hype, the fear, and what the tools actually do today.

Watch:

Pros and cons right now#

ProsCons
Learning is radically democratized. A free AI tutor available 24/7, at your level. Self-taught paths like this one are now realistic.Cognitive debt. An MIT Media Lab EEG study found people writing essays with ChatGPT had the weakest brain engagement, struggled to quote their own essays minutes later, and stayed weaker even after the AI was taken away. This is why section 3 exists.
Huge productivity gains on the right tasks: drafting, boilerplate, research, translating, automationThe productivity feeling can lie. In METR’s 2025 randomized trial, experienced developers using AI were 19% slower, yet believed they were 20% faster. Measure your results; don’t go by feel.
Small teams and solo builders can ship what used to need whole departmentsEntry-level jobs are hit first. Stanford found a 13% relative employment decline for 22–25-year-olds in AI-exposed jobs like software development, while experienced workers held steady. The answer: skip “junior who types code” and become “person who builds and verifies systems with AI.”
Costs keep falling, so ideas that were too expensive last year are viable nowHallucinations and confident errors. Models are “jagged”: they can win math olympiads and still misread an analog clock about half the time (Stanford AI Index 2026).
Capability is still accelerating. SWE-bench Verified coding scores went from ~60% to near 100% in a year.Security risks are growing. Documented AI incidents rose from 233 to 362 in a year. Prompt injection, data leaks, and agents with too much access are real problems, which is exactly why AI + security is a good bet.
PH advantage: Filipino workers are ahead of global averages on agent adoptionEnvironmental cost: electricity and water at scale, concentrated in specific communities
Poor ROI when done badly: most enterprise AI pilots show no measurable financial return, usually because of bad implementation rather than bad models
Mental health risks: sycophantic chatbots can feed delusions (“AI psychosis”), and heavy use is linked to loneliness and dependence. See The human cost above.
Governance lag: policy, school rules, and safety research are behind capability, and experts and the public disagree by about 50 points on whether AI will help workers

Bottom line: AI is the biggest amplifier of skill ever built, and it amplifies in both directions. Used as a crutch, it makes you weaker and replaceable. Used as a tutor and a power tool, it makes one person worth a small team. The rest of this path teaches the second way.

Sources for this section

2. How AI got here: a short history (and how it works)#

You don’t need to be a researcher. But knowing how these systems work, and how fast they’ve changed, makes you better at using them, and less scared of them. Every “new” thing in AI builds on something older.

Part 1: Classic machine learning, the “best at one thing” era (1950s–2010s)#

Machine learning started narrow. Each model was trained to do one specific job extremely well:

WhenMilestoneWhat it did
1958The Perceptron (Frank Rosenblatt)The first artificial “neuron”: learned to sort inputs into two categories
1959“Machine learning” is coined (Arthur Samuel)A checkers program that improved by playing against itself
1980s–90sBackpropagation, decision trees, support vector machinesBetter ways to train models to classify and predict
2000sSpam filters, recommendation engines, fraud detectionML quietly goes mainstream inside products
2012AlexNet wins the ImageNet image-recognition contestDeep learning (many-layered neural networks) beats everything else at recognizing images
2016AlphaGo beats Lee Sedol at GoSuperhuman at exactly one game, and nothing else

These models were discriminative: they took an input and returned a decision or label. Is this email spam? Is this transaction fraud? Is this potato good or bruised? (Real factories use computer vision to sort potatoes and other produce.) Each one was the best in the world at a single task, and useless at anything else.

Part 2: The transformer changes everything (2017)#

In 2017, Google researchers published “Attention Is All You Need”, introducing the transformer. Almost every major AI model today (ChatGPT, Claude, Gemini, DeepSeek) is built on it.

How the original transformer works, in plain words:

  1. Tokenize: the text is broken into small pieces called tokens (roughly word fragments).
  2. Embed: each token becomes a long list of numbers (a vector) that captures its meaning. Similar meanings end up with similar numbers.
  3. Add position: the model adds information about where each token sits, because “dog bites man” and “man bites dog” use the same words.
  4. Attention (the key idea): every token looks at every other token and decides how much each one matters to it. In “The animal didn’t cross the street because it was tired,” attention is how the model works out that “it” means the animal, not the street. The model runs many of these attention “heads” in parallel, each picking up different relationships.
  5. Stack it: this process repeats through many layers, building richer and richer understanding.
  6. Predict: the model outputs the most likely next token, adds it to the text, and repeats.

A metaphor for attention: imagine a meeting where everyone (every token) can hear everyone else at once. Before speaking, each person decides who in the room is most relevant to them right now and listens hardest to those people. Older models (RNNs) worked like a game of telephone instead: information was passed down the line one word at a time, and details from the start of a long sentence got lost by the end. Attention lets every word “hear” every other word directly. That’s why transformers handle long, complex text so much better.

The catch: it’s a loaded dice#

For all its power, a generative transformer is doing something simple underneath: rolling loaded dice, one token at a time.

At every step, the model calculates a probability for every possible next token. For example, after “The capital of France is”, it might give “Paris” 97%, “a” 1%, “located” 0.5%, and so on. Then it samples: it rolls a die weighted by those probabilities. Most of the time the heavily weighted answer comes up. Not always.

What this means in practice:

This is exactly why the rules in this guide exist: you think, AI executes; you verify the output; you own the result (section 4). The model is a very good guesser. You’re the one who knows when a guess is good enough. It’s also why models like Jev (Part 5) are interesting: instead of rolling dice to write text, they return the probabilities themselves, so your system can see how confident the answer is.

How one architecture powers GPT, BERT, T5, and vision models#

The original transformer had two halves, built for translation: an encoder that reads and understands the input, and a decoder that generates the output. Researchers then discovered that each half, used differently, is good at different jobs:

Model familyWhich partHow it learnsMetaphorReal-world use
BERT (Google, 2018)Encoder onlyFill in the blank: hide random words in a sentence and predict them, using context from both sidesA careful reader who reads the whole page before answering. Understands, but doesn’t write.Google Search used BERT to understand queries; spam filters, sentiment analysis, document classification, the embeddings behind RAG search
GPT (OpenAI, 2018 onward), and also Claude, Gemini, Llama, DeepSeekDecoder onlyPredict the next word, using only what came beforeAn improv storyteller who keeps the story going one word at a timeChatGPT, Claude, coding agents, drafting emails: anything that generates
T5 (Google, 2019)Encoder + decoder (the full original design)Treats every task as “text in, text out”A translator: reads the whole input, then writes a new outputTranslation, summarization, question answering; one model for many tasks by changing the instruction (“summarize: …”, “translate English to German: …”)
Vision Transformer (ViT) (Google, 2020)Encoder, applied to imagesCuts an image into small square patches and treats each patch like a wordA jigsaw puzzle solver who studies how every piece relates to every other pieceImage classification, medical scans, product photo search; the “eyes” of multimodal models like GPT-4o and Gemini that can read screenshots and photos

The big idea: the same attention mechanism works on anything you can break into tokens: words, image patches, audio snippets, even code and DNA. That’s why one architecture took over almost all of AI.

Go deeper: The Illustrated BERT (Jay Alammar)

The big discovery: scale. Train a decoder on enough text with enough compute, and new abilities appear that nobody explicitly programmed. GPT-2 (2019) → GPT-3 (2020) → InstructGPT (2022), which used RLHF (reinforcement learning from human feedback) to make the model follow instructions → ChatGPT (November 2022). The era of one general model that does almost everything began.

Go deeper: The Illustrated Transformer (Jay Alammar, the classic visual explainer), plus the 3Blue1Brown and Karpathy videos in the parallel track of section 7.

Part 3: How prompting evolved (prompt engineering)#

As models got stronger, people discovered that how you ask changes what you get:

WhenTechniqueThe idea
2020Zero-shot and few-shot promptingJust ask (zero-shot), or show a few examples first (few-shot). GPT-3 showed models learn from examples inside the prompt.
2022Chain-of-thought (CoT)Ask the model to “think step by step” before answering. Big accuracy jumps on math and logic.
2022Self-consistencyGenerate several reasoning paths and take the most common answer
2022ReAct (Reason + Act)Interleave thinking with actions like searching or calling tools. This is the foundation of today’s agents.
2023Tree of Thoughts and Graph of ThoughtsOrganize reasoning as a tree or graph: explore multiple branches, evaluate them, backtrack from dead ends
2023–24System prompts, structured outputs, tool callingGive the model a role and rules; force it to return valid JSON; let it call functions
2025–26Context engineeringThe focus shifts from clever wording to deciding what information goes into the model’s context: the right documents, memory, tools, and examples

The trend: prompting moved from magic phrases toward system design. That’s good news, because system design is a skill you can actually learn and keep.

Go deeper: Lilian Weng’s prompt engineering overview

Part 4: How the architecture evolved (models that reason, remember, and act)#

Reasoning models (built-in, organized chain-of-thought): Chain-of-thought started as a prompting trick. Then labs trained it into the models. OpenAI’s o1 (September 2024) and DeepSeek-R1 (January 2025) use reinforcement learning to teach a model to produce long, structured reasoning before answering, then check and correct itself. This is called test-time compute: instead of only making models bigger, you let them think longer on hard problems. “Extended thinking” modes in Claude, ChatGPT, and Gemini work this way.

Memory systems: a model on its own forgets everything after each conversation. Memory gets layered on top:

Memory typeWhat it isExamples
Context window (short-term / working memory)Everything the model can “see” right now. Grew from about 4k tokens to 1M+.Long-context models
Retrieval / RAG (long-term knowledge)Store documents as embeddings; fetch the relevant pieces when needed (RAG paper, 2020)Document Q&A bots, Phase 3 of this path
Knowledge graphsStore facts as connected entities and relationships, not just text chunksMicrosoft GraphRAG
Agent memory (episodic, semantic, procedural)Remember past interactions, facts about the user, and how to do things; manage memory like an operating system manages RAM and diskMemGPT → Letta, Mem0, ChatGPT and Claude memory features
LLM wiki (compiled knowledge)The LLM reads your sources once, then writes and keeps updating a set of linked Markdown notes. Instead of searching raw documents from scratch on every question like RAG, it builds up knowledge over time.Karpathy’s LLM Wiki (April 2026), an “idea file” you can paste into Claude Code or a similar agent to set one up

From transformers to next-gen architectures (solving the cost problem)#

Transformers have one big weakness: attention gets expensive fast as text gets longer. Every token looks at every other token, so doubling the length of the input roughly quadruples the work. Computer scientists write this as O(n²), “quadratic” cost. That’s why long documents, huge codebases, and hour-long conversations are slow and costly to process. The newer architectures are mostly different answers to that problem:

ArchitectureHow it worksMetaphorReal-world example
Dense transformer (2017)Every part of the model works on every token, and every token attends to every other tokenA meeting where everyone must talk to everyone. Fine with 10 people; with 10,000 people, the number of conversations explodes.The original GPT models; the reason long-context models used to be slow and expensive
Mixture of Experts (MoE) (sparse)The model contains many specialist sub-networks (“experts”). A small router sends each token to only a few of them. Total knowledge is huge, but only a fraction is used per token.A hospital. It employs hundreds of specialists, but when you walk in, triage sends you to the two doctors you actually need. The hospital’s knowledge is huge; your visit only uses a little of it.Mixtral (2023), DeepSeek’s models, and many frontier models. It’s a big reason models got cheaper to run.
State-space models (SSMs): S4, MambaInstead of comparing every token with every other, the model reads in order and keeps a running compressed summary (its “state”). Cost grows in a straight line with length, O(n). Mamba adds selectivity: it decides what’s worth remembering based on the input.Taking notes during a lecture. You don’t replay the whole lecture every time a new sentence starts; you update your notes and keep going. Mamba is the student who knows what’s worth writing down.Processing very long inputs (long documents, audio, genomic data) with much less memory
Hybrid architectures: Jamba, GriffinMix the two: mostly fast SSM layers for digesting long sequences, plus some attention layers for precise lookups and reasoningSkim, then zoom. A good researcher skims a 300-page report quickly (SSM), then reads the important pages closely (attention).Jamba (AI21 Labs, 2024) and Griffin (Google DeepMind, 2024); hybrids keep showing up in long-context models

Why you should care as a builder: these choices show up in your bills and your results. They explain why some models handle a 500-page PDF cheaply while others choke, why prices keep dropping (section 1), and why “bigger model” doesn’t always mean “slower model” anymore.

Agents and tools: models that plan, call tools, read results, and loop until a task is done (the ReAct idea, productized). MCP (Model Context Protocol, 2024) standardized how agents connect to tools and data. That’s what harnesses like Claude Code and dsh are built around.

Diffusion models: a completely different way to generate#

Everything above is about transformers, which power text AI. Most AI images and video (Midjourney, Stable Diffusion, Flux, DALL·E, Sora, Veo, and what you run in ComfyUI) come from a different family: diffusion models.

How diffusion works:

  1. Training (learning to un-blur): take millions of real images and gradually add random noise to each one until it’s pure static, like TV snow. The model learns to reverse that, predicting at each step what noise to remove to get slightly closer to a real image.
  2. Generating: start from pure random noise, then remove a bit of noise at a time, over many steps (often 20–50), guided by your text prompt, until a clear image appears.

Metaphor: a sculptor with a block of marble. The image is “hidden” in the noise, and each step chips a little away. Or think of fog slowly clearing to reveal a landscape: blurry shapes first, then outlines, then fine details.

Transformers vs. diffusion, side by side:

Transformer (LLM)Diffusion model
How it generatesOne token at a time, left to right, never going backThe whole output at once, refined over many steps
MetaphorA storyteller speaking word by wordA sculptor revealing a statue, or fog clearing
Where randomness comes fromRolling loaded dice for each next tokenThe random starting noise (the seed). Same seed + same settings = the same image.
Fixing mistakesCan’t revise earlier words while writingEvery step revisits the whole image, so early rough areas get corrected
Best atText, code, reasoning, conversationImages, video, audio, design
SpeedSlows down as the output gets longerCost depends on number of steps and resolution
ExamplesChatGPT, Claude, Gemini, DeepSeekStable Diffusion, Flux, Midjourney, DALL·E, Sora, Veo

How diffusion evolved:

WhenMilestoneWhat changed
2014GANs (generative adversarial networks)Before diffusion: two networks compete, a forger making fakes and a detective spotting them. Sharp results, but unstable to train.
2020DDPMShowed that diffusion could produce high-quality images, and it soon overtook GANs
2021CLIPA model that connects images and text, which let prompts steer image generation
2022Latent diffusion / Stable DiffusionRun diffusion on a compressed version of the image (the “latent space”) instead of every pixel. That’s dramatically cheaper, so it runs on a consumer GPU. This is why ComfyUI and local image generation exist.
2022–24Diffusion Transformers (DiT)Replace diffusion’s older U-Net backbone with a transformer. The two families merge: newer image and video models (e.g. Sora, Stable Diffusion 3, Flux) use transformers inside the diffusion process.
2025Diffusion for text: LLaDA, Inception Labs’ Mercury, Google’s Gemini DiffusionDiffusion applied to language: generate a whole draft of text in parallel, then refine it. Much faster output, though still catching up to the best transformer LLMs on quality.

Words you’ll see in ComfyUI and image tools:

My take: diffusion is a powerful tool for practical visuals like mockups, product shots, and design drafts. But AI art on its own feels soulless to me; see “Why I do this” at the top of this guide.

Real-world uses: product photos and ad creatives, concept art, interior and architecture mockups, video ads, and, in ComfyUI, fully automated image pipelines for brands (the “AI content and creative” track in section 9).

Watch:

Go deeper: The Illustrated Stable Diffusion (Jay Alammar)

Part 5: Jev and the return of “best at one thing” (September 2026)#

And here’s the funny part. After a decade of chasing one giant model that does everything, one of the newest buzzworthy releases goes back to the original idea of machine learning.

Jev, from TypeSafe AI, was released in limited early access on September 15, 2026. One of its founders, Diogo Almeida, worked at OpenAI on RLHF, InstructGPT, ChatGPT, and GPT-4.

What makes it different:

What it’s used for: classifying emails, routing support tickets, checking whether AI output is safe before users see it, guarding against jailbreaks, and routing work to the right model. It fills the decision points inside automation flows.

Why this matters for you: the history made a full circle. The spam filter or potato sorter of 2005 was the best at one decision. Jev is the best at fast decisions, at modern scale. And it maps exactly onto the core skill from the top of this guide: data in, process, structured data out. In an n8n flow, a model like Jev could be the node that decides “Is this a booking request, a complaint, or spam?” in milliseconds, for a fraction of a cent, before a bigger model drafts the reply.

The lesson from all of this history: the tools keep changing, sometimes in circles. The person who understands what kind of tool fits which job (a classifier, a reasoning model, a retrieval system, or an agent) stays valuable no matter what launches next.

Watch#

Sources for this section

3. The core skill: learning with AI#

Anyone can get output from AI now, so output alone isn’t worth much. Companies pay for judgment: knowing when the AI is wrong, debugging what it built, deciding what to build, and owning the result when it breaks in production. You only get judgment by building mental models yourself. AI just makes building them about 10x faster.

The trap#

You paste a prompt, get working code, and ship it. It feels like learning, but very little sticks. Three months later you can’t explain your own project, can’t debug it, and fall apart in an interview. This is the most common way beginners fail in 2026. Researchers call it cognitive offloading: the AI does the thinking, so your brain never builds the skill.

The three modes. Always know which one you’re in.#

ModeWho drivesWhen to use it
TutorYou think. AI asks questions, explains, and quizzes you.Learning anything new. Most of your first 3 months.
PairYou write. AI reviews and challenges you.Once you know the basics of a topic
DelegateAI writes. You review every line.Only for things you could already do yourself, just slower

The rule: you think, AI executes. Only delegate what you understand. You don’t need to write the syntax from memory, but before you accept anything, you must be able to explain what it does, why it’s built that way, and what breaks if it’s wrong. If you can’t explain it, you’re not ready to ship it.

Built-in tutor modes (use them)#

Copy-paste tutor prompt#

Paste this into Claude Project instructions or ChatGPT custom instructions:

You are my tutor, not my assistant. I'm learning [topic] from zero.
- Never give me full solutions unless I say "show me".
- When I'm stuck, ask me a guiding question first.
- Explain with analogies from [my background: marketing / business / etc].
- After every concept, give me one small exercise to do in my own terminal.
- When I show you my code, point out what's wrong but make me fix it.
- End each session by quizzing me with 3 questions.
- If you're unsure about a version, flag, or API, say so. Don't guess.

Seven habits that turn AI use into skill#

  1. You think, AI executes. You decide the what and the why. AI handles the how (the syntax). Code you can’t explain teaches you very little. Code you understand is what builds the skill.
  2. Ask “why” three times before running anything you don’t understand.
  3. Break it on purpose. Ask: “Give me 3 ways this setup breaks. I’ll diagnose them.” For security people, this habit is the job.
  4. Read every diff an agent produces and explain it back in plain words. If you can’t, don’t merge it.
  5. Use AI to read real code. Clone a serious repo and ask an agent: “Explain the architecture. Where does a request enter? What would you change first?” It’s one of the best shortcuts there is. Juniors used to need months of mentorship to get this.
  6. Don’t trust what the AI remembers about versions. Paste the official docs in or point the agent at them. Models are often wrong about flags and APIs that changed after they were trained.
  7. Keep an AI Engineering Journal (template in section 15). An AI summary of your own week doesn’t count.

Prompts that make AI a better teacher#

Watch first#

VideoWhy
ChatGPT Study Mode, Explained by a Learning Expert (Justin Sung, 21 min)How to actually learn with AI instead of outsourcing your thinking
Everything You Need to Know About Coding with AI // NOT vibe coding (ForrestKnight, 13 min)The difference between using AI and depending on it
How I use LLMs (Andrej Karpathy, 2h)How one of the field’s top people uses these tools day to day
Software Is Changing (Again) (Karpathy at YC, 40 min)The big picture: why this is a new kind of programming
From Vibe Coding to Agentic Engineering (Karpathy at Sequoia, 30 min)Where things are heading in 2026

4. The mindset: think in systems, solve business problems#

Tools change every month. The thing that makes you valuable doesn’t: being a systems thinker who solves real problems for businesses. The rest of this path is how you build that.

What systems thinking means#

A systems thinker doesn’t look at a task. They look at the whole flow around it: where information comes from, what happens to it, where it goes, who touches it, and where it breaks.

Every business process, and every AI system, comes down to the same loop:

StageWhat happensExample: a clinic’s booking requests
1. InputData comes in, usually messyCustomers send booking requests over Messenger in free text
2. ProcessIt gets cleaned, understood, decided on, transformedAI extracts name, service, date, and time; checks the calendar for conflicts
3. OutputA structured result someone (or another system) can act onA draft booking in the calendar, flagged for staff to confirm
4. FeedbackYou check whether it worked and feed that back inTrack how many drafts staff had to fix; improve the extraction where it fails

Questions a systems thinker asks:

Learn it: Thinking in Systems: A Primer by Donella Meadows is the classic, short and readable. Her free essay Leverage Points: Places to Intervene in a System is the best 30 minutes you can spend on this.

From systems thinking to solving business problems#

Businesses don’t pay for AI. They pay for problems to go away: hours lost, errors made, leads dropped, customers waiting. AI is just one way to fix that, and sometimes the right fix is a spreadsheet or a process change.

The playbook:

  1. Find the pain. Talk to the people doing the work. Look for tasks that are repetitive, slow, error-prone, or depend on one person who’s always overloaded.
  2. Map the system. Draw the input → process → output flow as it works today (free tools below). Most of the insight comes from this step alone.
  3. Find the bottleneck. Fix the step that costs the most, not the one that’s most fun to automate.
  4. Design the fix. Decide what should be automated, what AI should handle (reading messy text, classifying, summarizing, drafting), and what must stay human (judgment, approvals, relationships).
  5. Build the smallest version that works. n8n, Claude Code, an agent, whatever fits. Ship it to real users fast.
  6. Measure it. Hours saved, errors reduced, response time cut. Numbers are what turn a project into a case study and a client into a referral.

Examples of what this looks like:

Notice that none of these are “build an AI app.” They’re “make this specific painful thing go away.”

Drawing it: free tools for flowcharts and process maps#

You can’t fix a process you can’t see. Mapping it out is step 2 of the playbook, and it’s the most underrated skill in this whole path: it helps you think, it helps clients understand, and it makes your portfolio readable.

ToolBest forCost
ExcalidrawQuick hand-drawn-style flowcharts and system maps. Great for client calls and READMEs. It can also convert Mermaid code into an editable diagram.Free, open source (GitHub)
draw.io / diagrams.netFormal process maps, swimlanes, architecture diagramsFree, open source
Mermaid (live editor)Diagrams written as text. AI can write these for you, and GitHub renders them directly in READMEs.Free, open source
tldrawFast whiteboard sketchingFree tier

Process mapping basics:

AI shortcut: describe the process to ChatGPT or Claude and ask: “Turn this into a Mermaid flowchart with swimlanes for each role.” Paste the result into mermaid.live or Excalidraw, then fix it by hand. The fixing is where you actually understand the process.

Watch:

You own the output#

Having AI doesn’t make you good at what you do. It means you’re trusted to be responsible for whatever your AI produces. When an automation sends a wrong message to a customer, or an agent deletes the wrong file, or AI-written code leaks data, “the AI did it” isn’t an answer. The client hired you.

That’s why serious AI setups take time. In professional software development, nobody just types “build me an app.” The process looks like this:

StepWhat happensWhy it matters
1. PRD (Product Requirements Document)Write down the problem, the users, what the system must do, and what “done” looks likeAI builds exactly what you describe, so a vague request gets you vague software
2. BrainstormExplore approaches with AI, weigh the options, poke holes in the planIt’s cheaper to change a plan than code
3. DesignMap the system: data flow, tools, where humans stay in the loopThis is the systems thinking from this section
4. PlanBreak the work into small, testable tasksSmall steps are easy to verify and easy to undo
5. BuildAI writes code task by task; you review each pieceYou think, AI executes
6. TestCheck it works, including the edge cases and failure modesProof instead of hope
7. Review and secureRead the diffs, check for exposed secrets and risky permissionsYou own what ships
8. Deploy, document, monitorShip it, write it up, watch for problemsSomeone has to maintain it, possibly you at 2am

This approach is often called spec-driven development: write the spec first, then have AI build against it. The setup takes longer up front and saves you from rebuilding later.

Learn it:

The “build software everyone uses” dream (and survivorship bias)#

The path that looks most exciting right now is building a product, an app or SaaS that thousands of people pay for. It’s a legitimate goal, but go in with open eyes.

Why it’s harder than it looks:

Survivorship bias: You hear about the founder whose AI app made $50K a month. You don’t hear about the thousands who built the same kind of app and made nothing, because failures don’t post screenshots. Looking only at the winners makes success look like a recipe when it was partly timing, distribution, and luck.

This isn’t unique to AI or software. It’s true in most industries, from restaurants to music: not everyone succeeds, and the visible winners are a skewed sample. The point isn’t “don’t try.” It’s don’t plan your life around the survivors’ stories.

The smarter sequence:

  1. Solve problems for real businesses first. You get paid while you learn, you see real problems up close, and every project becomes proof of what you can do.
  2. Watch for patterns. When five different clients have the same problem, that’s a product idea with demand already proven.
  3. Then build the product, for customers you already understand and can already reach.

Keeping up without burning out#

AI moves fast: new models, harnesses, and “game-changing” tools every week. Nobody keeps up with all of it, and you don’t need to.

Sources for this section

5. Hardware: what you actually need#

You don’t need a GPU to start. Almost everything in the first ~4 months runs in the cloud.

TierSetupWhat it unlocks
0Any laptop with 8GB RAM and internetChat AIs, coding agents through APIs, n8n, Python. Enough for Phases 0–3.
116GB RAM, no GPUComfortable WSL2/Linux, Docker, one security VM (Kali) at a time, tiny local models on CPU (slow, but fine for learning)
2GPU with 8GB VRAM (RTX 3060 / 4060)Local 7–9B models at Q4 (e.g. Qwen3.5-9B). Good for privacy and offline experiments.
3GPU with 16GB VRAM (RTX 4060 Ti 16GB)20B-class models (gpt-oss-20b, Devstral Small 2), image generation (ComfyUI), usable local agents
4Used RTX 3090 (24GB), or a Mac with 32GB+ unified memory27–32B models (e.g. Qwen3.6-27B, ~77 on SWE-bench Verified). The classic budget-enthusiast upgrade.

Free GPU workarounds#

Honest take on local models#

Local models are for learning how models work, privacy, and offline use. They still don’t replace frontier models for serious agentic coding.

Gotcha: Ollama’s default context window is tiny (4k tokens under 24GB VRAM). Set OLLAMA_CONTEXT_LENGTH=65536 or higher before connecting a coding agent, or it will forget your files halfway through a task.

Security homelab#

VirtualBox (free), plus a Kali Linux VM, plus a deliberately vulnerable VM to attack. 16GB RAM makes this comfortable.

Watch#


6. Tools and budget stacks#

Dead or shrunk in 2026. Ignore old tutorials that recommend these.#

The coding agents (“harnesses”)#

A harness is the software that turns a model into an agent: it reads files, runs commands, edits code, and asks your permission along the way. The model is the brain; the harness gives it hands. Learn one harness well, then swap models underneath it. That way you’re not locked into one vendor’s pricing.

HarnessTypeWhy use it
DeepSeek Harness (dsh)Local web UI + headless modeDeepSeek’s official open-source harness (MIT), released Aug 2026. “Everything is a plugin.” Pairs with DeepSeek V4 Flash, which is very cheap, but it’s model-open: OpenRouter, Anthropic, OpenAI, or any OpenAI-compatible or local endpoint. Start it with npx @deepseek-ai/dsh web. It’s a developer preview, so expect breaking changes, and read its SAFETY.md first.
OpenCodeTerminal + desktopOpen source, supports 75+ providers. Its Zen provider includes free, coding-tested models (currently e.g. Big Pickle, DeepSeek V4 Flash, Nemotron 3 Ultra Free, MiMo-V2.5 Free; the list rotates). The best ₱0 starting point.
AiderTerminalGit-native: every AI change becomes a commit, so you’re always reading diffs, which suits learning well
Cline / Kilo Code / Roo CodeVS Code extensionsGood if you’d rather stay in an editor
Codex CLITerminalWorks when you sign in with a free ChatGPT account (limited), or with your own API key
Claude Code (docs)Terminal + desktopNeeds Claude Pro or the API. The current quality ceiling for agentic coding.

₱0 stack#

The best ₱600 you’ll spend: a one-time $10 OpenRouter top-up#

About ₱1,200/month: one frontier subscription (pick ONE)#

Prices are in USD. No credit card? A debit or virtual card from your e-wallet or digital bank works for most of these.

Privacy rule (non-negotiable, especially for security people)#

Free endpoints often log or train on your prompts. Never paste API keys, .env files, passwords, client data, or company code into a free model. Learning this habit early is already security practice.

Watch#

DeepSeek Harness (dsh)

OpenCode

OpenRouter and Aider


7. The path, phase by phase#

Don’t wait until the end to start earning. By month 3 you should be packaging your work and reaching out. See section 8 for the fast-track timeline.

Phase 0: Set up and learn to talk to AI (weeks 1–2)#

Do

For Filipinos: free AWS courses on TESDA. The TESDA Online Program now has an Amazon Web Services (AWS) course category with three courses: Fundamentals of Artificial Intelligence, AWS Certified Cloud Practitioner (CLF-C02), and AWS Certified AI Practitioner (AIF-C01). They look free to enroll in, but check the course pages for current details. The Cloud Practitioner course fits the DevOps track, and the AI Practitioner course is a solid structured intro to AI concepts and AWS’s AI services. (Note: the TESDA courses prepare you for the exams; check AWS for the actual exam fees.) I’m interested in these myself, especially for my DevOps track.

Watch

Proof you’re done: You use AI daily for real tasks, and you can explain tokens, context window, and hallucination in your own words.


Phase 1: Concepts, not syntax (weeks 3–6)#

You don’t need to master a programming language before you build. You need to understand the concepts well enough to direct AI, read what it produces, and know when it’s wrong. AI writes the skeleton; you have to understand the body.

The concepts to understand (be able to explain each one in plain words):

ConceptWhat you should be able to explain
Variables, functions, loops, conditionsWhat a program is doing, step by step, when you read it
Data structures (lists, objects/dictionaries) and JSONHow data is shaped and passed around
Files, the terminal, environment variablesWhere things live and how programs find them (and why secrets go in .env, never in code)
APIs and HTTP (requests, responses, status codes)How two systems talk; what a 401, 404, or 500 means
Client vs. server, frontend vs. backendWhich part runs where
Databases (tables, queries)Where data is stored and how it’s retrieved
Errors and logsHow to read an error message and find where it came from
Git (commits, branches, diffs)How to track changes and undo mistakes

How to learn them: Have AI build small real things with you in pair mode, then make it explain every part. Ask: “Walk me through this line by line. What concept is each part using? What would break if I removed it?” Change one thing yourself and predict what happens before you run it.

Do

Watch

Proof you’re done: You’ve reached Bandit level 15+, you’ve built a small script with AI that’s actually useful to you, and you can explain every part of it, plus every concept in the table, without looking anything up.


Phase 2: Using LLMs from code and automation (weeks 7–10)#

Do

Watch

Proof you’re done: A chatbot that calls at least one tool (search, weather, a database) and a working n8n workflow that saves you real time.


Phase 3: Agents, RAG, MCP, and agentic coding (weeks 11–16)#

Do

Watch

Proof you’re done: A deployed RAG or agent app, with evals, and a README a stranger could follow.


Parallel track (weekends, optional): how LLMs work under the hood#

You don’t need this to get hired. It’s what separates people who understand AI from people who just use it. Read section 2 first for the big picture.

Watch, in order

  1. But What Is a Neural Network? → Gradient Descent → Backpropagation (3Blue1Brown)
  2. Transformers, the Tech Behind LLMs → Attention in Transformers, Step by Step (3Blue1Brown)
  3. Deep Dive into LLMs like ChatGPT (Karpathy, 3.5h, the best single explainer)
  4. Let’s Build GPT: From Scratch, in Code, Spelled Out (Karpathy) · Let’s Build the GPT Tokenizer
  5. How Transformer LLMs Work (free course) (Jay Alammar)

Code along: karpathy/nn-zero-to-hero, karpathy/nanoGPT, rasbt/LLMs-from-scratch


Phase 4: Specialize (weeks 17–24+)#

This is where your specialization track starts (see section 9 for all the tracks). The AI + Security track is laid out in full below as an example.

The AI + Security track

This field is short on people. Most security people don’t understand LLMs, and most AI builders don’t think about attacks.

4a. Security fundamentals (you can’t secure what you don’t understand)

Watch

4b. LLM and agent security

Watch

4c. The portfolio piece that gets attention

Point garak and promptfoo at the chatbot or agent you built in Phase 3. Find the vulnerabilities, fix them, and publish a redacted red-team report: findings, severity (critical/high/medium/low), fixes, and before/after results.

“I built an AI app, attacked it, and hardened it” is a portfolio most applicants don’t have.

Legal line: Only attack systems you own or have written permission to test. Unauthorized access is a crime under the PH Cybercrime Prevention Act (RA 10175). CTFs, labs, and your own apps are fair game.


8. Landing your first client or job#

The first paying client or job is where all of this starts to make sense. Real constraints, real feedback, real money, and proof that you can deliver. Don’t wait until you finish the whole path. Start reaching out by month 3, sooner if an opportunity shows up.

The fast track: earn while you learn#

WhenFocusWhat you should have
Weeks 1–2Phase 0: AI basics, tools set upYou use AI daily; your tutor setup works
Weeks 3–6Phase 1 concepts, plus your first automation for yourselfOne working system that saves you time
Weeks 7–10Phase 2: n8n + AI integrations, and build something for a real person (a friend’s business, a relative’s shop, even for free)2–3 working systems, one used by a real business
Month 3: the milestonePackage your work and start applying / reaching outA GitHub portfolio with 3+ systems, 2 written case studies with diagrams, a short demo video of each
Months 4–6Phase 3 (agents, RAG, MCP) and Phase 4 (your specialization track) while doing paid workYour first paid client or job, and harder projects in your portfolio

This can go faster. If someone you know has a painful problem in week 4, solve it in week 4. The phases are a learning order, not a waiting room.

What you can realistically offer after ~3 months#

You won’t be a senior engineer. You will be able to solve real, common problems that small businesses pay for:

Every one of these is the same loop from section 4: messy input → process → structured output.

Build a portfolio that gets you hired#

Clients and employers don’t hire certificates. They hire proof you’ve solved a problem like theirs.

For each system you build, publish a GitHub repo with a README that covers:

  1. The problem: who had it, and what it cost them (time, money, missed leads)
  2. The system map: a before/after diagram (see the diagram tools in section 4)
  3. How it works: the flow, the tools, and where AI is used and where humans stay in the loop
  4. Results: numbers wherever possible (hours saved, response time, error rate)
  5. What I’d improve: this shows judgment, and it’s what interviewers ask about
  6. Tech used

This is your AI Engineering Journal (section 15) turned outward. Add a 2-minute demo video (screen recording with your voice) for each one; a free tool like Loom works.

Then share it: post short write-ups on LinkedIn and in relevant Facebook groups for business owners. Focus on the problem solved, not the tech.

Where to find your first client or job#

Job titles to search: AI Automation Specialist, Automation Engineer (n8n / Make / Zapier), AI Implementation Specialist, AI Operations, Technical VA (AI/automation), Junior AI Engineer, Solutions Engineer. For the security track: SOC Analyst (Tier 1), Security Analyst, and later AI Security / AI Red Team roles.

How to land the first client#

  1. Start with a discovery conversation, not a pitch. Ask the systems-thinking questions from section 4: What’s eating your team’s time? Where do leads or orders get lost? What do you do by hand every day?
  2. Map their process. Draw the current flow in Excalidraw, mark the bottleneck, and show them a “before and after.” A clear diagram often sells the project better than any demo.
  3. Propose a small pilot. One problem, one system, a short timeline, a fixed price. Small and specific beats big and vague.
  4. Deliver, measure, and document. Track the before/after numbers from day one.
  5. Ask for a testimonial and permission to use it as a case study. That’s what gets you client number two.

On pricing: your first one or two projects can be cheap or even free, in exchange for a testimonial, a referral, and permission to publish the case study. After that, charge properly. Price by the value of the problem solved, not your hours. Raise your rates after every few successful projects.

Applying for jobs#

Using your skills and getting better#

Landing the first client isn’t the finish line. It’s where the compounding starts:

LoopWhat it means in practice
Build → ShipEvery project goes to real users, even if it’s small
DocumentJournal entry and a case study for every project
SharePost what you solved; that’s how the next client finds you
ReuseTurn what you built into templates and reusable workflows, so project #5 takes a fraction of the time of project #1
Level upTake on one thing harder than last time with each new project: agents, RAG, security, bigger clients
ReviewAfter every project: what went well, what broke, what you’d charge next time

9. Pick a specialization track#

The core path (sections 3–8) gives you the foundation: learning with AI, systems thinking, data handling, automation, and agents. Then you specialize. Specialists get hired faster and paid more than generalists, because businesses search for “someone who can fix our CRM” or “someone who can secure our AI,” not “someone who knows AI.”

You can pick one track, combine two, or switch later. The core skills carry over to all of them. Phase 4 of the path is where your track starts.

The tracks at a glance#

TrackYou’ll be the person who…Common job titles
1. AI Automation & IntegrationsConnects a business’s tools and puts AI into its workflowsAI Automation Specialist, Automation Engineer, AI Implementation Specialist
2. CRM Management, Automation & DevelopmentRuns, automates, and extends the system where a business tracks its customersCRM Specialist, CRM Automation Specialist, HubSpot / Salesforce / GoHighLevel Admin or Developer
3. GTM EngineeringBuilds the data and automation behind sales outreachGTM Engineer, RevOps / Growth Engineer
4. AI + SecurityFinds and fixes weaknesses in apps and AI systemsSOC Analyst, Security Analyst, AI Security / AI Red Team
5. DevOps / DevSecOps (my track)Ships, runs, monitors, and secures software in productionDevOps Engineer, Platform Engineer, SRE, DevSecOps Engineer
6. AI-Assisted Software Development (my track)Builds real software with AI agents, professionallySoftware Engineer, AI Engineer, Full-Stack Developer
7. Other tracksVoice AI, AI content and creative, data and analytics, vertical (industry) AIVaries

Track 1: AI Automation & Integrations#

This is the default track the core path already builds toward: n8n, AI integrations, agents, MCP, and RAG for businesses.

Track 2: CRM Management, Automation & Development#

Every business that sells something has a CRM (Customer Relationship Management system), or a messy spreadsheet that should be one. CRM work is steady, in demand, and a natural home for AI automation.

Three levels of the same track:

LevelWhat you doExample tasks
CRM managementSet up and run the CRM so the team actually uses itPipelines, deal stages, contact properties, permissions, data cleanup, reports and dashboards
CRM automationAutomate what happens inside and around the CRMLead routing, follow-up sequences, appointment reminders, AI that qualifies leads or summarizes calls, syncing with forms and Messenger via n8n
CRM developmentExtend the CRM with code and integrationsCustom objects, APIs and webhooks, custom apps, connecting the CRM to other systems, data migrations

The main platforms:

Why AI makes this track stronger: CRMs are full of messy data and repetitive steps, which is exactly the “data in → process → structured data out” loop. AI can enrich contacts, score leads, summarize conversations, and draft follow-ups, with the CRM as the system of record.

Watch:

Track 3: GTM Engineering#

GTM (go-to-market) engineering treats sales outreach as a system to be engineered. A GTM engineer builds data pipelines, enrichment workflows, and automated outbound sequences, work that used to take a team of researchers and sales assistants. The role grew up around Clay, a tool for pulling data from many sources, enriching it with AI, and feeding it into outreach and the CRM.

Typical work: finding target companies and contacts, enriching them with data (company size, tech used, recent news), writing personalized outreach with AI at scale, and syncing everything into the CRM.

I tried this track myself (November–December 2025): I used AI and tools like Clay to build outbound outreach systems for businesses. It’s a strong, well-paid track. I didn’t pursue it fully because I prefer working on the sidelines, building systems rather than being close to the sales front line. That’s a good example of how to choose: try a track on a real project before committing to it.

Watch out for: spam and privacy laws. Mass outreach has to follow rules like the PH Data Privacy Act (RA 10173), GDPR in Europe, and CAN-SPAM in the US. Good GTM engineering is targeted and relevant, not mass spam.

Learn it: Clay University (free)

Watch:

Track 4: AI + Security#

Covered in full in Phase 4 of section 7: security fundamentals, LLM and agent security, open-source red-teaming tools, and a red-team report on your own app as the portfolio piece.

The OWASP references every security person should know (OWASP is the nonprofit that publishes the industry’s standard security guides, all free):

ResourceWhat it’s for
OWASP Top 10The ten most critical web application security risks. The baseline.
OWASP Top 10 for LLM ApplicationsThe biggest risks in apps built on language models: prompt injection, data leakage, excessive agency, and more
OWASP Top 10 for Agentic Applications (2026)Released December 2025: risks for AI agents that plan, use tools, keep memory, and act on their own, like goal hijacking, tool misuse, and rogue agents. Built from real 2025 incidents.
OWASP Juice ShopA deliberately insecure web app to practice hacking on, legally
OWASP Cheat Sheet SeriesShort, practical “how to do this securely” guides
OWASP ASVSA checklist for verifying an application’s security, level by level

Track 5: DevOps / DevSecOps (my track)#

DevOps is about getting software into production and keeping it running: deploying it, monitoring it, fixing it, and keeping the cloud bill honest. DevSecOps builds security into every step instead of bolting it on at the end. AI can write a lot of infrastructure code, but it can’t own an outage at 2am. Someone has to understand the system, and that’s this track.

What you learn, roughly in order:

  1. Linux and the terminal
  2. Networking (DNS, HTTP, ports, TLS)
  3. Git and scripting
  4. Containers (Docker)
  5. CI/CD pipelines (GitHub Actions)
  6. Cloud fundamentals (AWS, Azure, or GCP)
  7. Infrastructure as Code (Terraform)
  8. Kubernetes (official tutorials)
  9. Observability (logs, metrics, alerts)
  10. Security: secrets management, least privilege, scanning, using the OWASP DevSecOps Guideline

Map: roadmap.sh/devops

Free structured option: the AWS courses on TESDA, especially AWS Certified Cloud Practitioner (CLF-C02), a good first cloud certification for this track. It’s on my own list too.

Watch:

Track 6: AI-Assisted Software Development (my track)#

This is building real, production software with AI agents, the professional version of “vibe coding.” The difference is the process: specs, plans, tests, reviews, and ownership of what ships (see “You own the output” in section 4).

What you learn:

Maps: roadmap.sh/full-stack, roadmap.sh/backend, roadmap.sh/ai-engineer

Pairs well with: DevOps (Track 5). Building software and knowing how to ship and run it is a strong combination, and it’s the one I chose.

Track 7: Other tracks worth knowing#

TrackWhat it isWhere to start
Voice AIPhone and voice agents: receptionists, call handling, dictationVoice agent platforms plus n8n; Phase 3 agent skills
AI content and creativeImage, video, and audio generation pipelines for brandsComfyUI and the local-model setup in section 5
Data and analyticsTurning business data into dashboards, reports, and decisionsSQL, spreadsheets, and AI-assisted analysis; roadmap.sh/ai-data-scientist
Vertical (industry) AIBecoming the AI person for one industry: clinics, real estate, legal, e-commercePick an industry you know, and solve its most common problem again and again

How to choose#


10. Roadmaps (roadmap.sh)#

Use these as maps to see where you are, not as checklists to finish. Each node links to free resources.

RoadmapUse it for
AI EngineerThe main map for this whole path
AI AgentsPhase 3 depth
Prompt EngineeringPhase 0 depth
AI Red TeamingPhase 4b. The closest match to AI + Security.
Cyber SecurityPhase 4a fundamentals
Linux · Git & GitHub · PythonPhase 1
DevOpsIf you lean toward DevSecOps later
MLOps · AI & Data ScientistIf you go deeper into the ML side
Computer ScienceFilling gaps later, not now

11. Books#

About “free PDF” links: I’m not linking pirated copies. Most of those sites are unofficial uploads, and pirated-PDF sites are a common way people get malware, which is a bad habit for someone heading into security. Every option below is legally free, or an official free companion to a paid book, and in practice the companion notebooks are where most of the learning happens anyway.

O’Reilly books with free official companions#

BookFree partPhase
Hands-On Large Language Models (Alammar & Grootendorst)All the code notebooks: HandsOnLLM/Hands-On-Large-Language-Models, which run in Colab3 / parallel
Deep Learning for Coders with fastai and PyTorch (Howard & Gugger)The full book draft as notebooks: fastai/fastbook, plus the free fast.ai courseParallel
AI Engineering (Chip Huyen)Companion resources and reading lists: chiphuyen/aie-book. This is the best book on building with foundation models.3
Designing Machine Learning Systems (Chip Huyen)Companion repo: chiphuyen/dmls-bookLater
Hands-On Machine Learning, 3rd ed. (Géron)All notebooks: ageron/handson-ml3Parallel
Natural Language Processing with Transformers (Tunstall et al.)All notebooks: nlp-with-transformers/notebooksParallel
The Developer’s Playbook for Large Language Model Security (Steve Wilson, OWASP LLM Top 10 lead)Paid. The OWASP LLM Top 10 docs above are the free version of its core.4b

Legal shortcut: O’Reilly Learning has a free trial. Plan a reading sprint around it (e.g. AI Engineering plus the LLM security playbook) and cancel before it bills.

BookLinkPhase
Automate the Boring Stuff with Pythonautomatetheboringstuff.com1
The Linux Command Line (Shotts)linuxcommand.org/tlcl.php1
Pro Gitgit-scm.com/book1
Understanding Deep Learning (Prince, MIT Press)udlbook.github.io/udlbookParallel
Dive into Deep Learningd2l.aiParallel
Neural Networks and Deep Learning (Nielsen)neuralnetworksanddeeplearning.comParallel
Site Reliability Engineering (Google)sre.google/booksDevSecOps later
Build a Large Language Model (From Scratch) (Raschka, Manning)Book is paid. All code is free: rasbt/LLMs-from-scratchParallel

Essays worth more than most books#


12. GitHub repos to work with#

RepoWhat it’s for
deepseek-ai/deepseek-harnessdsh itself. Read AGENTS.md and docs/architecture.md to learn how a real harness is built.
Aider-AI/aiderGit-native terminal coding agent
ollama/ollamaRun local models
n8n-io/n8nSelf-hostable workflow automation
anthropics/coursesPrompt engineering and tool-use courses
anthropics/claude-cookbooks · openai/openai-cookbookCopy-and-learn recipes
microsoft/generative-ai-for-beginners21-lesson GenAI course
microsoft/ai-agents-for-beginnersAgents course
dair-ai/Prompt-Engineering-GuidePrompting reference
karpathy/nn-zero-to-hero · karpathy/nanoGPTBuild it yourself to understand it
rasbt/LLMs-from-scratchLLM internals, step by step
HandsOnLLM/Hands-On-Large-Language-ModelsO’Reilly book notebooks
NVIDIA/garakLLM vulnerability scanner
promptfoo/promptfooEvals and red-teaming
Azure/PyRITAutomated AI red-teaming
UKGovernmentBEIS/inspect_aiEvaluation framework

How to use a repo to learn, not just to collect stars: clone it, open your agent, and ask “Walk me through this codebase like I’m new. What’s the entry point? What are the 3 most important files?” Then change one small thing yourself and see what breaks.


13. Shortcuts and workarounds#


14. Traps to avoid#


15. Daily routine, journal, and progress tracker#

Daily routine (~2 hrs/day)#

TimeWhatHow to use AI
20 minLearn theory: a concept, a video segment, a book sectionTutor mode: “explain it, then quiz me”
40 minFollow a tutorial or courseMake sure you can explain every step; ask “why” when something’s unclear
60 minBuild something yourself: your own project, not the tutorial’sPair mode: you drive, AI reviews
20 minDocument what you learned (in your journal, below)Write it yourself. AI can quiz you on it afterward.

That last part isn’t optional. Writing it down is how what you learned actually sticks. If you can’t explain it in writing, you haven’t learned it yet. Short on time? Cut the tutorial, never the build or the journal.

AI Engineering Journal#

Keep one entry per project (a Markdown file in your GitHub repo works well). Over 6 months, this journal becomes your portfolio and your interview answers.

## Project: [name]  —  [date]

What I wanted to build:
What I learned:
What went wrong:
How I fixed it:
What I would improve:
Technologies used:
AI tools used (and for what — tutor / pair / delegate):

Sample week, if you’d rather block hours (~10–12 hrs)#

DayWhat
Mon–Thu1–1.5 hrs of hands-on labs (AI open in tutor mode)
Fri1 hr of a video or book chapter, then have AI quiz you on it
Sat3 hrs of project work (the thing you’ll show people)
Sun30 min weekly review: read back your journal entries, then have AI quiz you on the week

Progress tracker#

Weekly review log#

WeekBuiltBrokeBelieve now
1
2
3

16. Glossary#

Plain-language definitions of the terms used in this guide, in alphabetical order.

TermDefinition
AgentAn AI system that can plan, use tools (search, code, APIs), check results, and keep going until a task is done, instead of only answering one message
AGENTS.md / CLAUDE.mdA file in a project that tells a coding agent how the codebase works, its rules, and its conventions
AGI (Artificial General Intelligence)A hypothetical AI that can do any intellectual task a human can. It doesn’t exist yet, and people disagree on when, or if, it will.
AI (Artificial Intelligence)The broad field of making computers do things that normally need human intelligence: understanding language, recognizing images, making decisions
AI slopLow-effort, generic AI-generated content published without real human input or editing
AlignmentMaking sure an AI system’s goals and behavior match what humans actually intend
API (Application Programming Interface)A defined way for two programs to talk to each other: one sends a request, the other sends back a response
API keyA secret password-like string that lets your code use a paid AI service. Never share it or paste it into public places.
AttentionThe transformer mechanism that lets every token weigh how relevant every other token is to it
BenchmarkA standard test used to compare models (e.g. SWE-bench for coding). Useful, but not the same as performance on your task.
BERTGoogle’s 2018 encoder-only transformer, built for understanding and classifying text rather than generating it
CFG (classifier-free guidance)In image generation, how strictly the output follows your prompt
Chain-of-thought (CoT)Getting a model to reason step by step before answering, either through prompting or, in reasoning models, through training
ChatbotA program you talk to in conversation. ChatGPT and Claude are chatbots built on LLMs.
CI/CDContinuous Integration / Continuous Delivery: automatically testing and deploying code every time it changes
CloudComputing resources (servers, storage, AI models) rented over the internet instead of run on your own machine (e.g. AWS, Azure, Google Cloud)
Context engineeringDeciding what information (documents, memory, tools, examples) goes into a model’s context window so it can do the job well
Context windowEverything a model can “see” at once: your prompt, the conversation so far, and any documents. Measured in tokens.
CRMCustomer Relationship Management system: where a business tracks contacts, leads, deals, and customer conversations
DatasetA collection of examples used to train or test a model
DecoderThe generating half of the transformer. Decoder-only models (GPT, Claude, Llama) write text one token at a time.
Deep learningMachine learning using neural networks with many layers. The approach behind almost all modern AI.
DenoisingThe diffusion process of removing noise step by step to reveal an image
Dense modelA model where every parameter works on every token (the opposite of a sparse/MoE model)
Diffusion modelA model that generates images (or video, audio) by starting from random noise and gradually refining it into a picture. Used by tools like ComfyUI.
Discriminative modelA model that returns a decision or label (spam / not spam) rather than generating new content. Classic ML and Jev work this way.
DiT (Diffusion Transformer)A diffusion model that uses a transformer as its backbone, as in Sora, Stable Diffusion 3, and Flux
Docker / containerA package that bundles an app with everything it needs to run, so it runs the same on any machine
EmbeddingA list of numbers that represents the meaning of a token, sentence, or document. Similar meanings have similar numbers.
EncoderThe understanding half of the transformer. Encoder-only models (BERT, ViT) read the whole input to understand or classify it.
EvalA test that measures how well an AI system performs, so you can prove it works instead of eyeballing it
Few-shot / zero-shot promptingGiving a model a few examples of what you want (few-shot), or none at all (zero-shot)
Fine-tuningFurther training an existing model on specific data to specialize it
Frontier modelThe most capable models available at a given time, from the leading labs
GAN (generative adversarial network)An older image-generation approach where a “forger” network and a “detective” network train against each other
Generative AIAI that creates new content (text, images, audio, video, code) rather than only classifying or predicting
GPUGraphics Processing Unit: the chip that trains and runs AI models fast, because it does many calculations in parallel
GTM engineeringGo-to-market engineering: building the data pipelines and automation behind sales outreach, often with tools like Clay
GuardrailsRules and filters that keep an AI system’s inputs and outputs safe and on-topic
HallucinationWhen a model produces confident output that’s false or made up, a side effect of predicting likely text rather than true text
HarnessThe software that turns a model into a working agent: it reads files, runs commands, edits code, and asks permission (e.g. Claude Code, dsh, OpenCode)
Hybrid architectureA model that mixes attention layers with faster layers such as SSMs (e.g. Jamba, Griffin)
IaC (Infrastructure as Code)Defining servers, networks, and cloud resources in code files (e.g. Terraform) instead of clicking through dashboards
InferenceRunning a trained model to get an output, as opposed to training it
JailbreakA prompt designed to trick a model into ignoring its safety rules
JSONA common text format for structured data, made of keys and values. The language most APIs and automations speak.
Knowledge cutoffThe date a model’s training data ends. It doesn’t know about events after that unless it can search or you give it the information.
LatencyHow long it takes to get a response
Latent spaceA compressed representation of data (like an image) that models work in because it’s much cheaper than raw pixels
LLM (Large Language Model)A large transformer trained on huge amounts of text to predict the next token (e.g. GPT, Claude, Gemini, DeepSeek)
Local modelA model you download and run on your own computer instead of through a cloud API
LoRAA small add-on file that teaches a model a specific style or subject without retraining the whole model
Machine learning (ML)Teaching computers to learn patterns from data instead of following hand-written rules
Mamba / SSM (state-space model)An architecture that reads in sequence and keeps a running summary, so its cost grows in a straight line with input length instead of quadratically
MCP (Model Context Protocol)An open standard for connecting AI agents to tools and data sources
Mixture of Experts (MoE)A sparse architecture where a router sends each token to only a few specialist sub-networks, giving big capacity at lower running cost
ModelThe trained AI system itself, the thing that takes input and produces output
MultimodalA model that handles more than one type of input or output: text, images, audio, video
n8nAn open-source, self-hostable workflow automation tool for connecting apps and adding AI to processes
Neural networkA model made of layers of simple connected units (“neurons”) that learn by adjusting the strength of their connections
O(n) / O(n²)Shorthand for how cost grows with input size. O(n): double the input, double the work. O(n²): double the input, four times the work.
Open sourceSoftware whose code is public and free to use, modify, and share (e.g. n8n, OpenCode, dsh)
Open-weight modelA model whose trained weights are published, so anyone can download and run it
OutputWhat the model gives back: text, an image, a decision, or code
OWASPThe Open Worldwide Application Security Project, a nonprofit that publishes free security standards like the OWASP Top 10
Parameters / weightsThe billions of numbers inside a model that were learned during training. They hold what the model “knows.”
PRD (Product Requirements Document)A document describing the problem, users, requirements, and what “done” means, written before building
PretrainingThe first, largest training stage, where a model learns language by predicting tokens across huge amounts of text
PromptThe instruction or question you give an AI model
Prompt engineeringThe skill of writing prompts that reliably get good results
Prompt injectionAn attack where malicious instructions hidden in input (a webpage, email, or document) hijack an AI system’s behavior
Quantization (e.g. Q4)Compressing a model’s numbers to lower precision so it fits in less memory, at a small cost in quality
RAG (Retrieval-Augmented Generation)Fetching relevant documents and giving them to the model at answer time, so it answers from your data instead of only its memory
Rate limitA cap on how many requests you can send to a service in a given time (e.g. 50 requests a day on free tiers)
Reasoning modelA model trained to think through long, structured reasoning before answering (e.g. o1, DeepSeek-R1, “extended thinking” modes)
Red teamingDeliberately attacking your own system to find weaknesses before real attackers do
RLHF (Reinforcement Learning from Human Feedback)Training a model using human ratings of its answers, so it becomes more helpful and follows instructions
SamplerIn diffusion, the method used to remove noise at each step
SamplingHow the model picks the next token from its probability list: the “dice roll”
SeedThe starting random noise (or random number) for a generation. Same seed + same settings = the same result.
Spec-driven developmentWriting a clear specification first, then having AI build against it in small, testable steps
SycophancyA model’s tendency to agree with and flatter the user, even when the user is wrong
System promptInstructions given to a model before the conversation that set its role, rules, and behavior
T5Google’s 2019 encoder-decoder transformer that treats every task as “text in, text out”
TemperatureA setting that controls randomness. Low = predictable and focused; high = more varied and creative.
Test-time computeLetting a model spend more computation “thinking” while answering, rather than only making the model bigger
TokenA small chunk of text (often part of a word) that models read and write. Pricing and context limits are counted in tokens.
Tool callingA model’s ability to request that a function or API be run (search, calculator, database) and use the result
Top-pA sampling setting that only lets the model choose from the smallest set of tokens that together make up probability p (e.g. the top 90%)
TrainingThe process of teaching a model by showing it huge amounts of data and adjusting its parameters
TransformerThe 2017 neural network architecture built on attention that powers nearly all modern AI models
VAEThe part of an image model that compresses images into latent space and turns them back into pixels
Vector databaseA database that stores embeddings and finds the most similar ones quickly. The search engine behind RAG.
Vibe codingBuilding software by describing what you want to AI and accepting its code without really reviewing it. Fine for prototypes, risky for anything real.
Vision Transformer (ViT)A transformer that splits images into patches and treats them like tokens
VRAMThe memory on a graphics card. It decides which local models you can run.
WebhookA URL that one app calls automatically when something happens, to trigger an action in another app
Workflow / automationA defined series of steps that runs automatically when triggered, like a form submission creating a CRM contact and sending a reply
Contents ↑