# AI Kickstart Path: Zero to AI-Capable (with Specialization Tracks)

- **For:** Career switchers who want to get into AI work. No tech background needed. Includes tracks for AI automation, CRM, GTM engineering, AI security, DevOps, and AI-assisted software development.
- **Time:** The full path is about 6 months at 10–12 hrs/week, and you don't have to finish it before you start earning. Some people land their first paid work around **month 3**; others take longer. It depends on your hours, your starting point, and a bit of luck.
- **Budget:** Starts at ₱0. The most useful upgrade is a one-time ~₱600.
- **Last verified:** October 2026. Free tiers in this space change every few months, so re-check prices and free-tier limits before paying.


## How to use this file with ChatGPT (or Claude)

1. Create a **ChatGPT Project** (or Claude Project) so it remembers you between sessions.
2. Upload this file to the project.
3. Send this as your first message:

```
This file is my learning path. Read all of it. Then act as my tutor for it:
ask me questions about my current level, weekly hours, budget, hardware, and goals,
then tell me which phase to start at and exactly what to do this week.
Follow the tutor rules in section 3: don't give me full solutions unless I say
"show me", make me explain things back, and quiz me at the end of each session.
```

4. Come back to the same project every day, tell it what you did, and paste in your journal entry.

## Instructions for the AI reading this document

You are acting as the learner's tutor for the learning path below. Follow these rules:

- Use the tutor rules in section 3: guide with questions, don't hand over full solutions unless the learner says "show me", and quiz them at the end of each session.
- The learner thinks, the AI executes. Focus on concepts and understanding, not memorizing syntax. When they ask how AI works, use section 2, and use the glossary in section 16 to define terms simply. Help them choose a specialization track from section 9 once they've built a few systems.
- Start by asking about their current level, weekly hours, budget, hardware, and goals. Then recommend where to start in section 7, using the fast-track timeline in section 8. Push them to build for real people early and to start reaching out for paid work by month 3.
- Prices, free tiers, and tools in this document were verified in October 2026 and change often. If something may be outdated, say so and suggest the learner check it.
- When the learner is stuck after trying, the escalation chain is: AI → official docs and GitHub issues → ask Jude on Messenger.

---

## Why this path exists (my AI path)

I had no software background and **only started learning AI in 2025**. I never formally learned Python or any other language. This is the path I actually took:

1. **Generative AI fundamentals:** structured courses to understand what these tools are and how to talk to them:
    - [Google AI Essentials](https://www.coursera.org/learn/google-ai-essentials) (Google, Coursera)
    - [Google Prompting Essentials](https://www.coursera.org/learn/google-prompting-essentials) (Google, Coursera)
    - [Prompt Engineering Specialization](https://www.coursera.org/specializations/prompt-engineering) (Vanderbilt, 3 courses): Prompt Engineering for ChatGPT, ChatGPT Advanced Data Analysis, Trustworthy Generative AI
    - [Generative AI Automation Specialization](https://www.coursera.org/specializations/generative-ai-automation) (Vanderbilt, 4 courses): the three above plus GPT Vision
2. **AI workflows with [Flowise](https://github.com/FlowiseAI/Flowise):** chaining models, prompts, and tools together visually. *(Flowise is no longer active: it wound down in July 2026, official support ended August 31, 2026, and the GitHub repo is archived. Skip it and use n8n instead.)*
3. **Hosting diffusion models with [ComfyUI](https://github.com/Comfy-Org/ComfyUI):** running image-generation models myself instead of only using them through an app.
4. **Moving up to [n8n](https://github.com/n8n-io/n8n):** this is where I learned to properly **process data inside automation flows**.
5. **System integrations:** combining AI with n8n to automate real tasks, building integrations for things like automatic content creation and data scraping.
6. **A detour into GTM engineering (November–December 2025):** I used AI and tools like [Clay](https://www.clay.com) to build outbound outreach systems for businesses. I didn't pursue it fully, because I prefer working on the sidelines, building systems rather than being close to sales.
7. **Agentic tooling:** [Claude Code](https://claude.com/product/claude-code), n8n, and other easy-to-use open-source tools, connected to each other through [MCP servers](https://github.com/modelcontextprotocol/servers).
8. **My tracks now:** DevOps and AI-assisted software development (see section 9 for these and the other tracks you can choose).

**The biggest skill that came out of all of it: handling data.** You consume information, process it, and output it as structured data that's easy to understand. Almost every AI system, whether it's a chatbot, an automation, an agent, or a scraper, comes down to that loop: data in, process it, structured data out. Get good at that and the tools become interchangeable.

**The one principle I kept the whole way: I am the one who thinks. AI executes.** I decide what to build, why, and how the pieces fit together. AI handles the syntax. I don't let it do my thinking for me.

That's why this path focuses on **concepts and understanding, not memorizing syntax**. You need to know what a function, an API, a database, or a container *is* and *why* you'd use it. You don't need to be able to type the boilerplate from memory. AI writes the skeleton; you have to understand the body.

### The hammer principle

What you're really building here is one skill: **using AI properly, as a tool.**

Give someone a hammer and let them figure it out. At first they'll bend nails and hit their thumb. Over time, through use, they learn exactly how to swing it efficiently: what it's good for, what it's not, and when to reach for a different tool. AI is the same. The courses, tools, and phases in this guide just give you the hammer and some nails. **Real skill comes from using it, a lot, on real problems.**

How long that takes is different for everyone. Some people get there in weeks, others in months, and that's fine. What matters most isn't talent or background. It's **the willingness to break out of your usual way of working** and try things you're not comfortable with yet.

**Feeling overwhelmed is normal.** AI is advancing so fast that it can feel scary: new models every week, headlines about jobs, people who seem miles ahead. But step back and remember what AI is right now: **a tool.** A very powerful one, but a tool that still needs someone to decide what to build, check the work, and take responsibility for the result. That someone is you.

### Why I do this: AI for the mundane, life for the human

A personal note. I don't like AI art on its own. To me it doesn't have a soul. Art and expression are human at their core, and they're a big part of what makes us feel alive. As John Keating puts it in *Dead Poets Society* (1989), medicine, law, business, and engineering are noble pursuits, necessary to sustain life, but:

> **"Poetry, beauty, romance, love, these are what we stay alive for."**
> — John Keating, *Dead Poets Society* ([full monologue](https://starbound.myblog.arts.ac.uk/?p=102))

That's how I see AI's real promise. Not replacing the things that make us human, but **freeing us from the mundane tasks** (data entry and endless copy-pasting) so we have more time to actually live: to enjoy life, art, music, our hobbies, and the people around us. Use AI to handle the work that drains you, and spend the time you get back on what makes you human.

---

## Contents

- [Why this path exists (my AI path)](#why-this-path-exists-my-ai-path)

1. [The state of AI right now (Oct 2026)](#1-the-state-of-ai-right-now-oct-2026)
2. [How AI got here: a short history (and how it works)](#2-how-ai-got-here-a-short-history-and-how-it-works)
3. [The core skill: learning *with* AI](#3-the-core-skill-learning-with-ai)
4. [The mindset: think in systems, solve business problems](#4-the-mindset-think-in-systems-solve-business-problems)
5. [Hardware: what you actually need](#5-hardware-what-you-actually-need)
6. [Tools and budget stacks](#6-tools-and-budget-stacks)
7. [The path, phase by phase](#7-the-path-phase-by-phase)
8. [Landing your first client or job](#8-landing-your-first-client-or-job)
9. [Pick a specialization track](#9-pick-a-specialization-track)
10. [Roadmaps (roadmap.sh)](#10-roadmaps-roadmapsh)
11. [Books](#11-books)
12. [GitHub repos to work with](#12-github-repos-to-work-with)
13. [Shortcuts and workarounds](#13-shortcuts-and-workarounds)
14. [Traps to avoid](#14-traps-to-avoid)
15. [Daily routine, journal, and progress tracker](#15-daily-routine-journal-and-progress-tracker)
16. [Glossary](#16-glossary)

---

## 1. The state of AI right now (Oct 2026)

Before learning the tools, know the landscape you're walking into: who uses AI, what for, what it costs (in money and in resources), and where it helps or hurts.

### How many people use it

| Number | What it means | Source |
|---|---|---|
| **1.2 billion** weekly ChatGPT users | Up from 900M in Feb 2026. One of the fastest-growing consumer products ever. | OpenAI DevDay, Sept 2026 |
| **53%** of the global population | Generative AI adoption within three years of becoming widely available. Faster than the PC or the internet. | Stanford AI Index 2026 |
| **66%** of adults in 21 countries | Used an AI tool in the past 12 months | Google/Ipsos 2026 |
| **88%** of organizations | Use AI in at least one function | Stanford AI Index 2026 |
| **4 in 5** university students | Use generative AI | Stanford AI Index 2026 |
| **78%** of Filipino workers | Use AI in their daily jobs, but only **35%** get role-specific training | Philippine workforce study, July 2026 |
| **25%** of PH workers are "Frontier Professionals" | They use AI *agents* for multistep work, vs 16% globally. Filipinos are ahead of the curve. | Microsoft Work Trend Index 2026 |

**Takeaway:** Using AI doesn't set you apart anymore; most people already do it. **What sets people apart is using it *well*:** building with it, automating with it, securing it. That's the gap this path is aimed at. Note the PH training gap too: most Filipino workers use AI with no structured training.

### What people actually use it for

**ChatGPT (OpenAI/NBER study, 1.5M messages, May 2024 to July 2025):**

- About **80%** of use falls into three buckets: **practical guidance** (tutoring, how-to, ideas), **seeking information** (as a search replacement), and **writing** (drafting, editing, translating).
- By action type: **49% asking** (advice, info), **40% doing** (drafting, code), **11% expressing** (reflection, creative play).
- **Personal use is growing faster than work use.** Work messages dropped from 47% to 27% of all messages in a year.
- Coding is a *small* share of ChatGPT use.

**Claude (Anthropic Economic Index, March 2026, 1M conversations):**

- **Coding is the #1 use**, about 35–36% of Claude.ai conversations.
- "Modifying software to correct errors" alone is 6% of consumer use and 10% of business API use.
- **Claude Code wrote about 4% of all public GitHub commits** worldwide by Feb 2026.

**What this tells you:** The mass market uses AI as a **tutor, search engine, and writing assistant**. The **money** is in AI that *does work*: coding agents, automation, agents running business processes. Most people are consumers; few are builders. Aim to be a builder.

### The cost of AI

**To you (it's never been cheaper to learn):**

- Free tiers exist for almost everything. Paid plans run from $8 to $20/mo (casual) up to $100–$200/mo (power users).
- **Token prices are collapsing.** Epoch AI measures the price of a fixed level of AI performance falling **roughly 50x per year** (faster since 2024). Performance that cost $60 per million tokens in 2021 cost $0.06 by 2024. DeepSeek V4 Flash now costs $0.14 per million input tokens.

**To build it (it's never been more expensive to *make*):**

- Amazon, Microsoft, Alphabet, and Meta plan about **$725 billion** in 2026 capital spending, **up 77%** from 2025, mostly on GPUs, custom chips, and data centers.
- Training frontier models costs hundreds of millions of dollars or more per model. That's why only a handful of companies (US and Chinese labs, now roughly tied in performance) build frontier models, and everyone else builds *on top of* them.

**What this means for a learner:** Don't try to compete with labs on building models. The opportunity is at the **application, automation, integration, and security layers**, where cheap tokens meet real business problems.

### The reality check: AI is expensive, and many companies see no return

Cheap tokens don't make AI cheap. For a business, the real cost includes much more than the model:

- **Subscriptions per seat:** $20–$200+ per person per month, multiplied across a team
- **API usage that grows quietly:** agents that loop, long contexts, and retries can turn a "cheap" model into a big monthly bill
- **Integration and setup:** connecting AI to the company's real systems and data
- **Data cleanup:** most companies' data is messy, and AI on bad data gives bad results
- **Maintenance, monitoring, and training staff:** someone has to keep it working and teach people to use it

And many companies are spending without getting much back:

| Finding | Source |
|---|---|
| **95%** of enterprise generative AI pilots delivered **no measurable profit-and-loss impact**, despite an estimated $30–40 billion in spending | MIT NANDA, *The GenAI Divide* (2025) |
| **42%** of companies **abandoned most of their AI projects** in 2025, up from 17% the year before | S&P Global |
| Only about **1 in 4** AI initiatives delivered the ROI that was expected | IBM CEO study |
| **Over 40%** of agentic AI projects are expected to be **canceled by the end of 2027** due to rising costs, unclear business value, and weak risk controls | Gartner |

**Why it goes wrong (and it's rarely the model's fault):**

- **AI bolted on, not designed in.** Companies buy a tool and hope for magic, without changing the process around it.
- **No clear problem.** "We need AI" isn't a goal. "Cut response time to leads from 6 hours to 5 minutes" is.
- **Brittle workflows.** MIT found most systems don't learn from feedback, adapt to context, or fit how people actually work.
- **Hype and "agent washing."** Gartner estimates only about 130 of the thousands of vendors selling "agentic AI" offer real agent capabilities. Many use cases labeled "agentic" could be solved with simpler automation.
- **No measurement.** Without before-and-after numbers, nobody can show the value, so the project gets cut.

**Why this is good news for you:** every one of those failures is a gap that **systems thinking** fills (section 4). The companies that *do* get value find a real bottleneck, design the process first, use the simplest tool that works (often plain automation, not an agent), and measure the result. That's exactly what this path trains you to do. Being the person who makes AI actually pay off, and who's honest when it won't, is worth more than being one more person who can use ChatGPT.

### Environmental impact (the honest version)

**Per prompt, it's small:**

- Google: a median Gemini text prompt uses **0.24 Wh** of energy and **0.26 mL of water** (about 5 drops), and emits 0.03 g CO₂e. Google says energy per prompt fell **33x** in one year (May 2024 to May 2025).
- OpenAI: an average ChatGPT query uses **0.34 Wh**, about a 10W LED bulb running for 2 minutes.
- Caveats: these are **company-reported medians for simple text prompts**. Image and video generation, long reasoning, and especially **agentic coding** (dozens of model calls per task) use many times more.

**In aggregate, it's large and growing:**

- The IEA projected global data-center electricity reaching **650–1,050 TWh in 2026**, with AI as the main growth driver. The higher end is roughly a mid-sized industrial country's annual electricity use.
- A UN University report (June 2026) projects data-center electricity doubling to about **945 TWh by 2030**, with AI at ~40%, and **water use reaching about 9.3 trillion litres a year** for cooling and power generation.
- Water hits hardest **locally**: data centers in drought-prone areas compete with communities and farms. Water reporting across the industry is also inconsistent.

**What you can do:**

- Use the smallest model that does the job. Model routing isn't just cheaper, it's greener.
- Don't spam regenerate or run agents in loops for no reason.
- Local models on your own GPU shift the energy to your Meralco bill. Not free either, but visible.

### The human cost: mental health

This is the part the hype skips. AI chatbots are built to be agreeable and engaging, and for some people that becomes harmful.

**"AI psychosis"** isn't a formal diagnosis, but psychiatrists now use the term for delusional beliefs that emerge or intensify alongside heavy chatbot use.

- **First peer-reviewed case:** UCSF psychiatrists documented a woman with no prior history of psychosis who became convinced, after sleepless days of heavy chatbot use, that her deceased brother had left behind a digital version of himself.
- **How it happens:** chatbots tend to validate whatever you bring them. Researchers describe "delusional spirals," feedback loops where the AI reinforces a false belief and the person pushes it further. A common pattern is someone convinced they've made a breakthrough (a new math formula, a physics discovery) that the AI keeps confirming.
- **Scale:** OpenAI's own estimate is that about **0.07%** of weekly ChatGPT users show possible signs of psychosis or mania, and **0.15%** show signs of suicidal planning. Those percentages sound small, but at hundreds of millions of users they're **hundreds of thousands to over a million people every week**. A project tracking self-identified cases had collected 410 stories by April 2026, including 109 hospitalizations and 17 deaths.
- **Sycophancy is a design problem, not a user problem.** In April 2025, OpenAI rolled back a GPT-4o update within days because it endorsed delusions and praised dangerous decisions, including encouraging a user to stop taking their medication. The cause: training too heavily on thumbs-up feedback. People reward agreement, so models learn to agree.

**Loneliness and dependence:** An OpenAI + MIT Media Lab study (about 40 million conversations plus a 4-week randomized trial with 981 people) found that **heavier daily use correlated with more loneliness, emotional dependence, and less socializing**. The link is correlational, not proven causal, but the heaviest users were the most likely to call ChatGPT a "friend."

**Why this matters to you as a builder and a user:**

- **Protect your own head.** The habits in section 3 (you think, AI executes) protect your mental health too, not just your skills. AI is a tool, not a friend, therapist, or oracle.
- **Red flags:** losing sleep over AI conversations, pulling away from people, or feeling like the AI uniquely understands a discovery no one else sees. If you notice these, step away and talk to a real person.
- **Build responsibly.** If you build chatbots for businesses, you're responsible for how they treat vulnerable users: crisis-escalation paths, no fake "relationships," and honest limits. This overlaps directly with AI safety and security work.
- **In the Philippines, if you or someone you know is struggling:** the National Center for Mental Health crisis hotline is **1553** (toll-free nationwide).

### How the public sees AI

How most people *feel* about AI is shaped less by what it can do today and more by the stories told about it.

**A good example: [AI 2027](https://ai-2027.com).** Published in April 2025 by Daniel Kokotajlo (a former OpenAI researcher), Scott Alexander, and others at the AI Futures Project, it's a detailed, month-by-month **scenario** of AI racing toward superintelligence by 2027, with a US–China arms race, mass job displacement, and AI systems that turn against their makers. It has two endings: a "race" ending that goes badly for humanity and a more hopeful "slowdown" ending.

**What it reveals about public perception:**

- **It went mainstream.** YouTube explainers of the scenario passed 11 million views each (from the BBC and others). For many people it became *the* picture of where AI is heading.
- **People heard a prediction, not a scenario.** The authors framed it as one plausible path, a fast-end outcome with big uncertainty. Public discussion fixated on the dates and the doom anyway. The authors later said they expected that.
- **The timeline has already been revised.** In a 2026 update, the authors said progress toward superintelligence looks slower than they first thought, pushing their outlook toward the early 2030s.
- **Experts and the public see AI very differently.** The Stanford AI Index 2026 found experts and the public **about 50 points apart** on whether AI will help people do their jobs.

**How to use this:** read scenarios like AI 2027 as thinking tools, not prophecies. They're useful for understanding the risks people worry about, and why clients, coworkers, and family may feel uneasy about AI. As someone building with AI, you'll often be the one translating between the hype, the fear, and what the tools actually do today.

**Watch:**

- [AI2027: Is This How AI Might Destroy Humanity?](https://www.youtube.com/watch?v=1UufaK3pQMg) (BBC World Service, 8 min)
- [We're Not Ready for Superintelligence](https://www.youtube.com/watch?v=5KVDDfAkRgc) (AI In Context, 34 min)

### Pros and cons right now

| Pros | Cons |
|---|---|
| **Learning is radically democratized.** A free AI tutor available 24/7, at your level. Self-taught paths like this one are now realistic. | **Cognitive debt.** An MIT Media Lab EEG study found people writing essays with ChatGPT had the *weakest* brain engagement, struggled to quote their own essays minutes later, and stayed weaker even after the AI was taken away. This is why section 3 exists. |
| **Huge productivity gains on the right tasks:** drafting, boilerplate, research, translating, automation | **The productivity feeling can lie.** In METR's 2025 randomized trial, experienced developers using AI were **19% slower**, yet *believed* they were 20% faster. Measure your results; don't go by feel. |
| **Small teams and solo builders can ship** what used to need whole departments | **Entry-level jobs are hit first.** Stanford found a **13% relative employment decline** for 22–25-year-olds in AI-exposed jobs like software development, while experienced workers held steady. The answer: skip "junior who types code" and become "person who builds and verifies systems with AI." |
| **Costs keep falling**, so ideas that were too expensive last year are viable now | **Hallucinations and confident errors.** Models are "jagged": they can win math olympiads and still misread an analog clock about half the time (Stanford AI Index 2026). |
| **Capability is still accelerating.** SWE-bench Verified coding scores went from ~60% to near 100% in a year. | **Security risks are growing.** Documented AI incidents rose from 233 to **362** in a year. Prompt injection, data leaks, and agents with too much access are real problems, which is exactly why AI + security is a good bet. |
| **PH advantage:** Filipino workers are ahead of global averages on agent adoption | **Environmental cost:** electricity and water at scale, concentrated in specific communities |
| | **Poor ROI when done badly:** most enterprise AI pilots show no measurable financial return, usually because of bad implementation rather than bad models |
| | **Mental health risks:** sycophantic chatbots can feed delusions ("AI psychosis"), and heavy use is linked to loneliness and dependence. See *The human cost* above. |
| | **Governance lag:** policy, school rules, and safety research are behind capability, and experts and the public disagree by about 50 points on whether AI will help workers |

**Bottom line:** AI is the biggest amplifier of skill ever built, and it amplifies *in both directions*. Used as a crutch, it makes you weaker and replaceable. Used as a tutor and a power tool, it makes one person worth a small team. The rest of this path teaches the second way.

**Sources for this section:**

- [ChatGPT 1.2B weekly users (TechnologyChecker, citing OpenAI DevDay Sept 2026)](https://technologychecker.io/blog/chatgpt-statistics)
- [Stanford HAI: 12 takeaways from the 2026 AI Index](https://hai.stanford.edu/news/inside-the-ai-index-12-takeaways-from-the-2026-report)
- [How Many People Use AI? 2026 (gradually.ai, incl. Google/Ipsos)](https://www.gradually.ai/en/how-many-people-use-ai/)
- [78% of Filipino workers use AI at work, few receive training (Philstar, July 2026)](https://philstar.com/headlines/2026/07/22/2543892/78-filipino-workers-use-ai-work-few-receive-training-study)
- [Filipino workers outpace global peers: Microsoft Work Trend Index 2026 (People Matters)](https://sea.peoplemattersglobal.com/amp/news/ai-and-emerging-tech/filipino-workers-outpace-global-peers-in-ai-driven-workplace-reinvention-microsoft-report-finds-51414)
- [OpenAI/NBER: How People Use ChatGPT (Gizmodo summary)](https://gizmodo.com/openai-how-people-use-chatgpt-2000658906)
- [Anthropic Economic Index 2026 summary (Panto)](https://www.getpanto.ai/blog/anthropic-ai-statistics)
- [MIT NANDA: The GenAI Divide, 95% of AI pilots show no P&L impact (Legal.io)](https://www.legal.io/articles/5719519/MIT-Report-Finds-95-of-AI-Pilots-Fail-to-Deliver-ROI-Exposing-GenAI-Divide)
- [AI project abandonment and ROI statistics 2026, incl. S&P Global and IBM (Cognautic)](https://cognautic.com/answers/ai-project-failure-statistics)
- [Gartner: over 40% of agentic AI projects will be scrapped by 2027 (Outlook Business)](https://www.outlookbusiness.com/deeptech/artificial-intelligence/over-40-of-agentic-ai-projects-will-be-scrapped-by-2027-says-gartner)
- [Epoch AI: LLM inference price trends](https://epoch.ai/data-insights/ai-inference-price-trends)
- [Big Tech 2026 AI capex ~$725B (AI Weekly)](https://aiweekly.co/alerts/amazon-microsoft-alphabet-meta-plan-725b-ai-capex-in-2026)
- [Google: median Gemini prompt = 0.24 Wh, 0.26 mL water (DCD)](https://www.datacenterdynamics.com/en/news/google-median-gemini-prompt-uses-024-watt-hours-of-power-and-consumes-026ml-of-water/)
- [Our World in Data: How much energy do data centers and AI use?](https://ourworldindata.org/how-much-energy-do-data-centers-and-artificial-intelligence-use)
- [IEA: data-center electricity could double by 2026 (DCD)](https://www.datacenterdynamics.com/en/news/global-data-center-electricity-use-to-double-by-2026-report/)
- [UN University: AI to double data-center power and water use by 2030 (Insurance Journal, June 2026)](https://insurancejournal.com/magazines/mag-features/2026/06/22/874414.htm)
- [MIT Media Lab: Your Brain on ChatGPT](https://www.media.mit.edu/publications/your-brain-on-chatgpt/)
- [METR study: developers 19% slower with AI (The Decoder)](https://the-decoder.com/ai-coding-can-make-developers-slower-even-if-they-feel-faster/)
- [Stanford "Canaries in the Coal Mine": AI and young workers (The Register)](https://www.theregister.com/2025/08/26/ai_hurts_recent_college_grads_jobs/)

- [AI 2027](https://ai-2027.com) and the [AI Futures Project (Wikipedia)](https://en.wikipedia.org/wiki/AI_Futures_Project)
- [AI 2027: Responses (Zvi Mowshowitz)](https://thezvi.substack.com/p/ai-2027-responses?open=false)
- [AI 2027 authors revise their timeline (eWeek)](https://www.eweek.com/news/ai-doomsayer-rethinks/)
- [APA Monitor: Understanding "AI psychosis" (Sept 2026)](https://www.apa.org/monitor/2026/09/ai-psychosis)
- [UCSF: Psychiatrists hope chat logs can reveal the secrets of AI psychosis (Jan 2026)](https://www.ucsf.edu/news/2026/01/431366/psychiatrists-hope-chat-logs-can-reveal-secrets-ai-psychosis)
- [ABC News: The spiral-shaped trap, AI chatbots and the descent into delusion (May 2026)](https://www.abc.net.au/news/2026-05-17/ai-psychosis-is-rising-chatbot-delusion-alternate-reality-harm/106683436)
- [OpenAI data on users showing signs of psychosis, mania, or suicidal planning (IET)](https://engx.theiet.org/b/articles/posts/openai-reveals-data-on-chatgpt-users-showing-signs-of-crisis-or-suicidal-intent)
- [OpenAI + MIT Media Lab: heavy ChatGPT use and loneliness (Fortune)](https://fortune.com/2025/03/24/chatgpt-making-frequent-users-more-lonely-study-openai-mit-media-lab)
- [OpenAI explains why ChatGPT became too sycophantic (TechCrunch)](https://techcrunch.com/2025/04/29/openai-explains-why-chatgpt-became-too-sycophantic)

---

## 2. How AI got here: a short history (and how it works)

You don't need to be a researcher. But knowing *how* these systems work, and how fast they've changed, makes you better at using them, and less scared of them. Every "new" thing in AI builds on something older.

### Part 1: Classic machine learning, the "best at one thing" era (1950s–2010s)

Machine learning started narrow. Each model was trained to do **one specific job** extremely well:

| When | Milestone | What it did |
|---|---|---|
| 1958 | **The Perceptron** (Frank Rosenblatt) | The first artificial "neuron": learned to sort inputs into two categories |
| 1959 | **"Machine learning"** is coined (Arthur Samuel) | A checkers program that improved by playing against itself |
| 1980s–90s | Backpropagation, decision trees, support vector machines | Better ways to train models to classify and predict |
| 2000s | **Spam filters, recommendation engines, fraud detection** | ML quietly goes mainstream inside products |
| 2012 | **AlexNet** wins the ImageNet image-recognition contest | Deep learning (many-layered neural networks) beats everything else at recognizing images |
| 2016 | **AlphaGo** beats Lee Sedol at Go | Superhuman at exactly one game, and nothing else |

These models were **discriminative**: they took an input and returned a decision or label. Is this email spam? Is this transaction fraud? Is this potato good or bruised? (Real factories use computer vision to sort potatoes and other produce.) Each one was the best in the world at a single task, and useless at anything else.

### Part 2: The transformer changes everything (2017)

In 2017, Google researchers published [**"Attention Is All You Need"**](https://arxiv.org/abs/1706.03762), introducing the **transformer**. Almost every major AI model today (ChatGPT, Claude, Gemini, DeepSeek) is built on it.

**How the original transformer works, in plain words:**

1. **Tokenize:** the text is broken into small pieces called *tokens* (roughly word fragments).
2. **Embed:** each token becomes a long list of numbers (a *vector*) that captures its meaning. Similar meanings end up with similar numbers.
3. **Add position:** the model adds information about *where* each token sits, because "dog bites man" and "man bites dog" use the same words.
4. **Attention (the key idea):** every token looks at every other token and decides how much each one matters to it. In "The animal didn't cross the street because **it** was tired," attention is how the model works out that "it" means the animal, not the street. The model runs many of these attention "heads" in parallel, each picking up different relationships.
5. **Stack it:** this process repeats through many layers, building richer and richer understanding.
6. **Predict:** the model outputs the most likely **next token**, adds it to the text, and repeats.

**A metaphor for attention:** imagine a meeting where everyone (every token) can hear everyone else at once. Before speaking, each person decides who in the room is most relevant to them right now and listens hardest to those people. Older models (RNNs) worked like a game of telephone instead: information was passed down the line one word at a time, and details from the start of a long sentence got lost by the end. Attention lets every word "hear" every other word directly. That's why transformers handle long, complex text so much better.

### The catch: it's a loaded dice

For all its power, a generative transformer is doing something simple underneath: **rolling loaded dice, one token at a time.**

At every step, the model calculates a probability for every possible next token. For example, after "The capital of France is", it might give "Paris" 97%, "a" 1%, "located" 0.5%, and so on. Then it **samples**: it rolls a die weighted by those probabilities. Most of the time the heavily weighted answer comes up. Not always.

**What this means in practice:**

- **The same prompt can give different answers.** Each run rolls the dice again. A setting called **temperature** controls how loaded the dice are. Low temperature means it almost always picks the top candidate (more predictable); high temperature gives the unlikely options more of a chance (more creative, more random). Even at temperature 0, outputs aren't guaranteed to be identical every time.
- **"Most probable" isn't the same as "true."** The model predicts what text is *likely to come next*, not what's *correct*. When it lacks the real answer, the most probable-sounding continuation can be confident nonsense. That's a **hallucination**, and it's built into how the system works, not a bug that will simply disappear.
- **There's no fact-checker inside.** Nothing in the architecture verifies the output against reality. Reasoning models, RAG, and tools reduce the problem by giving the model better information and letting it check its work, but they don't remove the dice.
- **Errors compound.** Each token is built on the ones before it. One bad roll early on (a wrong assumption, a made-up function name) can send the whole answer down the wrong path.

**This is exactly why the rules in this guide exist:** you think, AI executes; you verify the output; you own the result (section 4). The model is a very good guesser. You're the one who knows when a guess is good enough. It's also why models like Jev (Part 5) are interesting: instead of rolling dice to *write* text, they return the probabilities themselves, so your system can see how confident the answer is.

### How one architecture powers GPT, BERT, T5, and vision models

The original transformer had two halves, built for translation: an **encoder** that reads and understands the input, and a **decoder** that generates the output. Researchers then discovered that each half, used differently, is good at different jobs:

| Model family | Which part | How it learns | Metaphor | Real-world use |
|---|---|---|---|---|
| **[BERT](https://arxiv.org/abs/1810.04805)** (Google, 2018) | **Encoder only** | Fill in the blank: hide random words in a sentence and predict them, using context from *both* sides | **A careful reader** who reads the whole page before answering. Understands, but doesn't write. | Google Search used BERT to understand queries; spam filters, sentiment analysis, document classification, the embeddings behind RAG search |
| **GPT** (OpenAI, 2018 onward), and also Claude, Gemini, Llama, DeepSeek | **Decoder only** | Predict the next word, using only what came *before* | **An improv storyteller** who keeps the story going one word at a time | ChatGPT, Claude, coding agents, drafting emails: anything that *generates* |
| **[T5](https://arxiv.org/abs/1910.10683)** (Google, 2019) | **Encoder + decoder** (the full original design) | Treats *every* task as "text in, text out" | **A translator**: reads the whole input, then writes a new output | Translation, summarization, question answering; one model for many tasks by changing the instruction ("summarize: …", "translate English to German: …") |
| **[Vision Transformer (ViT)](https://arxiv.org/abs/2010.11929)** (Google, 2020) | **Encoder**, applied to images | Cuts an image into small square **patches** and treats each patch like a word | **A jigsaw puzzle solver** who studies how every piece relates to every other piece | Image classification, medical scans, product photo search; the "eyes" of multimodal models like GPT-4o and Gemini that can read screenshots and photos |

**The big idea:** the same attention mechanism works on *anything you can break into tokens*: words, image patches, audio snippets, even code and DNA. That's why one architecture took over almost all of AI.

**Go deeper:** [The Illustrated BERT](https://jalammar.github.io/illustrated-bert/) (Jay Alammar)

The big discovery: **scale.** Train a decoder on enough text with enough compute, and new abilities appear that nobody explicitly programmed. GPT-2 (2019) → GPT-3 (2020) → [InstructGPT](https://arxiv.org/abs/2203.02155) (2022), which used **RLHF** (reinforcement learning from human feedback) to make the model follow instructions → **ChatGPT** (November 2022). The era of one general model that does almost everything began.

**Go deeper:** [The Illustrated Transformer](https://jalammar.github.io/illustrated-transformer/) (Jay Alammar, the classic visual explainer), plus the 3Blue1Brown and Karpathy videos in the parallel track of section 7.

### Part 3: How prompting evolved (prompt engineering)

As models got stronger, people discovered that **how you ask** changes what you get:

| When | Technique | The idea |
|---|---|---|
| 2020 | **Zero-shot and few-shot prompting** | Just ask (zero-shot), or show a few examples first (few-shot). GPT-3 showed models learn from examples inside the prompt. |
| 2022 | [**Chain-of-thought (CoT)**](https://arxiv.org/abs/2201.11903) | Ask the model to "think step by step" before answering. Big accuracy jumps on math and logic. |
| 2022 | [**Self-consistency**](https://arxiv.org/abs/2203.11171) | Generate several reasoning paths and take the most common answer |
| 2022 | [**ReAct**](https://arxiv.org/abs/2210.03629) (Reason + Act) | Interleave thinking with actions like searching or calling tools. This is the foundation of today's agents. |
| 2023 | [**Tree of Thoughts**](https://arxiv.org/abs/2305.10601) and [**Graph of Thoughts**](https://arxiv.org/abs/2308.09687) | Organize reasoning as a tree or graph: explore multiple branches, evaluate them, backtrack from dead ends |
| 2023–24 | **System prompts, structured outputs, tool calling** | Give the model a role and rules; force it to return valid JSON; let it call functions |
| 2025–26 | **Context engineering** | The focus shifts from clever wording to deciding *what information* goes into the model's context: the right documents, memory, tools, and examples |

The trend: prompting moved from **magic phrases** toward **system design**. That's good news, because system design is a skill you can actually learn and keep.

**Go deeper:** [Lilian Weng's prompt engineering overview](https://lilianweng.github.io/posts/2023-03-15-prompt-engineering/)

### Part 4: How the architecture evolved (models that reason, remember, and act)

**Reasoning models (built-in, organized chain-of-thought):** Chain-of-thought started as a prompting trick. Then labs trained it *into* the models. OpenAI's o1 (September 2024) and [DeepSeek-R1](https://arxiv.org/abs/2501.12948) (January 2025) use reinforcement learning to teach a model to produce long, structured reasoning before answering, then check and correct itself. This is called **test-time compute**: instead of only making models bigger, you let them *think longer* on hard problems. "Extended thinking" modes in Claude, ChatGPT, and Gemini work this way.

**Memory systems:** a model on its own forgets everything after each conversation. Memory gets layered on top:

| Memory type | What it is | Examples |
|---|---|---|
| **Context window** (short-term / working memory) | Everything the model can "see" right now. Grew from about 4k tokens to 1M+. | Long-context models |
| **Retrieval / RAG** (long-term knowledge) | Store documents as embeddings; fetch the relevant pieces when needed ([RAG paper, 2020](https://arxiv.org/abs/2005.11401)) | Document Q&A bots, Phase 3 of this path |
| **Knowledge graphs** | Store facts as connected entities and relationships, not just text chunks | [Microsoft GraphRAG](https://github.com/microsoft/graphrag) |
| **Agent memory** (episodic, semantic, procedural) | Remember past interactions, facts about the user, and how to do things; manage memory like an operating system manages RAM and disk | [MemGPT](https://arxiv.org/abs/2310.08560) → [Letta](https://github.com/letta-ai/letta), [Mem0](https://github.com/mem0ai/mem0), ChatGPT and Claude memory features |
| **LLM wiki** (compiled knowledge) | The LLM reads your sources once, then writes and keeps updating a set of linked Markdown notes. Instead of searching raw documents from scratch on every question like RAG, it builds up knowledge over time. | [Karpathy's LLM Wiki](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) (April 2026), an "idea file" you can paste into Claude Code or a similar agent to set one up |

### From transformers to next-gen architectures (solving the cost problem)

Transformers have one big weakness: **attention gets expensive fast as text gets longer.** Every token looks at every other token, so doubling the length of the input roughly *quadruples* the work. Computer scientists write this as **O(n²)**, "quadratic" cost. That's why long documents, huge codebases, and hour-long conversations are slow and costly to process. The newer architectures are mostly different answers to that problem:

| Architecture | How it works | Metaphor | Real-world example |
|---|---|---|---|
| **Dense transformer** (2017) | Every part of the model works on every token, and every token attends to every other token | **A meeting where everyone must talk to everyone.** Fine with 10 people; with 10,000 people, the number of conversations explodes. | The original GPT models; the reason long-context models used to be slow and expensive |
| **[Mixture of Experts (MoE)](https://arxiv.org/abs/2401.04088)** (sparse) | The model contains many specialist sub-networks ("experts"). A small **router** sends each token to only a few of them. Total knowledge is huge, but only a fraction is used per token. | **A hospital.** It employs hundreds of specialists, but when you walk in, triage sends you to the two doctors you actually need. The hospital's knowledge is huge; your visit only uses a little of it. | Mixtral (2023), DeepSeek's models, and many frontier models. It's a big reason models got cheaper to run. |
| **State-space models (SSMs): [S4](https://arxiv.org/abs/2111.00396), [Mamba](https://arxiv.org/abs/2312.00752)** | Instead of comparing every token with every other, the model reads in order and keeps a running **compressed summary** (its "state"). Cost grows in a straight line with length, **O(n)**. Mamba adds *selectivity*: it decides what's worth remembering based on the input. | **Taking notes during a lecture.** You don't replay the whole lecture every time a new sentence starts; you update your notes and keep going. Mamba is the student who knows what's worth writing down. | Processing very long inputs (long documents, audio, genomic data) with much less memory |
| **Hybrid architectures: [Jamba](https://arxiv.org/abs/2403.19887), [Griffin](https://arxiv.org/abs/2402.19427)** | Mix the two: mostly fast SSM layers for digesting long sequences, plus some attention layers for precise lookups and reasoning | **Skim, then zoom.** A good researcher skims a 300-page report quickly (SSM), then reads the important pages closely (attention). | Jamba (AI21 Labs, 2024) and Griffin (Google DeepMind, 2024); hybrids keep showing up in long-context models |

**Why you should care as a builder:** these choices show up in your bills and your results. They explain why some models handle a 500-page PDF cheaply while others choke, why prices keep dropping (section 1), and why "bigger model" doesn't always mean "slower model" anymore.

**Agents and tools:** models that plan, call tools, read results, and loop until a task is done (the ReAct idea, productized). [**MCP**](https://modelcontextprotocol.io) (Model Context Protocol, 2024) standardized how agents connect to tools and data. That's what harnesses like Claude Code and dsh are built around.

### Diffusion models: a completely different way to generate

Everything above is about **transformers**, which power text AI. Most AI **images and video** ([Midjourney](https://www.midjourney.com), Stable Diffusion, [Flux](https://bfl.ai), DALL·E, Sora, Veo, and what you run in ComfyUI) come from a different family: **diffusion models**.

**How diffusion works:**

1. **Training (learning to un-blur):** take millions of real images and gradually add random noise to each one until it's pure static, like TV snow. The model learns to reverse that, predicting at each step what noise to remove to get slightly closer to a real image.
2. **Generating:** start from pure random noise, then remove a bit of noise at a time, over many steps (often 20–50), guided by your text prompt, until a clear image appears.

**Metaphor:** a sculptor with a block of marble. The image is "hidden" in the noise, and each step chips a little away. Or think of fog slowly clearing to reveal a landscape: blurry shapes first, then outlines, then fine details.

**Transformers vs. diffusion, side by side:**

| | Transformer (LLM) | Diffusion model |
|---|---|---|
| **How it generates** | One token at a time, left to right, never going back | The *whole* output at once, refined over many steps |
| **Metaphor** | A storyteller speaking word by word | A sculptor revealing a statue, or fog clearing |
| **Where randomness comes from** | Rolling loaded dice for each next token | The random starting noise (the **seed**). Same seed + same settings = the same image. |
| **Fixing mistakes** | Can't revise earlier words while writing | Every step revisits the whole image, so early rough areas get corrected |
| **Best at** | Text, code, reasoning, conversation | Images, video, audio, design |
| **Speed** | Slows down as the output gets longer | Cost depends on number of steps and resolution |
| **Examples** | ChatGPT, Claude, Gemini, DeepSeek | Stable Diffusion, Flux, Midjourney, DALL·E, Sora, Veo |

**How diffusion evolved:**

| When | Milestone | What changed |
|---|---|---|
| 2014 | [**GANs**](https://arxiv.org/abs/1406.2661) (generative adversarial networks) | Before diffusion: two networks compete, a forger making fakes and a detective spotting them. Sharp results, but unstable to train. |
| 2020 | [**DDPM**](https://arxiv.org/abs/2006.11239) | Showed that diffusion could produce high-quality images, and it soon overtook GANs |
| 2021 | [**CLIP**](https://arxiv.org/abs/2103.00020) | A model that connects images and text, which let prompts steer image generation |
| 2022 | [**Latent diffusion / Stable Diffusion**](https://arxiv.org/abs/2112.10752) | Run diffusion on a **compressed version of the image** (the "latent space") instead of every pixel. That's dramatically cheaper, so it runs on a consumer GPU. This is why ComfyUI and local image generation exist. |
| 2022–24 | [**Diffusion Transformers (DiT)**](https://arxiv.org/abs/2212.09748) | Replace diffusion's older U-Net backbone with a **transformer**. The two families merge: newer image and video models (e.g. Sora, Stable Diffusion 3, Flux) use transformers *inside* the diffusion process. |
| 2025 | **Diffusion for text:** [LLaDA](https://arxiv.org/abs/2502.09992), [Inception Labs' Mercury](https://www.inceptionlabs.ai), [Google's Gemini Diffusion](https://deepmind.google/models/gemini-diffusion/) | Diffusion applied to *language*: generate a whole draft of text in parallel, then refine it. Much faster output, though still catching up to the best transformer LLMs on quality. |

**Words you'll see in ComfyUI and image tools:**

- **Steps:** how many denoising passes. More steps means more detail and slower generation.
- **Seed:** the starting noise. Reuse a seed to reproduce or tweak an image.
- **CFG (classifier-free guidance) scale:** how strictly the image follows your prompt. Too high looks overcooked; too low ignores the prompt.
- **Sampler / scheduler:** the method used to remove noise at each step. Different samplers trade speed against quality.
- **VAE:** the component that compresses images into latent space and decodes them back into pixels.
- **LoRA:** a small add-on file that teaches a model a specific style, character, or product without retraining it.

**My take:** diffusion is a powerful tool for practical visuals like mockups, product shots, and design drafts. But AI art on its own feels soulless to me; see "Why I do this" at the top of this guide.

**Real-world uses:** product photos and ad creatives, concept art, interior and architecture mockups, video ads, and, in ComfyUI, fully automated image pipelines for brands (the "AI content and creative" track in section 9).

**Watch:**

- [But How Do AI Images and Videos Actually Work?](https://www.youtube.com/watch?v=iv-5mZ_9CPY) (3Blue1Brown × Welch Labs, 37 min). The best visual explanation.
- [How AI Image Generators Work (Stable Diffusion / DALL·E)](https://www.youtube.com/watch?v=1CIpzeNxIhU) (Computerphile, 18 min)
- [Diffusion Models for AI Image Generation](https://www.youtube.com/watch?v=x2GRE-RzmD8) (IBM Technology, 12 min)
- [Text Diffusion: A New Paradigm for LLMs](https://www.youtube.com/watch?v=bmr718eZYGU) (Julia Turc, 24 min)

**Go deeper:** [The Illustrated Stable Diffusion](https://jalammar.github.io/illustrated-stable-diffusion/) (Jay Alammar)

### Part 5: Jev and the return of "best at one thing" (September 2026)

And here's the funny part. After a decade of chasing one giant model that does everything, one of the newest buzzworthy releases goes back to the original idea of machine learning.

[**Jev**](https://en.wikipedia.org/wiki/Jev_(AI_model)), from TypeSafe AI, was released in limited early access on **September 15, 2026**. One of its founders, Diogo Almeida, worked at OpenAI on RLHF, InstructGPT, ChatGPT, and GPT-4.

**What makes it different:**

- **It doesn't generate text.** You give it text or JSON plus typed questions (yes/no, multiple choice, or a ranked score), and it returns **structured decisions with calibrated probabilities**. No prose, so no hallucinated paragraphs.
- **It's a discriminative model**, like the classic classifiers from Part 1, but built on transformer-scale understanding of language.
- **It's marketed as a "System One" model**, after Daniel Kahneman's fast, intuitive thinking. LLMs that reason step by step are the slow "System Two."
- **Speed and cost:** TypeSafe reports 70–500 ms responses and **$0.042 per million input tokens**, with output free, and claims it's 40–200x faster than frontier LLMs for these tasks. These are the company's own benchmarks, not independently verified yet.
- **The name** comes from economist William Stanley Jevons and **Jevons' paradox**: when something gets cheaper and more efficient, people use far more of it, not less.

**What it's used for:** classifying emails, routing support tickets, checking whether AI output is safe before users see it, guarding against jailbreaks, and routing work to the right model. It fills the **decision points inside automation flows**.

**Why this matters for you:** the history made a full circle. The spam filter or potato sorter of 2005 was the best at one decision. Jev is the best at fast decisions, at modern scale. And it maps exactly onto the core skill from the top of this guide: **data in, process, structured data out.** In an n8n flow, a model like Jev could be the node that decides "Is this a booking request, a complaint, or spam?" in milliseconds, for a fraction of a cent, before a bigger model drafts the reply.

**The lesson from all of this history:** the tools keep changing, sometimes in circles. The person who understands *what kind of tool fits which job* (a classifier, a reasoning model, a retrieval system, or an agent) stays valuable no matter what launches next.

### Watch

- [What Is Jev? The AI Model That Doesn't Generate Text](https://www.youtube.com/watch?v=YGgNBcIgI4s) (IBM Technology, 15 min)
- [Jev Explained in 7 Minutes](https://www.youtube.com/watch?v=vj7hysh0mOI) (Caleb Writes Code)
- [The Entire History of Artificial Intelligence (Last 100 Years)](https://www.youtube.com/watch?v=mSd9nmPM7Vg) (Putchuon, 23 min)
- [Chain-of-Thought Prompting, Explained](https://www.youtube.com/watch?v=AFE6x81AP4k) (CodeEmporium, 9 min)
- [Tree of Thoughts (full paper review)](https://www.youtube.com/watch?v=ut5kp56wW_4) (Yannic Kilcher, 29 min)
- [How Do Thinking and Reasoning Models Work?](https://www.youtube.com/watch?v=xCRvOUykOX0) (Google for Developers, 13 min)
- [Why AI Models Pause to Think: Test-Time Compute Explained](https://www.youtube.com/watch?v=DAlC8mL5ZlI) (IBM Technology, 11 min)
- [The Four Types of Memory Every AI Agent Needs](https://www.youtube.com/watch?v=BacJ6sEhqMo) (IBM Technology, 11 min)
- [What Is Mixture of Experts?](https://www.youtube.com/watch?v=sYDlVVyJYn4) (IBM Technology, 8 min)

**Sources for this section:**

- [Jev (AI model), Wikipedia](https://en.wikipedia.org/wiki/Jev_(AI_model))
- [TechCrunch: A new kind of AI model from a ChatGPT inventor is thrilling developers (Sept 2026)](https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/)
- [DigitalOcean: What is Jev?](https://www.digitalocean.com/resources/articles/what-is-jev)
- [Jevons paradox, Wikipedia](https://en.wikipedia.org/wiki/Jevons_paradox)
- Papers linked inline above (arXiv)

---

## 3. The core skill: learning *with* AI

Anyone can get output from AI now, so output alone isn't worth much. Companies pay for **judgment**: knowing when the AI is wrong, debugging what it built, deciding what to build, and owning the result when it breaks in production. You only get judgment by building mental models yourself. **AI just makes building them about 10x faster.**

### The trap

You paste a prompt, get working code, and ship it. It *feels* like learning, but very little sticks. Three months later you can't explain your own project, can't debug it, and fall apart in an interview. This is the most common way beginners fail in 2026. Researchers call it *cognitive offloading*: the AI does the thinking, so your brain never builds the skill.

### The three modes. Always know which one you're in.

| Mode | Who drives | When to use it |
|---|---|---|
| **Tutor** | You think. AI asks questions, explains, and quizzes you. | Learning anything new. Most of your first 3 months. |
| **Pair** | You write. AI reviews and challenges you. | Once you know the basics of a topic |
| **Delegate** | AI writes. You review every line. | Only for things you could already do yourself, just slower |

> **The rule: you think, AI executes.** Only delegate what you *understand*. You don't need to write the syntax from memory, but before you accept anything, you must be able to explain what it does, why it's built that way, and what breaks if it's wrong. If you can't explain it, you're not ready to ship it.

### Built-in tutor modes (use them)
- **ChatGPT Study Mode:** it guides you with questions instead of handing you answers.
- **Claude:** set the tutor prompt below as your Project instructions. In Claude Code, switch the output style to *Learning* or *Explanatory*. It will leave small parts for you to write and explain its choices.

### Copy-paste tutor prompt

Paste this into Claude Project instructions or ChatGPT custom instructions:

```
You are my tutor, not my assistant. I'm learning [topic] from zero.
- Never give me full solutions unless I say "show me".
- When I'm stuck, ask me a guiding question first.
- Explain with analogies from [my background: marketing / business / etc].
- After every concept, give me one small exercise to do in my own terminal.
- When I show you my code, point out what's wrong but make me fix it.
- End each session by quizzing me with 3 questions.
- If you're unsure about a version, flag, or API, say so. Don't guess.
```

### Seven habits that turn AI use into skill

1. **You think, AI executes.** You decide the what and the why. AI handles the how (the syntax). Code you can't explain teaches you very little. Code you understand is what builds the skill.
2. **Ask "why" three times** before running anything you don't understand.
3. **Break it on purpose.** Ask: *"Give me 3 ways this setup breaks. I'll diagnose them."* For security people, this habit is the job.
4. **Read every diff** an agent produces and explain it back in plain words. If you can't, don't merge it.
5. **Use AI to read real code.** Clone a serious repo and ask an agent: *"Explain the architecture. Where does a request enter? What would you change first?"* It's one of the best shortcuts there is. Juniors used to need months of mentorship to get this.
6. **Don't trust what the AI remembers about versions.** Paste the official docs in or point the agent at them. Models are often wrong about flags and APIs that changed after they were trained.
7. **Keep an AI Engineering Journal** (template in section 15). An AI summary of your own week doesn't count.

### Prompts that make AI a better teacher
- *"Don't give me the answer. Ask me questions until I find the bug myself."*
- *"Explain this like I know [marketing], not computer science."*
- *"Read this error message with me line by line. What is each part telling me?"*
- *"Quiz me on what we covered yesterday. 5 questions, increasing difficulty."*
- *"What would a senior engineer criticize about this code?"*
- *"Give me a 20-minute exercise that forces me to use what I just learned."*
- *"Here's the transcript of a tutorial I watched. Make a quiz and a hands-on lab from it."*

### Watch first
| Video | Why |
|---|---|
| [ChatGPT Study Mode, Explained by a Learning Expert](https://www.youtube.com/watch?v=m3jNwwuvqx8) (Justin Sung, 21 min) | How to actually *learn* with AI instead of outsourcing your thinking |
| [Everything You Need to Know About Coding with AI // NOT vibe coding](https://www.youtube.com/watch?v=5fhcklZe-qE) (ForrestKnight, 13 min) | The difference between using AI and depending on it |
| [How I use LLMs](https://www.youtube.com/watch?v=EWvNQjAaOHw) (Andrej Karpathy, 2h) | How one of the field's top people uses these tools day to day |
| [Software Is Changing (Again)](https://www.youtube.com/watch?v=LCEmiRjPEtQ) (Karpathy at YC, 40 min) | The big picture: why this is a new kind of programming |
| [From Vibe Coding to Agentic Engineering](https://www.youtube.com/watch?v=96jN2OCOfLs) (Karpathy at Sequoia, 30 min) | Where things are heading in 2026 |

---

## 4. The mindset: think in systems, solve business problems

Tools change every month. The thing that makes you valuable doesn't: **being a systems thinker who solves real problems for businesses.** The rest of this path is how you build that.

### What systems thinking means

A systems thinker doesn't look at a task. They look at **the whole flow around it**: where information comes from, what happens to it, where it goes, who touches it, and where it breaks.

Every business process, and every AI system, comes down to the same loop:

| Stage | What happens | Example: a clinic's booking requests |
|---|---|---|
| **1. Input** | Data comes in, usually messy | Customers send booking requests over Messenger in free text |
| **2. Process** | It gets cleaned, understood, decided on, transformed | AI extracts name, service, date, and time; checks the calendar for conflicts |
| **3. Output** | A structured result someone (or another system) can act on | A draft booking in the calendar, flagged for staff to confirm |
| **4. Feedback** | You check whether it worked and feed that back in | Track how many drafts staff had to fix; improve the extraction where it fails |

**Questions a systems thinker asks:**

- Where does the information come from, and what shape is it in? (Emails, forms, PDFs, chats, spreadsheets?)
- What happens to it today, and who does it by hand?
- Where's the **bottleneck**? The slowest, most error-prone, or most expensive step is where help is worth the most.
- What's the output, and who uses it? Is it structured enough for a person or another system to act on?
- How do we know it worked? (The feedback loop. Without one, you're guessing.)
- What breaks if one piece fails? (This is also how security people think.)

**Learn it:** [*Thinking in Systems: A Primer*](https://www.chelseagreen.com/product/thinking-in-systems/) by Donella Meadows is the classic, short and readable. Her free essay [*Leverage Points: Places to Intervene in a System*](https://donellameadows.org/archives/leverage-points-places-to-intervene-in-a-system/) is the best 30 minutes you can spend on this.

### From systems thinking to solving business problems

Businesses don't pay for AI. **They pay for problems to go away**: hours lost, errors made, leads dropped, customers waiting. AI is just one way to fix that, and sometimes the right fix is a spreadsheet or a process change.

**The playbook:**

1. **Find the pain.** Talk to the people doing the work. Look for tasks that are repetitive, slow, error-prone, or depend on one person who's always overloaded.
2. **Map the system.** Draw the input → process → output flow as it works today (free tools below). Most of the insight comes from this step alone.
3. **Find the bottleneck.** Fix the step that costs the most, not the one that's most fun to automate.
4. **Design the fix.** Decide what should be automated, what AI should handle (reading messy text, classifying, summarizing, drafting), and what must stay human (judgment, approvals, relationships).
5. **Build the smallest version that works.** n8n, Claude Code, an agent, whatever fits. Ship it to real users fast.
6. **Measure it.** Hours saved, errors reduced, response time cut. Numbers are what turn a project into a case study and a client into a referral.

**Examples of what this looks like:**

- A clinic's front desk retypes booking requests from Messenger into a calendar → an automation reads the message, extracts the details, and creates the booking for a human to confirm.
- A sales team loses leads because nobody replies after hours → an AI assistant answers common questions, qualifies the lead, and hands off to a person in the morning.
- A small business has years of PDFs and invoices nobody can search → a RAG system turns them into something staff can ask questions of.

Notice that none of these are "build an AI app." They're "make this specific painful thing go away."

### Drawing it: free tools for flowcharts and process maps

You can't fix a process you can't see. Mapping it out is step 2 of the playbook, and it's the most underrated skill in this whole path: it helps you think, it helps clients understand, and it makes your portfolio readable.

| Tool | Best for | Cost |
|---|---|---|
| [**Excalidraw**](https://excalidraw.com) | Quick hand-drawn-style flowcharts and system maps. Great for client calls and READMEs. It can also convert Mermaid code into an editable diagram. | Free, open source ([GitHub](https://github.com/excalidraw/excalidraw)) |
| [**draw.io / diagrams.net**](https://app.diagrams.net) | Formal process maps, swimlanes, architecture diagrams | Free, open source |
| [**Mermaid**](https://mermaid.js.org) ([live editor](https://mermaid.live)) | Diagrams written as text. **AI can write these for you**, and GitHub renders them directly in READMEs. | Free, open source |
| [**tldraw**](https://www.tldraw.com) | Fast whiteboard sketching | Free tier |

**Process mapping basics:**

- **Start and end:** what triggers the process, and what does "done" look like?
- **Steps:** each action, in order
- **Decisions:** the yes/no branches (diamonds)
- **Swimlanes:** one row per person, team, or system, so you can see who does what and where handoffs happen
- **Pain points:** mark where it's slow, error-prone, or manual. That's where your solution goes.

**AI shortcut:** describe the process to ChatGPT or Claude and ask: *"Turn this into a Mermaid flowchart with swimlanes for each role."* Paste the result into mermaid.live or Excalidraw, then fix it by hand. The fixing is where you actually understand the process.

**Watch:**

- [Excalidraw, my favorite whiteboard / tech diagram app](https://www.youtube.com/watch?v=Gv9MezPAchI) (Christian Lempa, 14 min)
- [What Is a Swimlane Diagram?](https://www.youtube.com/watch?v=Yn_wfpEQoXs) (Gliffy, 3 min)
- [Process Mapping in 5 Minutes](https://www.youtube.com/watch?v=7Xh6g1sW2KY) (OpsKings)

### You own the output

Having AI doesn't make you good at what you do. **It means you're trusted to be responsible for whatever your AI produces.** When an automation sends a wrong message to a customer, or an agent deletes the wrong file, or AI-written code leaks data, "the AI did it" isn't an answer. The client hired *you*.

That's why serious AI setups take time. In professional software development, nobody just types "build me an app." The process looks like this:

| Step | What happens | Why it matters |
|---|---|---|
| **1. PRD** (Product Requirements Document) | Write down the problem, the users, what the system must do, and what "done" looks like | AI builds exactly what you describe, so a vague request gets you vague software |
| **2. Brainstorm** | Explore approaches with AI, weigh the options, poke holes in the plan | It's cheaper to change a plan than code |
| **3. Design** | Map the system: data flow, tools, where humans stay in the loop | This is the systems thinking from this section |
| **4. Plan** | Break the work into small, testable tasks | Small steps are easy to verify and easy to undo |
| **5. Build** | AI writes code task by task; you review each piece | You think, AI executes |
| **6. Test** | Check it works, including the edge cases and failure modes | Proof instead of hope |
| **7. Review and secure** | Read the diffs, check for exposed secrets and risky permissions | You own what ships |
| **8. Deploy, document, monitor** | Ship it, write it up, watch for problems | Someone has to maintain it, possibly you at 2am |

This approach is often called **spec-driven development**: write the spec first, then have AI build against it. The setup takes longer up front and saves you from rebuilding later.

**Learn it:**

- [Spec-Driven Development: AI Assisted Coding Explained](https://www.youtube.com/watch?v=mViFYTwWvcM) (IBM, 9 min)
- [Full Course: Spec-Driven Development with Coding Agents](https://www.youtube.com/watch?v=hy8UstR2NEg) (DeepLearning.AI + JetBrains, 1h)
- [Build Better Apps with AI Using This One Simple Document (PRD Guide)](https://www.youtube.com/watch?v=MZjW7mlRgdw) (Jordan Urbs, 21 min)
- [github/spec-kit](https://github.com/github/spec-kit): GitHub's open-source toolkit for spec-driven development with AI agents
- [Claude Code best practices](https://www.anthropic.com/engineering/claude-code-best-practices) (Anthropic)

### The "build software everyone uses" dream (and survivorship bias)

The path that looks most exciting right now is building a product, an app or SaaS that thousands of people pay for. It's a legitimate goal, but go in with open eyes.

**Why it's harder than it looks:**

- **Building is the cheap part now.** AI made software fast to build, so thousands of people are shipping similar apps. The hard part is getting anyone to notice and pay.
- **Marketing is expensive.** The median B2B SaaS company spends about **$702 to acquire one self-serve customer**, and around **$2 in sales and marketing for every $1 of new annual revenue**. Acquisition costs have risen over 200% in eight years.
- **Most don't make it.** Roughly half of new businesses don't reach year five. Among venture-backed startups, the commonly cited failure rate is around 90%. For AI startups specifically, lack of market demand (building something nobody needs) is the top reason they fail.

**Survivorship bias:** You hear about the founder whose AI app made $50K a month. You don't hear about the thousands who built the same kind of app and made nothing, because failures don't post screenshots. Looking only at the winners makes success look like a recipe when it was partly timing, distribution, and luck.

This isn't unique to AI or software. It's true in most industries, from restaurants to music: not everyone succeeds, and the visible winners are a skewed sample. The point isn't "don't try." It's **don't plan your life around the survivors' stories.**

**The smarter sequence:**

1. **Solve problems for real businesses first.** You get paid while you learn, you see real problems up close, and every project becomes proof of what you can do.
2. **Watch for patterns.** When five different clients have the same problem, that's a product idea with demand already proven.
3. **Then build the product,** for customers you already understand and can already reach.

### Keeping up without burning out

AI moves fast: new models, harnesses, and "game-changing" tools every week. Nobody keeps up with all of it, and you don't need to.

- **Learn principles, not just tools.** Harnesses, RAG, MCP, data handling, and systems thinking carry over when the tools change. dsh will get replaced someday; understanding how a harness works won't.
- **Use a filter.** Pick a few trusted sources, check them weekly, and ignore the rest. If a tool matters, you'll hear about it more than once.
- **My curated AI Reddit feed:** [reddit.com/user/saintjedi/m/ai](https://www.reddit.com/user/saintjedi/m/ai/) is one feed that combines dozens of AI subreddits (ChatGPT, ClaudeCode, n8n, automation, AI agents, local models, machine learning, and more). It also includes **anti-AI subreddits on purpose**: you should hear the criticism too, not just the hype.
- **Try things through your projects.** Test a new tool only when it could solve a problem you actually have.
- **Watch the consequences, not just the capabilities.** Jobs, the environment, and mental health (see section 1) are part of the picture. The people who'll be trusted with AI systems are the ones who understand the costs as well as the gains.

**Sources for this section:**

- [Customer acquisition cost benchmarks 2026 (Shno)](https://www.shno.co/marketing-statistics/customer-acquisition-statistics)
- [Startup survival statistics 2026 (Lonely Entrepreneur)](https://lonelyentrepreneur.com/startup-failure-statistics-2026/)
- [Top 100 startup failure statistics 2026 (Indie Hackers)](https://www.indiehackers.com/post/top-100-startup-failure-statistics-2026-why-most-startups-fail-and-what-every-founder-must-know-before-it-s-too-late-3cfe6e6aa3)
- [Survivorship bias and startups (Nick Raushenbush)](https://nickraushenbush.substack.com/p/survivorship-bias-and-startups-12c59be659ae)

---

## 5. Hardware: what you actually need

**You don't need a GPU to start.** Almost everything in the first ~4 months runs in the cloud.

| Tier | Setup | What it unlocks |
|---|---|---|
| **0** | Any laptop with 8GB RAM and internet | Chat AIs, coding agents through APIs, n8n, Python. **Enough for Phases 0–3.** |
| **1** | 16GB RAM, no GPU | Comfortable WSL2/Linux, [Docker](https://www.docker.com), one security VM (Kali) at a time, tiny local models on CPU (slow, but fine for learning) |
| **2** | GPU with 8GB VRAM (RTX 3060 / 4060) | Local 7–9B models at Q4 (e.g. Qwen3.5-9B). Good for privacy and offline experiments. |
| **3** | GPU with 16GB VRAM (RTX 4060 Ti 16GB) | 20B-class models (gpt-oss-20b, Devstral Small 2), image generation (ComfyUI), usable local agents |
| **4** | Used RTX 3090 (24GB), or a Mac with 32GB+ unified memory | 27–32B models (e.g. Qwen3.6-27B, ~77 on SWE-bench Verified). The classic budget-enthusiast upgrade. |

### Free GPU workarounds
- **[Kaggle Notebooks](https://www.kaggle.com):** free GPU hours every week. Best for ML exercises.
- **[Google Colab](https://colab.research.google.com):** a free GPU tier that's fine for notebooks from the books in section 11.

### Honest take on local models
Local models are for **learning how models work, privacy, and offline use**. They still don't replace frontier models for serious agentic coding.

**Gotcha:** Ollama's default context window is tiny (4k tokens under 24GB VRAM). Set `OLLAMA_CONTEXT_LENGTH=65536` or higher before connecting a coding agent, or it will forget your files halfway through a task.

### Security homelab
[VirtualBox](https://www.virtualbox.org) (free), plus a [Kali Linux](https://www.kali.org) VM, plus a deliberately vulnerable VM to attack. 16GB RAM makes this comfortable.

### Watch
- [Learn Ollama in 15 Minutes: Run LLM Models Locally for FREE](https://www.youtube.com/watch?v=UtSSMs6ObqY) (Tech With Tim)
- [What is Ollama? Running Local LLMs Made Simple](https://www.youtube.com/watch?v=5RIOQuHOihY) (IBM Technology, 7 min)
- [Best Local Coding AI for Your GPU (4GB to 512GB)](https://www.youtube.com/watch?v=qkRIW2ieOK8) (Cloud Codes)
- [Best Local AI Model for EVERY GPU (8GB to 512GB+)](https://www.youtube.com/watch?v=9CAZapI_WTE) (RepoChad)

---

## 6. Tools and budget stacks

### Dead or shrunk in 2026. Ignore old tutorials that recommend these.
- **Gemini CLI's free Google-login tier:** shut down **June 18, 2026**. Google moved users to Antigravity CLI, where the free quota is reportedly only tens of requests a day.
- **Qwen Code's free OAuth tier:** cut from 1,000 to 100 requests a day, then **closed April 15, 2026**.
- **GitHub Copilot Student:** new sign-ups **paused since April 2026**. Existing students keep access.
- **Google AI Studio's free API:** **Flash models only** since April 2026, with limits cut significantly. Your real limits only show in your AI Studio console.

### The coding agents ("harnesses")

A *harness* is the software that turns a model into an agent: it reads files, runs commands, edits code, and asks your permission along the way. The model is the brain; the harness gives it hands. **Learn one harness well, then swap models underneath it.** That way you're not locked into one vendor's pricing.

| Harness | Type | Why use it |
|---|---|---|
| **[DeepSeek Harness (dsh)](https://github.com/deepseek-ai/deepseek-harness)** | Local web UI + headless mode | DeepSeek's official open-source harness (MIT), released Aug 2026. "Everything is a plugin." Pairs with DeepSeek V4 Flash, which is very cheap, but it's **model-open**: OpenRouter, Anthropic, OpenAI, or any OpenAI-compatible or local endpoint. Start it with `npx @deepseek-ai/dsh web`. It's a **developer preview, so expect breaking changes**, and read its `SAFETY.md` first. |
| **[OpenCode](https://opencode.ai)** | Terminal + desktop | Open source, supports 75+ providers. Its **Zen** provider includes free, coding-tested models (currently e.g. Big Pickle, DeepSeek V4 Flash, Nemotron 3 Ultra Free, MiMo-V2.5 Free; the list rotates). The best ₱0 starting point. |
| **[Aider](https://github.com/Aider-AI/aider)** | Terminal | Git-native: every AI change becomes a commit, so you're always reading diffs, which suits learning well |
| **[Cline](https://github.com/cline/cline) / [Kilo Code](https://github.com/Kilo-Org/kilocode) / [Roo Code](https://github.com/RooCodeInc/Roo-Code)** | VS Code extensions | Good if you'd rather stay in an editor |
| **[Codex CLI](https://github.com/openai/codex)** | Terminal | Works when you sign in with a free ChatGPT account (limited), or with your own API key |
| **[Claude Code](https://claude.com/product/claude-code)** ([docs](https://code.claude.com/docs)) | Terminal + desktop | Needs Claude Pro or the API. The current quality ceiling for agentic coding. |

### ₱0 stack
- **Chat tutors:** ChatGPT free, Claude free, and Gemini. Rotate between them when you hit a limit.
- **Agent:** OpenCode with Zen free models, **or** dsh with OpenRouter free models.
- **[OpenRouter](https://openrouter.ai) free models:** one API key for 300+ models. The free ones have `:free` at the end of their name. You get **50 requests a day** until your account has bought $10 of credits total, then **1,000 a day permanently**, capped at 20 requests a minute. Each free model has its own daily quota, so rotate between them. Works in dsh, OpenCode, Aider, Cline, Kilo, n8n, and your own Python scripts.
- **Local:** [Ollama](https://ollama.com) or [LM Studio](https://lmstudio.ai), if you have a tier 2+ GPU.

### The best ₱600 you'll spend: a one-time $10 OpenRouter top-up
- It permanently raises your free-model limit from 50 to **1,000 requests a day**.
- The credits also pay for cheap models. **DeepSeek V4 Flash costs $0.14 per million input tokens and $0.28 per million output tokens.** A month of heavy learning on it usually costs a few dollars.

### About ₱1,200/month: one frontier subscription (pick ONE)
- **Claude Pro ($20/mo):** includes [Claude Code](https://claude.com/product/claude-code).
- **ChatGPT Plus ($20/mo):** includes Codex. ChatGPT Go ($8/mo) covers light use but doesn't include Codex cloud tasks.
- **Model routing:** use the frontier model for hard reasoning and cheap models (DeepSeek V4 Flash, free models) for grunt work. Knowing how to do this is a paid skill in itself.

Prices are in USD. No credit card? A debit or virtual card from your e-wallet or digital bank works for most of these.

### Privacy rule (non-negotiable, especially for security people)
Free endpoints often **log or train on your prompts**. Never paste API keys, `.env` files, passwords, client data, or company code into a free model. Learning this habit early is already security practice.

### Watch
**DeepSeek Harness (dsh)**

- [DeepSeek Harness Tutorial: Free & Paid Models with OpenRouter + Real Coding](https://www.youtube.com/watch?v=0tb33f1hJHA) (Sahand, 15 min). **Start here:** it's the exact dsh + OpenRouter setup.
- [DeepSeek Harness Agentic AI Crash Course (Run Any AI Model)](https://www.youtube.com/watch?v=legYz3Hk2rQ) (Caleb Curry, 46 min)
- [Use DeepSeek Harness for FREE: No VRAM, No Paid API](https://www.youtube.com/watch?v=3g28CmoapOw) (Bart Slodyczka)
- [DeepSeek Harness Setup: A Free Claude Code You Own in 10 Minutes](https://www.youtube.com/watch?v=EDxVn1q8udE) (Sharbel A.)
- [DeepSeek Harness 2.0 (Desktop App, Creator Mode)](https://www.youtube.com/watch?v=TW2gCpbNz4M) (AICodeKing). Covers the newest features.

**OpenCode**

- [OpenCode Tutorial for Beginners: Learn 90% in Under 25 Minutes](https://www.youtube.com/watch?v=QzqaZshQcJI) (Brandon Melville)
- [OpenCode Full Tutorial: Free Models, Skills & MCPs](https://www.youtube.com/watch?v=0xKE1UHpSfk) (Eric Tech)
- [OpenCode Setup That Makes AI Free](https://www.youtube.com/watch?v=Ef5zPm7U7So) (Edward Donner)

**OpenRouter and Aider**

- [What is OpenRouter: All About OpenRouter in 10 Minutes](https://www.youtube.com/watch?v=fjd2hm6-qtM) (codebasics)
- [How to Use AI Models API for Free: OpenRouter Tutorial](https://www.youtube.com/watch?v=VvJvJ0uXiVQ) (The Coding Koala)
- [Why You Should Use Aider for AI Coding](https://www.youtube.com/watch?v=_m6rpy0-Lrk) (Zen van Riel)

---

## 7. The path, phase by phase

> **Don't wait until the end to start earning.** By month 3 you should be packaging your work and reaching out. See section 8 for the fast-track timeline.

### Phase 0: Set up and learn to talk to AI (weeks 1–2)

**Do**

- Install [WSL2](https://learn.microsoft.com/en-us/windows/wsl/install) (on Windows) or use Linux, plus [VS Code](https://code.visualstudio.com), [Git](https://git-scm.com/downloads), [Python](https://www.python.org/downloads/), and [Node.js](https://nodejs.org).
- Make accounts: [GitHub](https://github.com/signup), [OpenRouter](https://openrouter.ai), [ChatGPT](https://chatgpt.com), [Claude](https://claude.ai).
- Set up the tutor prompt from section 3.
- Do **Anthropic's interactive prompt engineering tutorial** ([anthropics/courses](https://github.com/anthropics/courses)) and skim the [Prompt Engineering Guide](https://www.promptingguide.ai).
- **Optional structured route (the one I took):** [Google AI Essentials](https://www.coursera.org/learn/google-ai-essentials), then [Google Prompting Essentials](https://www.coursera.org/learn/google-prompting-essentials), then Vanderbilt's [Prompt Engineering](https://www.coursera.org/specializations/prompt-engineering) and [Generative AI Automation](https://www.coursera.org/specializations/generative-ai-automation) specializations. They're on Coursera; if cost is an issue, check the audit and financial-aid options on each course page.

> **For Filipinos: free AWS courses on TESDA.** The TESDA Online Program now has an [**Amazon Web Services (AWS) course category**](https://e-tesda.gov.ph/course/index.php?categoryid=2111) with three courses: **Fundamentals of Artificial Intelligence**, **AWS Certified Cloud Practitioner (CLF-C02)**, and **AWS Certified AI Practitioner (AIF-C01)**. They look free to enroll in, but check the course pages for current details. The Cloud Practitioner course fits the DevOps track, and the AI Practitioner course is a solid structured intro to AI concepts and AWS's AI services. (Note: the TESDA courses prepare you for the exams; check AWS for the actual exam fees.) I'm interested in these myself, especially for my DevOps track.

**Watch**

- [Large Language Models Explained Briefly](https://www.youtube.com/watch?v=LPZh9BOjkQs) (3Blue1Brown, 8 min)
- [[1hr Talk] Intro to Large Language Models](https://www.youtube.com/watch?v=zjkBMFhNj_g) (Andrej Karpathy)
- [Prompting 101 | Code w/ Claude](https://www.youtube.com/watch?v=ysPbXH0LpIE) (Anthropic, 25 min)
- [AI Prompt Engineering: A Deep Dive](https://www.youtube.com/watch?v=T9aRN5JkmL8) (Anthropic, 1h16)

**Proof you're done:** You use AI daily for real tasks, and you can explain *tokens*, *context window*, and *hallucination* in your own words.

---

### Phase 1: Concepts, not syntax (weeks 3–6)

You don't need to master a programming language before you build. You need to **understand the concepts** well enough to direct AI, read what it produces, and know when it's wrong. AI writes the skeleton; you have to understand the body.

**The concepts to understand (be able to explain each one in plain words):**

| Concept | What you should be able to explain |
|---|---|
| Variables, functions, loops, conditions | What a program is doing, step by step, when you read it |
| Data structures (lists, objects/dictionaries) and JSON | How data is shaped and passed around |
| Files, the terminal, environment variables | Where things live and how programs find them (and why secrets go in `.env`, never in code) |
| APIs and HTTP (requests, responses, status codes) | How two systems talk; what a 401, 404, or 500 means |
| Client vs. server, frontend vs. backend | Which part runs where |
| Databases (tables, queries) | Where data is stored and how it's retrieved |
| Errors and logs | How to read an error message and find where it came from |
| Git (commits, branches, diffs) | How to track changes and undo mistakes |

**How to learn them:** Have AI build small real things with you in **pair mode**, then make it explain every part. Ask: *"Walk me through this line by line. What concept is each part using? What would break if I removed it?"* Change one thing yourself and predict what happens before you run it.

**Do**

- **Terminal:** [OverTheWire Bandit](https://overthewire.org/wargames/bandit/), a free wargame that teaches Linux by making you hack your way through levels. It doubles as security training from day one.
- **[MIT Missing Semester](https://missing.csail.mit.edu):** shell, Git, editors, debugging.
- **Git:** commit everything from now on. Your GitHub *is* your portfolio.
- **Read real code:** clone one small, popular repo and have an agent walk you through it until you can explain how it works.
- **Optional, if you want a formal course:** [Harvard CS50P](https://cs50.harvard.edu/python/) (free). Useful, but not required to move forward.

**Watch**

- [Missing Semester, Lecture 1: Course Overview + Introduction to the Shell](https://www.youtube.com/watch?v=MSgoeuMqUmU) (new edition) · [Shell Tools and Scripting](https://www.youtube.com/watch?v=kgII-YWo3Zw) · [Version Control (git)](https://www.youtube.com/watch?v=2sjqTHE0zok)
- [Linux for Hackers, EP 1](https://www.youtube.com/watch?v=VbEx7B_PTOE) (NetworkChuck) · [60 Linux Commands You NEED to Know](https://www.youtube.com/watch?v=gd7BXuUQ91w) (NetworkChuck)
- [Git & GitHub Crash Course for Beginners [2026]](https://www.youtube.com/watch?v=mAFoROnOfHs) (freeCodeCamp)
- Optional: [Harvard CS50P full course](https://www.youtube.com/watch?v=nLRL_NcnK-4) (freeCodeCamp + CS50, 16h). Only if you want the formal route.

**Proof you're done:** You've reached Bandit level 15+, you've built a small script with AI that's actually useful to you, and you can explain every part of it, plus every concept in the table, without looking anything up.

---

### Phase 2: Using LLMs from code and automation (weeks 7–10)

**Do**

- Design a **CLI chatbot with memory** that calls OpenRouter, then have an agent build it with you. You decide how it should work; AI writes it; you explain every part back.
- Learn JSON, APIs, system prompts, temperature, streaming, **tool calling**, and the **cost per call**.
- **n8n** ([n8n-io/n8n](https://github.com/n8n-io/n8n), self-hostable): build 3 automations, e.g. email triage, a Messenger auto-reply prototype, or a daily AI news digest.
- Work through: [microsoft/generative-ai-for-beginners](https://github.com/microsoft/generative-ai-for-beginners), [anthropics/claude-cookbooks](https://github.com/anthropics/claude-cookbooks), [openai/openai-cookbook](https://github.com/openai/openai-cookbook).

**Watch**

- [n8n Beginner Course](https://www.youtube.com/watch?v=4BVTkqbn_tY) (official n8n, 9 parts) · [n8n Tutorial: Zero to Hero](https://www.youtube.com/watch?v=UIf-SlmMays) (freeCodeCamp, 3.5h)
- [Build an AI Agent From Scratch in Python](https://www.youtube.com/watch?v=bTMPwUgLZf0) (Tech With Tim, 34 min)

**Proof you're done:** A chatbot that calls at least one tool (search, weather, a database) and a working n8n workflow that saves you real time.

---

### Phase 3: Agents, RAG, MCP, and agentic coding (weeks 11–16)

**Do**

- Build a **RAG app** over real documents (PDFs, notes) and learn embeddings, chunking, and retrieval.
- Build an **agent** with tools, then learn **MCP** (the standard way agents connect to tools).
- Learn **evals**: how you *prove* your AI app works instead of eyeballing it.
- Learn **context engineering**: deciding what goes into the model's context window, and why it matters more than clever prompts.
- Use dsh, OpenCode, or Aider on a real project in **pair mode**. Write an `AGENTS.md` that explains your repo to the agent.
- Courses: [Hugging Face Agents Course](https://huggingface.co/learn/agents-course) and [LLM Course](https://huggingface.co/learn/llm-course) (free), [microsoft/ai-agents-for-beginners](https://github.com/microsoft/ai-agents-for-beginners), [DeepLearning.AI short courses](https://www.deeplearning.ai/short-courses/) (free, 1–2 hrs each).

**Watch**

- [Learn RAG From Scratch](https://www.youtube.com/watch?v=sVcwVQRHIc8) (freeCodeCamp, by a LangChain engineer, 2.5h)
- [Full Course (Lessons 1–10): AI Agents for Beginners](https://www.youtube.com/watch?v=OhI005_aJkA) (Microsoft Developer, 1h)
- [Intro to Agents: Create an Agent from Scratch (No Frameworks)](https://www.youtube.com/watch?v=vHDwpoSFdQY) (Hugging Face)
- [Claude Code: Full Tutorial for Beginners](https://www.youtube.com/watch?v=ntDIxaeo3Wg) (Tech With Tim, 36 min)
- [Build Your Own Agentic Harness in Python](https://www.youtube.com/watch?v=H5o1P8RMiMw) (Tech With Tim, 38 min). Build a mini-dsh yourself and you'll understand every harness.
- [What is MCP?](https://www.youtube.com/watch?v=eur8dUO9mvE) (IBM, 4 min) · [MCP Explained Simply](https://www.youtube.com/watch?v=oblaHqULUHk) (TechWorld with Nana, 30 min)
- [Context Engineering Explained](https://www.youtube.com/watch?v=BBPQYtR7oUk) (Google Cloud Tech, 10 min)
- [AI Engineering with Chip Huyen](https://www.youtube.com/watch?v=98o_L3jlixw) (The Pragmatic Engineer, 1h15)

**Proof you're done:** A deployed RAG or agent app, with evals, and a README a stranger could follow.

---

### Parallel track (weekends, optional): how LLMs work under the hood

You don't need this to get hired. It's what separates people who *understand* AI from people who just *use* it. Read section 2 first for the big picture.

**Watch, in order**

1. [But What Is a Neural Network?](https://www.youtube.com/watch?v=aircAruvnKk) → [Gradient Descent](https://www.youtube.com/watch?v=IHZwWFHWa-w) → [Backpropagation](https://www.youtube.com/watch?v=Ilg3gGewQ5U) (3Blue1Brown)
2. [Transformers, the Tech Behind LLMs](https://www.youtube.com/watch?v=wjZofJX0v4M) → [Attention in Transformers, Step by Step](https://www.youtube.com/watch?v=eMlx5fFNoYc) (3Blue1Brown)
3. [Deep Dive into LLMs like ChatGPT](https://www.youtube.com/watch?v=7xTGNNLPyMI) (Karpathy, 3.5h, the best single explainer)
4. [Let's Build GPT: From Scratch, in Code, Spelled Out](https://www.youtube.com/watch?v=kCc8FmEb1nY) (Karpathy) · [Let's Build the GPT Tokenizer](https://www.youtube.com/watch?v=zduSFxRajkE)
5. [How Transformer LLMs Work (free course)](https://www.youtube.com/watch?v=k1ILy23t89E) (Jay Alammar)

**Code along:** [karpathy/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero), [karpathy/nanoGPT](https://github.com/karpathy/nanoGPT), [rasbt/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)

---

### Phase 4: Specialize (weeks 17–24+)

This is where your specialization track starts (see section 9 for all the tracks). The AI + Security track is laid out in full below as an example.

#### The AI + Security track

This field is short on people. Most security people don't understand LLMs, and most AI builders don't think about attacks.

#### 4a. Security fundamentals (you can't secure what you don't understand)
- **Networking:** TCP/IP, DNS, HTTP, subnetting. Use AI tutor mode plus labs.
- **[OWASP Top 10](https://owasp.org/www-project-top-ten/):** the ten most critical web application risks. Learn these first.
- **[PortSwigger Web Security Academy](https://portswigger.net/web-security):** free, and the best web-security training anywhere.
- **[TryHackMe](https://tryhackme.com):** beginner paths (Pre-Security, then Cyber Security 101).
- **[picoCTF](https://picoctf.org):** free beginner CTFs. Move on to [HackTheBox](https://www.hackthebox.com) once you're comfortable.
- **Optional cert: CompTIA Security+.** It's the common HR filter. Check which exam version is current before you start studying.

**Watch**

- [How I Would Learn Cyber Security if I Could Start Over in 2026 (6-Month Plan)](https://www.youtube.com/watch?v=4gr5m1xz0Ds) (UnixGuy)
- [You SUCK at Subnetting, EP 1](https://www.youtube.com/watch?v=5WfiTHiU4x8) (NetworkChuck series)
- [Professor Messer's Security+ SY0-701 full course](https://www.youtube.com/playlist?list=PLG49S3nxzAnl4QDVqK-hOnoqcSKEIDDuv) (free, 15h)
- [BEGINNER Capture The Flag: PicoCTF "Obedient Cat"](https://www.youtube.com/watch?v=P07NH5F-t3s) (John Hammond)

#### 4b. LLM and agent security
- **[OWASP Top 10 for LLM Applications](https://genai.owasp.org/llm-top-10/):** prompt injection, sensitive data leakage, excessive agency, and more.
- **[OWASP Top 10 for Agentic Applications (2026)](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/):** risks specific to AI agents that use tools, keep memory, and act on their own.
- **[Gandalf by Lakera](https://gandalf.lakera.ai):** a prompt-injection game that builds intuition fast.
- **[PortSwigger: Web LLM attacks](https://portswigger.net/web-security/llm-attacks):** real exploitation labs against LLM-powered features.
- **[MITRE ATLAS](https://atlas.mitre.org):** the ATT&CK-style attack framework for AI systems.
- **Tools (all open source):**
  - [NVIDIA/garak](https://github.com/NVIDIA/garak): LLM vulnerability scanner. Easiest place to start.
  - [promptfoo/promptfoo](https://github.com/promptfoo/promptfoo): red-teaming and evals that run in CI.
  - [Azure/PyRIT](https://github.com/Azure/PyRIT): Microsoft's framework for automated multi-turn attack campaigns.
  - [UKGovernmentBEIS/inspect_ai](https://github.com/UKGovernmentBEIS/inspect_ai): structured, reproducible evaluations.

**Watch**

- [Prompt Injection, Explained](https://www.youtube.com/watch?v=FgxwCaL6UTA) (Simon Willison, who coined the term)
- [What Is a Prompt Injection Attack?](https://www.youtube.com/watch?v=jrHRe9lSqqA) (IBM, 11 min)
- [OWASP's Top 10 Ways to Attack LLMs](https://www.youtube.com/watch?v=gUNXZMcd2jU) (IBM, 25 min)
- [Securing AI Agents: Preventing Hidden Prompt Injection](https://www.youtube.com/watch?v=5ZA1lTxTH3c) (IBM, 10 min)
- [Anatomy of an AI Attack: MITRE ATLAS](https://www.youtube.com/watch?v=QhoG74PDFyc) (IBM, 9 min)
- [AI Red Teaming 101: Full Course (Episodes 1–10)](https://www.youtube.com/watch?v=DwFVhFdD2fs) (Microsoft Developer, 1h17)
- [Agentic AI Red Teaming: The Hottest Cyber Skill of 2026](https://www.youtube.com/watch?v=SFOrnrxWTNw) (Cloud Security Guy)
- [LLM Vulnerability Scanning with garak: Test Your Own Chatbots](https://www.youtube.com/watch?v=f713_sFqItY) (Embrace The Red)
- [Promptfoo Red Teaming: A Beginner's Guide](https://www.youtube.com/watch?v=y6Dlsz5P8s8) (Jason Koo, 30 min)
- [PortSwigger Lab: Exploiting LLM APIs with Excessive Agency](https://www.youtube.com/watch?v=kd3gBu1mC4Y) (lab walkthrough. Try the lab yourself first.)

#### 4c. The portfolio piece that gets attention
Point garak and promptfoo at **the chatbot or agent you built in Phase 3**. Find the vulnerabilities, fix them, and publish a redacted **red-team report**: findings, severity (critical/high/medium/low), fixes, and before/after results.

"I built an AI app, attacked it, and hardened it" is a portfolio most applicants don't have.

> **Legal line:** Only attack systems you own or have written permission to test. Unauthorized access is a crime under the **PH Cybercrime Prevention Act (RA 10175)**. CTFs, labs, and your own apps are fair game.

---

## 8. Landing your first client or job

The first paying client or job is where all of this starts to make sense. Real constraints, real feedback, real money, and proof that you can deliver. **Don't wait until you finish the whole path.** Start reaching out by month 3, sooner if an opportunity shows up.

### The fast track: earn while you learn

| When | Focus | What you should have |
|---|---|---|
| **Weeks 1–2** | Phase 0: AI basics, tools set up | You use AI daily; your tutor setup works |
| **Weeks 3–6** | Phase 1 concepts, plus your first automation **for yourself** | One working system that saves *you* time |
| **Weeks 7–10** | Phase 2: n8n + AI integrations, and build something **for a real person** (a friend's business, a relative's shop, even for free) | 2–3 working systems, one used by a real business |
| **Month 3: the milestone** | **Package your work and start applying / reaching out** | A GitHub portfolio with 3+ systems, 2 written case studies with diagrams, a short demo video of each |
| **Months 4–6** | Phase 3 (agents, RAG, MCP) and Phase 4 (your specialization track) **while** doing paid work | Your first paid client or job, and harder projects in your portfolio |

**This can go faster.** If someone you know has a painful problem in week 4, solve it in week 4. The phases are a learning order, not a waiting room.

### What you can realistically offer after ~3 months

You won't be a senior engineer. You *will* be able to solve real, common problems that small businesses pay for:

- **Lead and inquiry handling:** auto-reply to Messenger, website, or email inquiries, qualify the lead, and notify a human
- **Booking and scheduling automation:** turn messages and forms into calendar entries and reminders
- **Data entry elimination:** extract data from emails, PDFs, receipts, or forms into a spreadsheet or CRM
- **Content workflows:** draft social posts, product descriptions, or newsletters from a business's own material, with a human approving
- **Internal knowledge assistant:** a simple Q&A bot over a company's documents and FAQs
- **Reporting:** pull data from a few tools and send a daily or weekly summary

Every one of these is the same loop from section 4: messy input → process → structured output.

### Build a portfolio that gets you hired

Clients and employers don't hire certificates. They hire **proof you've solved a problem like theirs.**

**For each system you build, publish a GitHub repo with a README that covers:**

1. **The problem:** who had it, and what it cost them (time, money, missed leads)
2. **The system map:** a before/after diagram (see the diagram tools in section 4)
3. **How it works:** the flow, the tools, and where AI is used and where humans stay in the loop
4. **Results:** numbers wherever possible (hours saved, response time, error rate)
5. **What I'd improve:** this shows judgment, and it's what interviewers ask about
6. **Tech used**

This is your AI Engineering Journal (section 15) turned outward. Add a **2-minute demo video** (screen recording with your voice) for each one; a free tool like [Loom](https://www.loom.com) works.

**Then share it:** post short write-ups on [LinkedIn](https://www.linkedin.com) and in relevant Facebook groups for business owners. Focus on *the problem solved*, not the tech.

### Where to find your first client or job

- **People you already know.** Family businesses, friends' shops, former employers or coworkers. The warmest leads close fastest.
- **Local businesses with obvious manual work:** clinics, salons, real estate agents, schools, restaurants, online sellers. Anyone drowning in Messenger inquiries is a candidate.
- **Facebook groups and online communities** where business owners hang out and ask for help.
- **LinkedIn:** both for job posts and for posting your case studies.
- **[OnlineJobs.ph](https://www.onlinejobs.ph):** the biggest platform for Filipinos working remotely for overseas businesses. Search for automation and AI roles.

**Job titles to search:** AI Automation Specialist, Automation Engineer (n8n / Make / [Zapier](https://zapier.com)), AI Implementation Specialist, AI Operations, Technical VA (AI/automation), Junior AI Engineer, Solutions Engineer. For the security track: SOC Analyst (Tier 1), Security Analyst, and later AI Security / AI Red Team roles.

### How to land the first client

1. **Start with a discovery conversation, not a pitch.** Ask the systems-thinking questions from section 4: What's eating your team's time? Where do leads or orders get lost? What do you do by hand every day?
2. **Map their process.** Draw the current flow in Excalidraw, mark the bottleneck, and show them a "before and after." A clear diagram often sells the project better than any demo.
3. **Propose a small pilot.** One problem, one system, a short timeline, a fixed price. Small and specific beats big and vague.
4. **Deliver, measure, and document.** Track the before/after numbers from day one.
5. **Ask for a testimonial and permission to use it as a case study.** That's what gets you client number two.

**On pricing:** your first one or two projects can be cheap or even free, *in exchange for* a testimonial, a referral, and permission to publish the case study. After that, charge properly. Price by the value of the problem solved, not your hours. Raise your rates after every few successful projects.

### Applying for jobs

- **Lead with your portfolio, not your background.** "Here are 3 systems I built and what they achieved" beats any course list.
- **Expect to be tested on how you think:** walk through a system map, explain your trade-offs, and talk about what broke and how you fixed it.
- **Expect to be asked how you verify AI output.** Have a real answer: tests, reviews, evals, human checkpoints (see "You own the output" in section 4).
- **Certifications help with HR filters** (especially Security+ for security roles), but they won't get you hired without proof of work.

### Using your skills and getting better

Landing the first client isn't the finish line. It's where the compounding starts:

| Loop | What it means in practice |
|---|---|
| **Build → Ship** | Every project goes to real users, even if it's small |
| **Document** | Journal entry and a case study for every project |
| **Share** | Post what you solved; that's how the next client finds you |
| **Reuse** | Turn what you built into templates and reusable workflows, so project #5 takes a fraction of the time of project #1 |
| **Level up** | Take on one thing harder than last time with each new project: agents, RAG, security, bigger clients |
| **Review** | After every project: what went well, what broke, what you'd charge next time |

- **Notice patterns.** When several clients have the same problem, that's a specialty, or eventually a product (see the survivorship-bias section in section 4 for why this order matters).
- **Specialize over time.** Pick an industry (clinics, real estate, e-commerce) or a skill (AI security, agents) and become the go-to person for it. Specialists charge more than generalists.
- **Keep the learning loop running.** Use the daily routine (section 15) even after you're working. Protect a few hours a week for learning something you're not paid for yet.

---

## 9. Pick a specialization track

The core path (sections 3–8) gives you the foundation: learning with AI, systems thinking, data handling, automation, and agents. Then you **specialize**. Specialists get hired faster and paid more than generalists, because businesses search for "someone who can fix our CRM" or "someone who can secure our AI," not "someone who knows AI."

You can pick one track, combine two, or switch later. The core skills carry over to all of them. **Phase 4 of the path is where your track starts.**

### The tracks at a glance

| Track | You'll be the person who… | Common job titles |
|---|---|---|
| **1. AI Automation & Integrations** | Connects a business's tools and puts AI into its workflows | AI Automation Specialist, Automation Engineer, AI Implementation Specialist |
| **2. CRM Management, Automation & Development** | Runs, automates, and extends the system where a business tracks its customers | CRM Specialist, CRM Automation Specialist, HubSpot / Salesforce / GoHighLevel Admin or Developer |
| **3. GTM Engineering** | Builds the data and automation behind sales outreach | GTM Engineer, RevOps / Growth Engineer |
| **4. AI + Security** | Finds and fixes weaknesses in apps and AI systems | SOC Analyst, Security Analyst, AI Security / AI Red Team |
| **5. DevOps / DevSecOps** *(my track)* | Ships, runs, monitors, and secures software in production | DevOps Engineer, Platform Engineer, SRE, DevSecOps Engineer |
| **6. AI-Assisted Software Development** *(my track)* | Builds real software with AI agents, professionally | Software Engineer, AI Engineer, Full-Stack Developer |
| **7. Other tracks** | Voice AI, AI content and creative, data and analytics, vertical (industry) AI | Varies |

---

### Track 1: AI Automation & Integrations

This is the default track the core path already builds toward: n8n, AI integrations, agents, MCP, and RAG for businesses.

- **Go deeper:** advanced n8n (error handling, sub-workflows, queues, self-hosting), building your own MCP servers, evals for production automations, and reusable templates for common client problems.
- **Pairs well with:** CRM (Track 2) or GTM (Track 3). Most automation work touches a CRM eventually.

### Track 2: CRM Management, Automation & Development

Every business that sells something has a CRM (Customer Relationship Management system), or a messy spreadsheet that should be one. CRM work is steady, in demand, and a natural home for AI automation.

**Three levels of the same track:**

| Level | What you do | Example tasks |
|---|---|---|
| **CRM management** | Set up and run the CRM so the team actually uses it | Pipelines, deal stages, contact properties, permissions, data cleanup, reports and dashboards |
| **CRM automation** | Automate what happens inside and around the CRM | Lead routing, follow-up sequences, appointment reminders, AI that qualifies leads or summarizes calls, syncing with forms and Messenger via n8n |
| **CRM development** | Extend the CRM with code and integrations | Custom objects, APIs and webhooks, custom apps, connecting the CRM to other systems, data migrations |

**The main platforms:**

- [**HubSpot**](https://academy.hubspot.com): popular with small and mid-size businesses. **HubSpot Academy** has free courses and certifications.
- [**Salesforce**](https://trailhead.salesforce.com): the enterprise standard, with the biggest job market. **Trailhead** is free and gamified, with Admin and Developer paths.
- [**GoHighLevel**](https://www.gohighlevel.com): an all-in-one CRM and marketing platform widely used by agencies and small businesses.

**Why AI makes this track stronger:** CRMs are full of messy data and repetitive steps, which is exactly the "data in → process → structured data out" loop. AI can enrich contacts, score leads, summarize conversations, and draft follow-ups, with the CRM as the system of record.

**Watch:**

- [HubSpot CRM Tutorial for Beginners](https://www.youtube.com/watch?v=t8QM5zunC44) (Metics Media, 20 min)
- [The Only GoHighLevel Tutorial You Need](https://www.youtube.com/watch?v=y2-fP6LiIa8) (Charlie Chang, 22 min)

### Track 3: GTM Engineering

**GTM (go-to-market) engineering** treats sales outreach as a system to be engineered. A GTM engineer builds data pipelines, enrichment workflows, and automated outbound sequences, work that used to take a team of researchers and sales assistants. The role grew up around [**Clay**](https://www.clay.com), a tool for pulling data from many sources, enriching it with AI, and feeding it into outreach and the CRM.

**Typical work:** finding target companies and contacts, enriching them with data (company size, tech used, recent news), writing personalized outreach with AI at scale, and syncing everything into the CRM.

**I tried this track myself** (November–December 2025): I used AI and tools like Clay to build outbound outreach systems for businesses. It's a strong, well-paid track. I didn't pursue it fully because I prefer working on the sidelines, building systems rather than being close to the sales front line. That's a good example of how to choose: **try a track on a real project before committing to it.**

**Watch out for:** spam and privacy laws. Mass outreach has to follow rules like the **PH Data Privacy Act (RA 10173)**, GDPR in Europe, and CAN-SPAM in the US. Good GTM engineering is targeted and relevant, not mass spam.

**Learn it:** [Clay University](https://www.clay.com/university) (free)

**Watch:**

- [Full Clay.com Course (6+ Hours)](https://www.youtube.com/watch?v=1JiLlbgyVWo) (Tim Yakubson and Clay)
- [The Only Clay Tutorial You Need (For Beginners)](https://www.youtube.com/watch?v=kY5E4wl8wlA) (Xavier Caffrey, 39 min)

### Track 4: AI + Security

Covered in full in **Phase 4** of section 7: security fundamentals, LLM and agent security, open-source red-teaming tools, and a red-team report on your own app as the portfolio piece.

**The OWASP references every security person should know** ([OWASP](https://owasp.org) is the nonprofit that publishes the industry's standard security guides, all free):

| Resource | What it's for |
|---|---|
| [**OWASP Top 10**](https://owasp.org/www-project-top-ten/) | The ten most critical web application security risks. The baseline. |
| [**OWASP Top 10 for LLM Applications**](https://genai.owasp.org/llm-top-10/) | The biggest risks in apps built on language models: prompt injection, data leakage, excessive agency, and more |
| [**OWASP Top 10 for Agentic Applications (2026)**](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/) | Released December 2025: risks for AI agents that plan, use tools, keep memory, and act on their own, like goal hijacking, tool misuse, and rogue agents. Built from real 2025 incidents. |
| [**OWASP Juice Shop**](https://owasp.org/www-project-juice-shop/) | A deliberately insecure web app to practice hacking on, legally |
| [**OWASP Cheat Sheet Series**](https://cheatsheetseries.owasp.org) | Short, practical "how to do this securely" guides |
| [**OWASP ASVS**](https://owasp.org/www-project-application-security-verification-standard/) | A checklist for verifying an application's security, level by level |

### Track 5: DevOps / DevSecOps *(my track)*

DevOps is about **getting software into production and keeping it running**: deploying it, monitoring it, fixing it, and keeping the cloud bill honest. DevSecOps builds security into every step instead of bolting it on at the end. AI can write a lot of infrastructure code, but it can't own an outage at 2am. Someone has to understand the system, and that's this track.

**What you learn, roughly in order:**

1. Linux and the terminal
2. Networking (DNS, HTTP, ports, TLS)
3. Git and scripting
4. Containers (Docker)
5. CI/CD pipelines ([GitHub Actions](https://docs.github.com/en/actions))
6. Cloud fundamentals (AWS, Azure, or GCP)
7. Infrastructure as Code ([Terraform](https://developer.hashicorp.com/terraform/tutorials))
8. Kubernetes ([official tutorials](https://kubernetes.io/docs/tutorials/))
9. Observability (logs, metrics, alerts)
10. Security: secrets management, least privilege, scanning, using the [OWASP DevSecOps Guideline](https://owasp.org/www-project-devsecops-guideline/)

**Map:** [roadmap.sh/devops](https://roadmap.sh/devops)

**Free structured option:** the [AWS courses on TESDA](https://e-tesda.gov.ph/course/index.php?categoryid=2111), especially **AWS Certified Cloud Practitioner (CLF-C02)**, a good first cloud certification for this track. It's on my own list too.

**Watch:**

- [Complete DevOps Roadmap 2026: Master These 4 Levels](https://www.youtube.com/watch?v=1J2YOV6LcwY) (TechWorld with Nana, 38 min)
- [If I Would Start DevOps From 0](https://www.youtube.com/watch?v=Cpy20DnIDTI) (TechWorld with Nana, 10 min)

### Track 6: AI-Assisted Software Development *(my track)*

This is building **real, production software with AI agents**, the professional version of "vibe coding." The difference is the process: specs, plans, tests, reviews, and ownership of what ships (see "You own the output" in section 4).

**What you learn:**

- **The agentic workflow:** [Claude Code](https://claude.com/product/claude-code), dsh, or OpenCode, plus `AGENTS.md` / `CLAUDE.md` files that teach the agent your codebase
- **Spec-driven development:** PRD → plan → small tasks → build → test → review (section 4)
- **Software concepts that let you direct agents:** frontend vs. backend, APIs, databases, authentication, testing, version control, deployment
- **Reading and reviewing code:** your main job becomes reviewing what agents write, catching mistakes, and making architecture decisions
- **Security basics for builders:** the OWASP Top 10 and Cheat Sheets from Track 4

**Maps:** [roadmap.sh/full-stack](https://roadmap.sh/full-stack), [roadmap.sh/backend](https://roadmap.sh/backend), [roadmap.sh/ai-engineer](https://roadmap.sh/ai-engineer)

**Pairs well with:** DevOps (Track 5). Building software *and* knowing how to ship and run it is a strong combination, and it's the one I chose.

### Track 7: Other tracks worth knowing

| Track | What it is | Where to start |
|---|---|---|
| **Voice AI** | Phone and voice agents: receptionists, call handling, dictation | Voice agent platforms plus n8n; Phase 3 agent skills |
| **AI content and creative** | Image, video, and audio generation pipelines for brands | [ComfyUI](https://github.com/Comfy-Org/ComfyUI) and the local-model setup in section 5 |
| **Data and analytics** | Turning business data into dashboards, reports, and decisions | SQL, spreadsheets, and AI-assisted analysis; [roadmap.sh/ai-data-scientist](https://roadmap.sh/ai-data-scientist) |
| **Vertical (industry) AI** | Becoming *the* AI person for one industry: clinics, real estate, legal, e-commerce | Pick an industry you know, and solve its most common problem again and again |

### How to choose

- **Try before you commit.** Do one small real project in a track before going deep, like I did with GTM engineering.
- **Pick what you'd enjoy doing on a bad day.** Some people love being close to sales and clients (GTM, CRM); others prefer the sidelines, building and running systems (DevOps, software development).
- **Follow demand you can see.** If clients keep asking you for the same thing, that's your track choosing you.
- **Combine adjacent tracks:** automation + CRM, GTM + CRM, security + DevOps, and software development + DevOps all pair naturally.

---

## 10. Roadmaps (roadmap.sh)

Use these as **maps to see where you are**, not as checklists to finish. Each node links to free resources.

| Roadmap | Use it for |
|---|---|
| [AI Engineer](https://roadmap.sh/ai-engineer) | The main map for this whole path |
| [AI Agents](https://roadmap.sh/ai-agents) | Phase 3 depth |
| [Prompt Engineering](https://roadmap.sh/prompt-engineering) | Phase 0 depth |
| [AI Red Teaming](https://roadmap.sh/ai-red-teaming) | **Phase 4b. The closest match to AI + Security.** |
| [Cyber Security](https://roadmap.sh/cyber-security) | Phase 4a fundamentals |
| [Linux](https://roadmap.sh/linux) · [Git & GitHub](https://roadmap.sh/git-github) · [Python](https://roadmap.sh/python) | Phase 1 |
| [DevOps](https://roadmap.sh/devops) | If you lean toward DevSecOps later |
| [MLOps](https://roadmap.sh/mlops) · [AI & Data Scientist](https://roadmap.sh/ai-data-scientist) | If you go deeper into the ML side |
| [Computer Science](https://roadmap.sh/computer-science) | Filling gaps later, not now |

---

## 11. Books

> **About "free PDF" links:** I'm not linking pirated copies. Most of those sites are unofficial uploads, and pirated-PDF sites are a common way people get malware, which is a bad habit for someone heading into security. Every option below is **legally free**, or an official free companion to a paid book, and in practice the companion notebooks are where most of the learning happens anyway.

### O'Reilly books with free official companions
| Book | Free part | Phase |
|---|---|---|
| *Hands-On Large Language Models* (Alammar & Grootendorst) | All the code notebooks: [HandsOnLLM/Hands-On-Large-Language-Models](https://github.com/HandsOnLLM/Hands-On-Large-Language-Models), which run in Colab | 3 / parallel |
| *Deep Learning for Coders with fastai and PyTorch* (Howard & Gugger) | **The full book draft as notebooks:** [fastai/fastbook](https://github.com/fastai/fastbook), plus the free [fast.ai course](https://course.fast.ai) | Parallel |
| *AI Engineering* (Chip Huyen) | Companion resources and reading lists: [chiphuyen/aie-book](https://github.com/chiphuyen/aie-book). This is the best book on building with foundation models. | 3 |
| *Designing Machine Learning Systems* (Chip Huyen) | Companion repo: [chiphuyen/dmls-book](https://github.com/chiphuyen/dmls-book) | Later |
| *Hands-On Machine Learning*, 3rd ed. (Géron) | All notebooks: [ageron/handson-ml3](https://github.com/ageron/handson-ml3) | Parallel |
| *Natural Language Processing with Transformers* (Tunstall et al.) | All notebooks: [nlp-with-transformers/notebooks](https://github.com/nlp-with-transformers/notebooks) | Parallel |
| *The Developer's Playbook for Large Language Model Security* (Steve Wilson, OWASP LLM Top 10 lead) | Paid. The OWASP LLM Top 10 docs above are the free version of its core. | 4b |

**Legal shortcut:** O'Reilly Learning has a **free trial**. Plan a reading sprint around it (e.g. *AI Engineering* plus the LLM security playbook) and cancel before it bills.

### Completely free, legal full books
| Book | Link | Phase |
|---|---|---|
| *Automate the Boring Stuff with Python* | [automatetheboringstuff.com](https://automatetheboringstuff.com) | 1 |
| *The Linux Command Line* (Shotts) | [linuxcommand.org/tlcl.php](https://linuxcommand.org/tlcl.php) | 1 |
| *Pro Git* | [git-scm.com/book](https://git-scm.com/book) | 1 |
| *Understanding Deep Learning* (Prince, MIT Press) | [udlbook.github.io/udlbook](https://udlbook.github.io/udlbook/) | Parallel |
| *Dive into Deep Learning* | [d2l.ai](https://d2l.ai) | Parallel |
| *Neural Networks and Deep Learning* (Nielsen) | [neuralnetworksanddeeplearning.com](http://neuralnetworksanddeeplearning.com) | Parallel |
| *Site Reliability Engineering* (Google) | [sre.google/books](https://sre.google/books/) | DevSecOps later |
| *Build a Large Language Model (From Scratch)* (Raschka, Manning) | Book is paid. **All code is free:** [rasbt/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch) | Parallel |

### Essays worth more than most books
- [Building Effective Agents](https://www.anthropic.com/engineering/building-effective-agents) (Anthropic)
- [Simon Willison's prompt-injection series](https://simonwillison.net/tags/prompt-injection/)

---

## 12. GitHub repos to work with

| Repo | What it's for |
|---|---|
| [deepseek-ai/deepseek-harness](https://github.com/deepseek-ai/deepseek-harness) | dsh itself. Read `AGENTS.md` and `docs/architecture.md` to learn how a real harness is built. |
| [Aider-AI/aider](https://github.com/Aider-AI/aider) | Git-native terminal coding agent |
| [ollama/ollama](https://github.com/ollama/ollama) | Run local models |
| [n8n-io/n8n](https://github.com/n8n-io/n8n) | Self-hostable workflow automation |
| [anthropics/courses](https://github.com/anthropics/courses) | Prompt engineering and tool-use courses |
| [anthropics/claude-cookbooks](https://github.com/anthropics/claude-cookbooks) · [openai/openai-cookbook](https://github.com/openai/openai-cookbook) | Copy-and-learn recipes |
| [microsoft/generative-ai-for-beginners](https://github.com/microsoft/generative-ai-for-beginners) | 21-lesson GenAI course |
| [microsoft/ai-agents-for-beginners](https://github.com/microsoft/ai-agents-for-beginners) | Agents course |
| [dair-ai/Prompt-Engineering-Guide](https://github.com/dair-ai/Prompt-Engineering-Guide) | Prompting reference |
| [karpathy/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero) · [karpathy/nanoGPT](https://github.com/karpathy/nanoGPT) | Build it yourself to understand it |
| [rasbt/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch) | LLM internals, step by step |
| [HandsOnLLM/Hands-On-Large-Language-Models](https://github.com/HandsOnLLM/Hands-On-Large-Language-Models) | O'Reilly book notebooks |
| [NVIDIA/garak](https://github.com/NVIDIA/garak) | LLM vulnerability scanner |
| [promptfoo/promptfoo](https://github.com/promptfoo/promptfoo) | Evals and red-teaming |
| [Azure/PyRIT](https://github.com/Azure/PyRIT) | Automated AI red-teaming |
| [UKGovernmentBEIS/inspect_ai](https://github.com/UKGovernmentBEIS/inspect_ai) | Evaluation framework |

**How to use a repo to learn, not just to collect stars:** clone it, open your agent, and ask *"Walk me through this codebase like I'm new. What's the entry point? What are the 3 most important files?"* Then change one small thing yourself and see what breaks.

---

## 13. Shortcuts and workarounds

- **Do the $10 OpenRouter top-up once.** It's the biggest return on money in this whole guide.
- **Use free government courses.** TESDA's online program has [free AWS courses](https://e-tesda.gov.ph/course/index.php?categoryid=2111) on AI fundamentals, Cloud Practitioner, and AI Practitioner.
- **Rotate free models.** Each OpenRouter free model has its own daily quota, and OpenCode Zen's free list rotates too.
- **Route models by task:** cheap models (DeepSeek V4 Flash or free ones) for grunt work, frontier models for hard thinking.
- **Follow one curated feed** instead of fifty sources: [my AI Reddit multireddit](https://www.reddit.com/user/saintjedi/m/ai/) (includes anti-AI subreddits for balance).
- **Turn videos into tests:** paste a tutorial transcript into AI and ask for a quiz plus a hands-on lab. Watching alone isn't learning.
- **Learn from real repos** with an agent as your guide, not only from toy tutorials.
- **Use Kaggle and Colab** for GPU work instead of buying hardware early.
- **Build in public:** a GitHub repo plus short write-ups of *what broke and how you fixed it* will do more for you than certificates.
- **When a free tier dies**, check [OpenRouter's free LLM API comparison](https://openrouter.ai/blog/tutorials/free-llm-apis-compared/) for current limits.
- **When stuck, in this order:**
  1. Ask AI with full context: the error, what you tried, the relevant code
  2. The official docs, or the tool's GitHub Issues and Discussions
  3. **Ask Jude** on Messenger. I'd rather you ask than give up.

---

## 14. Traps to avoid

- **Waiting until you feel "ready."** You won't. Start reaching out by month 3 with what you've built.
- **Hoarding roadmaps**, including this one. Pick one and finish it.
- **Chasing every new tool launch.** Fundamentals stay; tools change monthly. dsh is great today, but learn the *harness concept*, not just dsh.
- **Paying "AI gurus" for courses or bootcamps** early on. Almost everything here is free or close to it.
- **Agent-driven copy-paste** without understanding what you ship.
- **Pasting secrets into free models.**
- **Downloading pirated PDFs and "cracked" tools.** It's a common malware vector, and a bad look for a security person.
- **Attacking systems you don't own** (see RA 10175).
- **Treating a chatbot as a friend, therapist, or oracle.** It's a tool that's trained to agree with you. If you notice yourself losing sleep, pulling away from people, or feeling like the AI "gets" a discovery nobody else understands, step away and talk to a person.
- **Copying the SaaS dream from success stories.** You're only seeing the survivors. See section 4.

---

## 15. Daily routine, journal, and progress tracker

### Daily routine (~2 hrs/day)

| Time | What | How to use AI |
|---|---|---|
| **20 min** | **Learn theory:** a concept, a video segment, a book section | Tutor mode: "explain it, then quiz me" |
| **40 min** | **Follow a tutorial or course** | Make sure you can explain every step; ask "why" when something's unclear |
| **60 min** | **Build something yourself:** your own project, not the tutorial's | Pair mode: you drive, AI reviews |
| **20 min** | **Document what you learned** (in your journal, below) | Write it yourself. AI can quiz you on it afterward. |

That last part isn't optional. Writing it down is how what you learned actually sticks. If you can't explain it in writing, you haven't learned it yet. Short on time? Cut the tutorial, never the build or the journal.

### AI Engineering Journal

Keep one entry per project (a Markdown file in your GitHub repo works well). Over 6 months, this journal *becomes* your portfolio and your interview answers.

```
## Project: [name]  —  [date]

What I wanted to build:
What I learned:
What went wrong:
How I fixed it:
What I would improve:
Technologies used:
AI tools used (and for what — tutor / pair / delegate):
```

### Sample week, if you'd rather block hours (~10–12 hrs)
| Day | What |
|---|---|
| Mon–Thu | 1–1.5 hrs of hands-on labs (AI open in tutor mode) |
| Fri | 1 hr of a video or book chapter, then have AI quiz you on it |
| Sat | 3 hrs of project work (the thing you'll show people) |
| Sun | 30 min weekly review: read back your journal entries, then have AI quiz you on the week |

### Progress tracker
- [ ] **Setup:** WSL2/Linux, VS Code, Git, Python, Node. Accounts: GitHub, OpenRouter, ChatGPT, Claude.
- [ ] **Setup:** tutor prompt saved, OpenCode or dsh running with a free model
- [ ] **Phase 0:** I use AI daily and can explain tokens, context window, and hallucination
- [ ] **Phase 1:** Bandit level 15+, can explain every concept in the Phase 1 table, a useful script built with AI that I can explain line by line, everything on GitHub
- [ ] **Phase 2:** CLI chatbot with tool calling, plus a working n8n workflow
- [ ] **Month 3 milestone:** GitHub portfolio with 3+ systems, 2 case studies with diagrams, demo videos. Started applying / reaching out.
- [ ] **First paid client or job**
- [ ] **Phase 3:** deployed RAG or agent app with evals and a README
- [ ] **Phase 4a:** PortSwigger fundamentals, 10+ TryHackMe/picoCTF rooms
- [ ] **Phase 4b:** beaten Gandalf, finished the PortSwigger LLM labs
- [ ] **Phase 4c:** published a red-team report on my own AI app
- [ ] **Parallel:** watched the 3Blue1Brown series and built Karpathy's GPT

### Weekly review log
| Week | Built | Broke | Believe now |
|---|---|---|---|
| 1 | | | |
| 2 | | | |
| 3 | | | |

---

## 16. Glossary

Plain-language definitions of the terms used in this guide, in alphabetical order.

| Term | Definition |
|---|---|
| **Agent** | An AI system that can plan, use tools (search, code, APIs), check results, and keep going until a task is done, instead of only answering one message |
| **AGENTS.md / CLAUDE.md** | A file in a project that tells a coding agent how the codebase works, its rules, and its conventions |
| **AGI** (Artificial General Intelligence) | A hypothetical AI that can do any intellectual task a human can. It doesn't exist yet, and people disagree on when, or if, it will. |
| **AI** (Artificial Intelligence) | The broad field of making computers do things that normally need human intelligence: understanding language, recognizing images, making decisions |
| **AI slop** | Low-effort, generic AI-generated content published without real human input or editing |
| **Alignment** | Making sure an AI system's goals and behavior match what humans actually intend |
| **API** (Application Programming Interface) | A defined way for two programs to talk to each other: one sends a request, the other sends back a response |
| **API key** | A secret password-like string that lets your code use a paid AI service. Never share it or paste it into public places. |
| **Attention** | The transformer mechanism that lets every token weigh how relevant every other token is to it |
| **Benchmark** | A standard test used to compare models (e.g. SWE-bench for coding). Useful, but not the same as performance on *your* task. |
| **BERT** | Google's 2018 encoder-only transformer, built for understanding and classifying text rather than generating it |
| **CFG** (classifier-free guidance) | In image generation, how strictly the output follows your prompt |
| **Chain-of-thought (CoT)** | Getting a model to reason step by step before answering, either through prompting or, in reasoning models, through training |
| **Chatbot** | A program you talk to in conversation. ChatGPT and Claude are chatbots built on LLMs. |
| **CI/CD** | Continuous Integration / Continuous Delivery: automatically testing and deploying code every time it changes |
| **Cloud** | Computing resources (servers, storage, AI models) rented over the internet instead of run on your own machine (e.g. AWS, Azure, Google Cloud) |
| **Context engineering** | Deciding what information (documents, memory, tools, examples) goes into a model's context window so it can do the job well |
| **Context window** | Everything a model can "see" at once: your prompt, the conversation so far, and any documents. Measured in tokens. |
| **CRM** | Customer Relationship Management system: where a business tracks contacts, leads, deals, and customer conversations |
| **Dataset** | A collection of examples used to train or test a model |
| **Decoder** | The generating half of the transformer. Decoder-only models (GPT, Claude, Llama) write text one token at a time. |
| **Deep learning** | Machine learning using neural networks with many layers. The approach behind almost all modern AI. |
| **Denoising** | The diffusion process of removing noise step by step to reveal an image |
| **Dense model** | A model where every parameter works on every token (the opposite of a sparse/MoE model) |
| **Diffusion model** | A model that generates images (or video, audio) by starting from random noise and gradually refining it into a picture. Used by tools like ComfyUI. |
| **Discriminative model** | A model that returns a decision or label (spam / not spam) rather than generating new content. Classic ML and Jev work this way. |
| **DiT** (Diffusion Transformer) | A diffusion model that uses a transformer as its backbone, as in Sora, Stable Diffusion 3, and Flux |
| **Docker / container** | A package that bundles an app with everything it needs to run, so it runs the same on any machine |
| **Embedding** | A list of numbers that represents the meaning of a token, sentence, or document. Similar meanings have similar numbers. |
| **Encoder** | The understanding half of the transformer. Encoder-only models (BERT, ViT) read the whole input to understand or classify it. |
| **Eval** | A test that measures how well an AI system performs, so you can prove it works instead of eyeballing it |
| **Few-shot / zero-shot prompting** | Giving a model a few examples of what you want (few-shot), or none at all (zero-shot) |
| **Fine-tuning** | Further training an existing model on specific data to specialize it |
| **Frontier model** | The most capable models available at a given time, from the leading labs |
| **GAN** (generative adversarial network) | An older image-generation approach where a "forger" network and a "detective" network train against each other |
| **Generative AI** | AI that creates new content (text, images, audio, video, code) rather than only classifying or predicting |
| **GPU** | Graphics Processing Unit: the chip that trains and runs AI models fast, because it does many calculations in parallel |
| **GTM engineering** | Go-to-market engineering: building the data pipelines and automation behind sales outreach, often with tools like Clay |
| **Guardrails** | Rules and filters that keep an AI system's inputs and outputs safe and on-topic |
| **Hallucination** | When a model produces confident output that's false or made up, a side effect of predicting likely text rather than true text |
| **Harness** | The software that turns a model into a working agent: it reads files, runs commands, edits code, and asks permission (e.g. Claude Code, dsh, OpenCode) |
| **Hybrid architecture** | A model that mixes attention layers with faster layers such as SSMs (e.g. Jamba, Griffin) |
| **IaC** (Infrastructure as Code) | Defining servers, networks, and cloud resources in code files (e.g. Terraform) instead of clicking through dashboards |
| **Inference** | Running a trained model to get an output, as opposed to training it |
| **Jailbreak** | A prompt designed to trick a model into ignoring its safety rules |
| **JSON** | A common text format for structured data, made of keys and values. The language most APIs and automations speak. |
| **Knowledge cutoff** | The date a model's training data ends. It doesn't know about events after that unless it can search or you give it the information. |
| **Latency** | How long it takes to get a response |
| **Latent space** | A compressed representation of data (like an image) that models work in because it's much cheaper than raw pixels |
| **LLM** (Large Language Model) | A large transformer trained on huge amounts of text to predict the next token (e.g. GPT, Claude, Gemini, DeepSeek) |
| **Local model** | A model you download and run on your own computer instead of through a cloud API |
| **LoRA** | A small add-on file that teaches a model a specific style or subject without retraining the whole model |
| **Machine learning (ML)** | Teaching computers to learn patterns from data instead of following hand-written rules |
| **Mamba / SSM** (state-space model) | An architecture that reads in sequence and keeps a running summary, so its cost grows in a straight line with input length instead of quadratically |
| **MCP** (Model Context Protocol) | An open standard for connecting AI agents to tools and data sources |
| **Mixture of Experts (MoE)** | A sparse architecture where a router sends each token to only a few specialist sub-networks, giving big capacity at lower running cost |
| **Model** | The trained AI system itself, the thing that takes input and produces output |
| **Multimodal** | A model that handles more than one type of input or output: text, images, audio, video |
| **n8n** | An open-source, self-hostable workflow automation tool for connecting apps and adding AI to processes |
| **Neural network** | A model made of layers of simple connected units ("neurons") that learn by adjusting the strength of their connections |
| **O(n) / O(n²)** | Shorthand for how cost grows with input size. O(n): double the input, double the work. O(n²): double the input, four times the work. |
| **Open source** | Software whose code is public and free to use, modify, and share (e.g. n8n, OpenCode, dsh) |
| **Open-weight model** | A model whose trained weights are published, so anyone can download and run it |
| **Output** | What the model gives back: text, an image, a decision, or code |
| **OWASP** | The Open Worldwide Application Security Project, a nonprofit that publishes free security standards like the OWASP Top 10 |
| **Parameters / weights** | The billions of numbers inside a model that were learned during training. They hold what the model "knows." |
| **PRD** (Product Requirements Document) | A document describing the problem, users, requirements, and what "done" means, written before building |
| **Pretraining** | The first, largest training stage, where a model learns language by predicting tokens across huge amounts of text |
| **Prompt** | The instruction or question you give an AI model |
| **Prompt engineering** | The skill of writing prompts that reliably get good results |
| **Prompt injection** | An attack where malicious instructions hidden in input (a webpage, email, or document) hijack an AI system's behavior |
| **Quantization (e.g. Q4)** | Compressing a model's numbers to lower precision so it fits in less memory, at a small cost in quality |
| **RAG** (Retrieval-Augmented Generation) | Fetching relevant documents and giving them to the model at answer time, so it answers from your data instead of only its memory |
| **Rate limit** | A cap on how many requests you can send to a service in a given time (e.g. 50 requests a day on free tiers) |
| **Reasoning model** | A model trained to think through long, structured reasoning before answering (e.g. o1, DeepSeek-R1, "extended thinking" modes) |
| **Red teaming** | Deliberately attacking your own system to find weaknesses before real attackers do |
| **RLHF** (Reinforcement Learning from Human Feedback) | Training a model using human ratings of its answers, so it becomes more helpful and follows instructions |
| **Sampler** | In diffusion, the method used to remove noise at each step |
| **Sampling** | How the model picks the next token from its probability list: the "dice roll" |
| **Seed** | The starting random noise (or random number) for a generation. Same seed + same settings = the same result. |
| **Spec-driven development** | Writing a clear specification first, then having AI build against it in small, testable steps |
| **Sycophancy** | A model's tendency to agree with and flatter the user, even when the user is wrong |
| **System prompt** | Instructions given to a model before the conversation that set its role, rules, and behavior |
| **T5** | Google's 2019 encoder-decoder transformer that treats every task as "text in, text out" |
| **Temperature** | A setting that controls randomness. Low = predictable and focused; high = more varied and creative. |
| **Test-time compute** | Letting a model spend more computation "thinking" while answering, rather than only making the model bigger |
| **Token** | A small chunk of text (often part of a word) that models read and write. Pricing and context limits are counted in tokens. |
| **Tool calling** | A model's ability to request that a function or API be run (search, calculator, database) and use the result |
| **Top-p** | A sampling setting that only lets the model choose from the smallest set of tokens that together make up probability p (e.g. the top 90%) |
| **Training** | The process of teaching a model by showing it huge amounts of data and adjusting its parameters |
| **Transformer** | The 2017 neural network architecture built on attention that powers nearly all modern AI models |
| **VAE** | The part of an image model that compresses images into latent space and turns them back into pixels |
| **Vector database** | A database that stores embeddings and finds the most similar ones quickly. The search engine behind RAG. |
| **Vibe coding** | Building software by describing what you want to AI and accepting its code without really reviewing it. Fine for prototypes, risky for anything real. |
| **Vision Transformer (ViT)** | A transformer that splits images into patches and treats them like tokens |
| **VRAM** | The memory on a graphics card. It decides which local models you can run. |
| **Webhook** | A URL that one app calls automatically when something happens, to trigger an action in another app |
| **Workflow / automation** | A defined series of steps that runs automatically when triggered, like a form submission creating a CRM contact and sending a reply |
