Learn AI, one skill at a time.
Six areas, from how models work to running AI in production. Follow them in order, or start wherever you need. Each skill lists the short reads that teach it, from beginner to advanced.
Understand the models
What a model is made of, how it learns, and what it can and cannot see.
Explain how a model learns
You can describe how data, a loss and small weight updates turn into a trained model.
- Intermediate
- Gradient descentBackpropagationOverfitting
Know what's inside a model
You can sketch the parts of a transformer and say what each one contributes.
Builds on: Explain how a model learns
- Beginner
- Neural networksTransformersLLM
- Intermediate
- AttentionSelf-attentionEmbeddings
Reason about tokens and context
You can predict what fits in a model's context and why long inputs cost more.
Builds on: Explain how a model learns
- Beginner
- TokensTokenizationContext window
- Intermediate
- Max tokensLong contextByte-pair encoding
Explain how a model becomes an assistant
You can explain how pretraining gives a model broad knowledge and how instruction tuning and human feedback turn it into an assistant that follows requests.
Builds on: Explain how a model learns
- Intermediate
- Supervised fine-tuningRLAIF
- Advanced
- DPOConstitutional AI
Read a training curve
You can spot overfitting and pick sensible training settings from a loss curve.
Builds on: Explain how a model learns
- Beginner
- EpochValidation setTest set
- Intermediate
- Learning rateBatch sizeBias–variance tradeoff
- Advanced
- Chinchilla scaling
Build with models
Prompts, structured outputs, retrieval and tools: the parts of a working AI feature.
Write reliable prompts
You can write prompts that give consistent answers and know which settings change them.
Builds on: Reason about tokens and context
- Intermediate
- Chain of thoughtPrompt chainingTemperature
- Advanced
- Top-p samplingReasoning models
Get structured outputs
You can make a model return data your code can parse every time.
Builds on: Write reliable prompts
- Beginner
- Structured outputs
- Intermediate
- JSON modeStop sequences
- Advanced
- Logits
Ground answers with retrieval
You can connect a model to your documents so answers cite what you actually have.
Builds on: Write reliable prompts
- Beginner
- RAGEmbeddingsSemantic search
- Intermediate
- ChunkingVector databasesHybrid searchReranking
- Advanced
- GraphRAGHyDEContextual retrieval
Give models tools
You can let a model call your functions safely and use the results.
Builds on: Get structured outputs
- Beginner
- Function callingTool use
- Intermediate
- Structured outputs
- Advanced
- MCP
Build agents
Loops, context, memory, harnesses and protocols for models that act.
Design an agent loop
You can build a model that reasons, acts, reads the result and decides again, and know when a fixed workflow is better.
Builds on: Give models tools
- Intermediate
- ReActAgentic workflows
- Advanced
- Agent planningSelf-reflection
Engineer the context
You can decide what goes into a model's context at each step, and what to leave out.
Builds on: Design an agent loop, Reason about tokens and context
- Intermediate
- Long contextPrompt caching
- Advanced
- Contextual retrieval
Keep long contexts useful
You can spot when a long conversation or agent run starts losing track, and trim, summarise or cache context to fix it.
Builds on: Engineer the context
- Beginner
- Context windowLong context
- Advanced
- KV cache
Give agents memory
You can keep what an agent learns across steps and sessions without flooding its context.
Builds on: Engineer the context
- Beginner
- Agent memory
- Intermediate
- Agent stateRAGVector databases
- Advanced
- Knowledge graphs
Build the harness
You can wrap a model in the code that dispatches tools, retries, limits and logs, so it behaves in production.
Builds on: Design an agent loop
- Beginner
- AI agentsSandboxing agents
- Advanced
- Code interpreters
Defend agents against injection
You can explain how instructions hidden in user input, documents or tool results can hijack an agent, and add guardrails and human approval to stop it.
Builds on: Design an agent loop
- Beginner
- Prompt injectionGuardrails
- Advanced
- Red teaming
Connect tools with MCP
You can expose tools and data to any compatible model through one protocol.
Builds on: Give models tools
- Beginner
- MCP
- Intermediate
- Function calling
- Advanced
- A2A protocol
Split work across agents
You can decide when several cooperating agents beat one, and how they hand off work.
Builds on: Build the harness, Design an agent loop, Evaluate agents
- Beginner
- Multi-agent systems
- Intermediate
- Sub-agentsAgent orchestration
- Advanced
- A2A protocol
Measure and improve
How to tell whether an AI system works, and find out why when it does not.
Design evals
You can choose metrics and test sets that tell you whether a change made things better.
Builds on: Write reliable prompts
- Intermediate
- F1 scoreConfusion matrixROC curve
Evaluate and trace an LLM app
You can build a test suite for an LLM feature, grade outputs with a model as judge, and read traces to find where it fails.
Builds on: Design evals
- Beginner
- LLM evaluationEvals
- Intermediate
- LLM-as-a-judgeBenchmarksLLM observability
- Advanced
- Benchmark contamination
Evaluate agents
You can measure whether an agent finishes real tasks, not just whether it sounds right.
Builds on: Design an agent loop, Design evals, Evaluate and trace an LLM app
- Beginner
- Agent evaluation
- Intermediate
- SWE-benchLLM observability
Do error analysis
You can sort failures into causes and fix the ones that matter most.
Builds on: Design evals
- Beginner
- Confusion matrix
- Intermediate
- Data leakageImbalanced data
Catch hallucinations
You can reduce confident wrong answers and detect the ones that slip through.
Builds on: Ground answers with retrieval, Design evals
- Beginner
- HallucinationGrounding
- Intermediate
- RAGKnowledge cutoff
- Advanced
- Sycophancy
Run it in production
Latency, cost, serving and safety once real users arrive.
Cut latency and cost
You can find where time and money go in a model call and bring both down.
Builds on: Reason about tokens and context
Serve models
You can run a model behind an API that stays fast under many users.
Builds on: Cut latency and cost
- Beginner
- Inference
- Intermediate
- Model servingSelf-hosting LLMs
- Advanced
- Speculative decoding
Isolate and secure
You can limit what a model or agent can reach, and keep one customer's data away from another's.
Builds on: Build the harness
- Beginner
- Sandboxing agents
- Intermediate
- Human in the loop
- Advanced
- Constitutional AI
Choose the right stack
Picking models, databases and hosting with the tradeoffs in view.
Pick a model
You can choose between open and closed, large and small, and general and reasoning models for a job.
Builds on: Know what's inside a model
- Advanced
- Frontier models
Pick a vector database
You can choose where to store embeddings for your scale and search needs.
Builds on: Ground answers with retrieval
- Beginner
- Vector databases
Decide when to self-host
You can weigh a hosted API against running open models yourself.
Builds on: Serve models, Pick a model
- Beginner
- Open-weight models
- Intermediate
- Self-hosting LLMsOn-device AI
Pick an AI framework
You can tell what LLM app frameworks and agent SDKs add over raw API calls, and choose one that fits your language, model and level of control.
Builds on: Design an agent loop
- Beginner
- LangChainLlamaIndex
Use a cloud AI platform
You can compare how AWS, Google Cloud and Microsoft offer models, agent hosting and training, and choose one for a project.
Builds on: Pick a model
- Advanced
- AWS TrainiumAWS InferentiaTPUs
New to all of this? Start with the beginner roadmap.