28 Sept 2026
New explainers published, each checked sentence by sentence against its official sources, for originality and for plain language.
- New explainer
It is the starting point for the site's introductory learning path.
- New explainer
Readers need a grounded explanation of tokens, training and generation before following model news.
- New explainer
It explains the model, tool and control loop behind agentic applications.
- New explainer
It gives readers a concrete path from indexed documents to a grounded answer.
- New explainer
Embeddings connect token representations, semantic search and retrieval systems.
- New explainer
It explains how AI applications connect to external tools and data through a shared protocol.
- New explainer
A GPU is a chip built to run many similar calculations at the same time, which suits the matrix maths inside neural networks.
- New explainer
FLOPs count the small arithmetic steps, like one multiply or one add on decimal numbers, that a computer does to train an AI model.
- New explainer
Agent state is the running record an AI agent keeps while it works, so a task can pause, resume, or recover from a crash.
- New explainer
Prompt injection is text that sneaks new instructions into what an AI model reads, so the model behaves in unintended ways.
- New explainer
A CPU is the chip that runs a computer's software, carrying out instructions with just a few cores backed by lots of cache memory.
- New explainer
TPUs are chips Google designed to speed up the maths of machine learning, rented out through Google Cloud.
- New explainer
An NPU is a part of a chip built to speed up AI models at low power.
- New explainer
AI accelerators are chips built to speed up the maths inside AI models, which is largely multiply-add sums.
- New explainer
Guardrails are checks placed around an AI model that inspect what goes in and what comes out, and can block, change or check inputs and replies that are unsafe, off-topic or break the app's rules.
- New explainer
An eval checks an AI system's work. You feed it an input and use a set of rules to score what comes back.
- New explainer
LLM evaluation means testing an AI feature with structured tests, so you can check how accurate and reliable it is even though its answers vary.
- New explainer
Agent evaluation tests an AI agent on set tasks, checking both its final result and the steps it took, over several tries.
- New explainer
Agent orchestration is how an app with several AI agents decides which agent works, in what order, and who picks what happens next.
- New explainer
Amazon Bedrock is an AWS service that lets developers call AI models from several companies through one set of APIs, with extra tools for search and safety.
- New explainer
A benchmark is a fixed set of test tasks with a scoring rule, so different AI models can be compared on the same exam.
- New explainer
Claude Code is Anthropic's coding assistant that reads a project, edits files and runs commands for you, asking before risky steps.
- New explainer
Context compaction swaps the older part of a long AI conversation for a short summary, so the chat can keep going inside the model's memory limit.
- New explainer
Contextual retrieval has a model write a short note that places each chunk in its document, then indexes the note with the chunk so search can find it.
- New explainer
CrewAI is an open-source toolkit, written in Python, for building teams of AI agents, each with a role, that work through a list of tasks together.
- New explainer
Cursor is a coding tool with an AI agent that can read your project, edit files and run commands for you.
- New explainer
Data privacy is about what an AI company may do with your chats, such as training models, and the settings that let you limit it.
- New explainer
Feature engineering turns raw data, like prices and colour names, into the lists of numbers a model can learn from.
- New explainer
GitHub Copilot is an AI helper for programmers that offers the next lines of code while you write, explains code when you ask, and can take on jobs you give it.
- New explainer
Google ADK is an open-source toolkit from Google for writing AI agents in code, giving them tools, a record of each chat and helper agents.
- New explainer
GraphRAG turns your documents into a map of people, places and links, groups and summarises it, then answers questions from that map.
- New explainer
An open-source Python toolkit from deepset that builds AI apps (search, RAG, agents) by linking small parts into a pipeline.
- New explainer
HNSW finds the stored vectors closest to a query quickly by searching a stack of linked graphs, from a sparse top layer down to a dense bottom one.
- New explainer
With this setup, a person checks an AI system's work at chosen points and can approve it, change it, reject it or step in.
- New explainer
A jailbreak is a prompt written to trick an AI model into ignoring its safety training and producing something it would normally refuse.
- New explainer
A knowledge graph stores facts as labelled links between real-world things, so software can follow the links to answer questions.
- New explainer
LangGraph is an open-source framework for building AI agents as graphs, where steps share one state that can be saved, paused and resumed.
- New explainer
LlamaIndex is an open-source toolkit that brings your own files and data to a language model when you ask a question, so apps and agents can answer from them.
- New explainer
The day-to-day work of running an app built on a language model, from managing prompts and testing answers to watching speed and cost.
- New explainer
Language models often use facts near the opening or closing of a long input better than facts buried in the middle.
- New explainer
Metadata filtering limits a vector search to records whose labels, such as category or year, match rules you set.
- New explainer
Microsoft Foundry is Microsoft's Azure platform for building AI apps and agents, with a large model catalogue, tools, testing and safety controls in one place.
- New explainer
Ollama is an open-source program that downloads open AI models and runs them on your own computer, with a simple command and a local API.
- New explainer
PyTorch is a free Python library for building and training neural networks, with fast maths on GPUs and automatic gradients.
- New explainer
A rate limit caps how much you can use an AI service in a set stretch of time, such as how many requests you send each minute.
- New explainer
ReAct is a prompting method in which a language model alternates written thoughts with tool actions, reading each result before it picks the next step.
- New explainer
Red teaming means deliberately testing an AI system the way an attacker would, to find its weak spots.
- New explainer
Sandboxing an agent means running its commands inside a fence, set up by the operating system, that limits which files and websites they can reach.
- New explainer
A sub-agent is a helper AI that a main agent sends off to do one side task in its own workspace, then report back a short result.
- New explainer
Supervised fine-tuning keeps training an already-trained model on example questions paired with good answers, so it learns to answer that way.
- New explainer
AI model APIs charge by the token, with separate prices for the text you send in and the text the model writes back.
- New explainer
LoRA adapts a large model to a new task by freezing its weights and training two thin matrices whose product is added to them.
- New explainer
PEFT is a free Hugging Face library that teaches a big AI model a new task by training a small set of extra weights instead of the whole model.
- New explainer
Vertex AI is Google Cloud’s platform for choosing AI models, tuning them on your own examples and running models and agents for real users.
- New explainer
On-device AI runs the model on your own phone, laptop or browser, so your data does not have to go to a server.
- New explainer
LLM-as-a-judge means asking a strong language model to grade answers to open-ended questions.
- New explainer
Summarization is when a model turns a long text into a shorter version that keeps the important information.
- New explainer
LLM observability means recording what an AI app does on each request, such as the prompt, each step, the time taken and the tokens used.
- New explainer
Prompt caching saves the work a model did on the unchanged start of a prompt, so the next request with the same start is faster and cheaper.
- New explainer
Pydantic AI is a Python framework for building AI agents whose tool inputs and final answers are checked against types you define.
- New explainer
Model serving means running a trained model on a server so apps can send it requests over the network and get answers back.
- New explainer
A KV cache lets a language model save work from earlier tokens, so each new token does not redo the same maths.
- New explainer
OpenAI's Agents SDK is a small open-source toolkit from OpenAI for building AI agents that use tools, pass work to each other and can record each run as a trace.
- New explainer
An LLM gateway is one server that sits between your apps and AI model providers, adding limits, caching, fallbacks and logs to model calls.
- New explainer
JSON mode is an API setting that asks a language model to reply in JSON, text a program can parse (read into data), and on some providers, such as OpenAI, enforces valid syntax, without fixing which keys or types appear.
- New explainer
Context engineering is choosing, trimming and updating the tokens a model sees at each step, so an AI agent works from a small, useful context.
- New explainer
A set of AWS services, called Amazon Bedrock AgentCore, for running AI agents in the cloud, with hosting, memory, sign-in, tool access and monitoring.
- New explainer
Agent memory is how an AI agent saves useful information outside the model and loads the right pieces back into its prompt later.
- New explainer
Anthropic's Claude Agent SDK is a Python and TypeScript library for building agents that run on the same loop, tools and context handling as Claude Code.
- New explainer
Batch inference sends many AI requests as one job that runs in the background, trading an instant reply for a lower price.
- New explainer
A2A is an open protocol for AI agents: one agent can find another, hand it a task and track the result, even when different companies built them.
- New explainer
Benchmark contamination happens when test questions, or close copies of them, end up in a model's training data, so its score can look better than its real skill.
- New explainer
LangChain is an open-source framework for building apps and agents on top of large language models, with one standard way to talk to many AI providers.
- New explainer
An agentic workflow runs a task as several language model and tool steps, joined by a path that your code fixes in advance.
- New explainer
AWS Inferentia is a family of computer chips that Amazon designed to run trained AI models cheaply and quickly in its cloud.
- New explainer
Azure OpenAI lets companies use OpenAI's models through Microsoft's Azure cloud, with Azure billing, safety filters and data controls.
- New explainer
Quantization stores a model's numbers with fewer bits, such as 8-bit integers instead of 32-bit decimals, so it needs less memory.
- New explainer
Agent planning is how an AI agent splits a big task into smaller steps first, then carries out those steps one by one.
- New explainer
AWS Trainium is a computer chip Amazon designed for training and running AI models, rented through special Amazon EC2 cloud servers.
- New explainer
A multi-agent system is a group of AI agents, each a language model using tools in its own loop, that work together on one job.
- New explainer
Approximate nearest neighbor search
Approximate nearest neighbor search finds stored vectors that are very close to a query quickly, by checking only a promising part of the collection and accepting a few misses.
- New explainer
Inference optimization is a set of methods, like caching, quantization and speculative decoding, that let a trained language model answer with less work or less memory.
- New explainer
Strands Agents is a free AWS toolkit, released as open source under Apache 2.0, for making AI agents with Python or TypeScript, where the model itself plans the steps and picks the tools.
- New explainer
A voice agent is an AI app you talk to out loud; it listens, works out what you want, can use tools, and answers in speech.
- New explainer
Agno is open-source software for building AI agents and running them as a live service. You can build single agents, teams of agents and step-by-step workflows.
- New explainer
Indirect prompt injection is when an attacker hides instructions in content an AI app reads later, such as an email or web page, instead of typing them in.
- New explainer
A code interpreter lets an AI model write a small program, run it in a sealed-off sandbox, and use the result in its answer.
- New explainer
AutoGen is an open-source Microsoft framework for building apps where several AI agents talk to each other, and sometimes to people, to finish a task.
- New explainer
Hugging Face runs the Hub, a website where people share AI models, datasets and small demo apps so others can find, download and reuse them.
- New explainer
ComfyUI is a free, open-source app where you build AI image, video and audio generators by wiring boxes called nodes together.
- New explainer
MLOps is a way of working that makes building and releasing machine learning models simpler and more automatic.
- New explainer
QLoRA fine-tunes a large language model by storing its frozen weights in 4 bits and training only a small add-on that sits beside them.
- New explainer
Fine-tuning takes a model that is already trained and keeps training it on a smaller set of examples for one task.
- New explainer
HyDE has a language model write a made-up answer to your question, then searches for real documents that look like that answer.
- New explainer
Model routing sends each prompt to the model that suits it, so simple requests go to cheaper models and hard ones to stronger models.
- New explainer
Open WebUI, a free chat app you host yourself, lets you talk to AI models running on your own computer or in the cloud, from one web page.
- New explainer
SGLang is open-source server software that runs large language models on GPUs and answers many requests quickly, partly by reusing work shared between prompts.
- New explainer
Mastra is an open-source TypeScript toolkit for making AI agents, step-by-step workflows and the apps that use them.
- New explainer
Self-reflection is when an AI model checks its own first answer, writes feedback on it, and then tries again using that feedback.
- New explainer
llama.cpp is an open-source program that runs large language models on your own computer.
- New explainer
Transformers is a free, open-source Python library that lets you download a ready-trained AI model by name and run or train it with a short piece of code.
- New explainer
Self-hosting means running a language model on a computer you control, so your prompts go to your own machine instead of a company's service.
- New explainer
Smolagents is a small open-source Python library from Hugging Face for building AI agents, including ones that act by writing short pieces of Python code.
- New explainer
SageMaker AI, from AWS, lets you build, train and run machine learning models without managing your own servers.
- New explainer
Unsloth is an open-source tool for training and running AI language models on your own computer, built to fine-tune them faster and with less graphics-card memory.
- New explainer
Semantic Kernel is an open-source Microsoft toolkit that connects AI models to your own code, so a model can ask your functions to do real work.
- New explainer
DSPy is a Python framework where you describe an AI job by naming its inputs and outputs, and it tunes the prompts for you.
- New explainer
The Vercel AI SDK is an open-source TypeScript toolkit, free to use, that lets apps talk to AI models from many companies through the same code.
- New explainer
vLLM is open-source software that runs large language models on a server and answers many requests at once, using GPU memory carefully so fewer bytes go to waste.