Site updates

What's new here

New explainers, updates and corrections, newest first. When we get something wrong, the fix is listed here too. Follow by RSS.

28 Sept 2026

New explainers published, each checked sentence by sentence against its official sources, for originality and for plain language.

  • New explainer

    Artificial intelligence

    It is the starting point for the site's introductory learning path.

  • New explainer

    LLM

    Readers need a grounded explanation of tokens, training and generation before following model news.

  • New explainer

    AI agents

    It explains the model, tool and control loop behind agentic applications.

  • New explainer

    RAG

    It gives readers a concrete path from indexed documents to a grounded answer.

  • New explainer

    Embeddings

    Embeddings connect token representations, semantic search and retrieval systems.

  • New explainer

    MCP

    It explains how AI applications connect to external tools and data through a shared protocol.

  • New explainer

    GPUs

    A GPU is a chip built to run many similar calculations at the same time, which suits the matrix maths inside neural networks.

  • New explainer

    FLOPs

    FLOPs count the small arithmetic steps, like one multiply or one add on decimal numbers, that a computer does to train an AI model.

  • New explainer

    Agent state

    Agent state is the running record an AI agent keeps while it works, so a task can pause, resume, or recover from a crash.

  • New explainer

    Prompt injection

    Prompt injection is text that sneaks new instructions into what an AI model reads, so the model behaves in unintended ways.

  • New explainer

    CPUs

    A CPU is the chip that runs a computer's software, carrying out instructions with just a few cores backed by lots of cache memory.

  • New explainer

    TPUs

    TPUs are chips Google designed to speed up the maths of machine learning, rented out through Google Cloud.

  • New explainer

    NPUs

    An NPU is a part of a chip built to speed up AI models at low power.

  • New explainer

    AI accelerators

    AI accelerators are chips built to speed up the maths inside AI models, which is largely multiply-add sums.

  • New explainer

    Guardrails

    Guardrails are checks placed around an AI model that inspect what goes in and what comes out, and can block, change or check inputs and replies that are unsafe, off-topic or break the app's rules.

  • New explainer

    Evals

    An eval checks an AI system's work. You feed it an input and use a set of rules to score what comes back.

  • New explainer

    LLM evaluation

    LLM evaluation means testing an AI feature with structured tests, so you can check how accurate and reliable it is even though its answers vary.

  • New explainer

    Agent evaluation

    Agent evaluation tests an AI agent on set tasks, checking both its final result and the steps it took, over several tries.

  • New explainer

    Agent orchestration

    Agent orchestration is how an app with several AI agents decides which agent works, in what order, and who picks what happens next.

  • New explainer

    Amazon Bedrock

    Amazon Bedrock is an AWS service that lets developers call AI models from several companies through one set of APIs, with extra tools for search and safety.

  • New explainer

    Benchmarks

    A benchmark is a fixed set of test tasks with a scoring rule, so different AI models can be compared on the same exam.

  • New explainer

    Claude Code

    Claude Code is Anthropic's coding assistant that reads a project, edits files and runs commands for you, asking before risky steps.

  • New explainer

    Context compaction

    Context compaction swaps the older part of a long AI conversation for a short summary, so the chat can keep going inside the model's memory limit.

  • New explainer

    Contextual retrieval

    Contextual retrieval has a model write a short note that places each chunk in its document, then indexes the note with the chunk so search can find it.

  • New explainer

    CrewAI

    CrewAI is an open-source toolkit, written in Python, for building teams of AI agents, each with a role, that work through a list of tasks together.

  • New explainer

    Cursor

    Cursor is a coding tool with an AI agent that can read your project, edit files and run commands for you.

  • New explainer

    Data privacy

    Data privacy is about what an AI company may do with your chats, such as training models, and the settings that let you limit it.

  • New explainer

    Feature engineering

    Feature engineering turns raw data, like prices and colour names, into the lists of numbers a model can learn from.

  • New explainer

    GitHub Copilot

    GitHub Copilot is an AI helper for programmers that offers the next lines of code while you write, explains code when you ask, and can take on jobs you give it.

  • New explainer

    Google ADK

    Google ADK is an open-source toolkit from Google for writing AI agents in code, giving them tools, a record of each chat and helper agents.

  • New explainer

    GraphRAG

    GraphRAG turns your documents into a map of people, places and links, groups and summarises it, then answers questions from that map.

  • New explainer

    Haystack

    An open-source Python toolkit from deepset that builds AI apps (search, RAG, agents) by linking small parts into a pipeline.

  • New explainer

    HNSW

    HNSW finds the stored vectors closest to a query quickly by searching a stack of linked graphs, from a sparse top layer down to a dense bottom one.

  • New explainer

    Human in the loop

    With this setup, a person checks an AI system's work at chosen points and can approve it, change it, reject it or step in.

  • New explainer

    Jailbreaks

    A jailbreak is a prompt written to trick an AI model into ignoring its safety training and producing something it would normally refuse.

  • New explainer

    Knowledge graphs

    A knowledge graph stores facts as labelled links between real-world things, so software can follow the links to answer questions.

  • New explainer

    LangGraph

    LangGraph is an open-source framework for building AI agents as graphs, where steps share one state that can be saved, paused and resumed.

  • New explainer

    LlamaIndex

    LlamaIndex is an open-source toolkit that brings your own files and data to a language model when you ask a question, so apps and agents can answer from them.

  • New explainer

    LLMOps

    The day-to-day work of running an app built on a language model, from managing prompts and testing answers to watching speed and cost.

  • New explainer

    Lost in the middle

    Language models often use facts near the opening or closing of a long input better than facts buried in the middle.

  • New explainer

    Metadata filtering

    Metadata filtering limits a vector search to records whose labels, such as category or year, match rules you set.

  • New explainer

    Microsoft Foundry

    Microsoft Foundry is Microsoft's Azure platform for building AI apps and agents, with a large model catalogue, tools, testing and safety controls in one place.

  • New explainer

    Ollama

    Ollama is an open-source program that downloads open AI models and runs them on your own computer, with a simple command and a local API.

  • New explainer

    PyTorch

    PyTorch is a free Python library for building and training neural networks, with fast maths on GPUs and automatic gradients.

  • New explainer

    Rate limits

    A rate limit caps how much you can use an AI service in a set stretch of time, such as how many requests you send each minute.

  • New explainer

    ReAct

    ReAct is a prompting method in which a language model alternates written thoughts with tool actions, reading each result before it picks the next step.

  • New explainer

    Red teaming

    Red teaming means deliberately testing an AI system the way an attacker would, to find its weak spots.

  • New explainer

    Sandboxing agents

    Sandboxing an agent means running its commands inside a fence, set up by the operating system, that limits which files and websites they can reach.

  • New explainer

    Sub-agents

    A sub-agent is a helper AI that a main agent sends off to do one side task in its own workspace, then report back a short result.

  • New explainer

    Supervised fine-tuning

    Supervised fine-tuning keeps training an already-trained model on example questions paired with good answers, so it learns to answer that way.

  • New explainer

    Token costs

    AI model APIs charge by the token, with separate prices for the text you send in and the text the model writes back.

  • New explainer

    LoRA

    LoRA adapts a large model to a new task by freezing its weights and training two thin matrices whose product is added to them.

  • New explainer

    PEFT

    PEFT is a free Hugging Face library that teaches a big AI model a new task by training a small set of extra weights instead of the whole model.

  • New explainer

    Vertex AI

    Vertex AI is Google Cloud’s platform for choosing AI models, tuning them on your own examples and running models and agents for real users.

  • New explainer

    On-device AI

    On-device AI runs the model on your own phone, laptop or browser, so your data does not have to go to a server.

  • New explainer

    LLM-as-a-judge

    LLM-as-a-judge means asking a strong language model to grade answers to open-ended questions.

  • New explainer

    Summarization

    Summarization is when a model turns a long text into a shorter version that keeps the important information.

  • New explainer

    LLM observability

    LLM observability means recording what an AI app does on each request, such as the prompt, each step, the time taken and the tokens used.

  • New explainer

    Prompt caching

    Prompt caching saves the work a model did on the unchanged start of a prompt, so the next request with the same start is faster and cheaper.

  • New explainer

    Pydantic AI

    Pydantic AI is a Python framework for building AI agents whose tool inputs and final answers are checked against types you define.

  • New explainer

    Model serving

    Model serving means running a trained model on a server so apps can send it requests over the network and get answers back.

  • New explainer

    KV cache

    A KV cache lets a language model save work from earlier tokens, so each new token does not redo the same maths.

  • New explainer

    OpenAI Agents SDK

    OpenAI's Agents SDK is a small open-source toolkit from OpenAI for building AI agents that use tools, pass work to each other and can record each run as a trace.

  • New explainer

    LLM gateways

    An LLM gateway is one server that sits between your apps and AI model providers, adding limits, caching, fallbacks and logs to model calls.

  • New explainer

    JSON mode

    JSON mode is an API setting that asks a language model to reply in JSON, text a program can parse (read into data), and on some providers, such as OpenAI, enforces valid syntax, without fixing which keys or types appear.

  • New explainer

    Context engineering

    Context engineering is choosing, trimming and updating the tokens a model sees at each step, so an AI agent works from a small, useful context.

  • New explainer

    Amazon Bedrock AgentCore

    A set of AWS services, called Amazon Bedrock AgentCore, for running AI agents in the cloud, with hosting, memory, sign-in, tool access and monitoring.

  • New explainer

    Agent memory

    Agent memory is how an AI agent saves useful information outside the model and loads the right pieces back into its prompt later.

  • New explainer

    Claude Agent SDK

    Anthropic's Claude Agent SDK is a Python and TypeScript library for building agents that run on the same loop, tools and context handling as Claude Code.

  • New explainer

    Batch inference

    Batch inference sends many AI requests as one job that runs in the background, trading an instant reply for a lower price.

  • New explainer

    A2A protocol

    A2A is an open protocol for AI agents: one agent can find another, hand it a task and track the result, even when different companies built them.

  • New explainer

    Benchmark contamination

    Benchmark contamination happens when test questions, or close copies of them, end up in a model's training data, so its score can look better than its real skill.

  • New explainer

    LangChain

    LangChain is an open-source framework for building apps and agents on top of large language models, with one standard way to talk to many AI providers.

  • New explainer

    Agentic workflows

    An agentic workflow runs a task as several language model and tool steps, joined by a path that your code fixes in advance.

  • New explainer

    AWS Inferentia

    AWS Inferentia is a family of computer chips that Amazon designed to run trained AI models cheaply and quickly in its cloud.

  • New explainer

    Azure OpenAI

    Azure OpenAI lets companies use OpenAI's models through Microsoft's Azure cloud, with Azure billing, safety filters and data controls.

  • New explainer

    Quantization

    Quantization stores a model's numbers with fewer bits, such as 8-bit integers instead of 32-bit decimals, so it needs less memory.

  • New explainer

    Agent planning

    Agent planning is how an AI agent splits a big task into smaller steps first, then carries out those steps one by one.

  • New explainer

    AWS Trainium

    AWS Trainium is a computer chip Amazon designed for training and running AI models, rented through special Amazon EC2 cloud servers.

  • New explainer

    Multi-agent systems

    A multi-agent system is a group of AI agents, each a language model using tools in its own loop, that work together on one job.

  • New explainer

    Approximate nearest neighbor search

    Approximate nearest neighbor search finds stored vectors that are very close to a query quickly, by checking only a promising part of the collection and accepting a few misses.

  • New explainer

    Inference optimization

    Inference optimization is a set of methods, like caching, quantization and speculative decoding, that let a trained language model answer with less work or less memory.

  • New explainer

    Strands Agents

    Strands Agents is a free AWS toolkit, released as open source under Apache 2.0, for making AI agents with Python or TypeScript, where the model itself plans the steps and picks the tools.

  • New explainer

    Voice agents

    A voice agent is an AI app you talk to out loud; it listens, works out what you want, can use tools, and answers in speech.

  • New explainer

    Agno

    Agno is open-source software for building AI agents and running them as a live service. You can build single agents, teams of agents and step-by-step workflows.

  • New explainer

    Indirect prompt injection

    Indirect prompt injection is when an attacker hides instructions in content an AI app reads later, such as an email or web page, instead of typing them in.

  • New explainer

    Code interpreters

    A code interpreter lets an AI model write a small program, run it in a sealed-off sandbox, and use the result in its answer.

  • New explainer

    AutoGen

    AutoGen is an open-source Microsoft framework for building apps where several AI agents talk to each other, and sometimes to people, to finish a task.

  • New explainer

    Hugging Face

    Hugging Face runs the Hub, a website where people share AI models, datasets and small demo apps so others can find, download and reuse them.

  • New explainer

    ComfyUI

    ComfyUI is a free, open-source app where you build AI image, video and audio generators by wiring boxes called nodes together.

  • New explainer

    MLOps

    MLOps is a way of working that makes building and releasing machine learning models simpler and more automatic.

  • New explainer

    QLoRA

    QLoRA fine-tunes a large language model by storing its frozen weights in 4 bits and training only a small add-on that sits beside them.

  • New explainer

    Fine-tuning

    Fine-tuning takes a model that is already trained and keeps training it on a smaller set of examples for one task.

  • New explainer

    HyDE

    HyDE has a language model write a made-up answer to your question, then searches for real documents that look like that answer.

  • New explainer

    Model routing

    Model routing sends each prompt to the model that suits it, so simple requests go to cheaper models and hard ones to stronger models.

  • New explainer

    Open WebUI

    Open WebUI, a free chat app you host yourself, lets you talk to AI models running on your own computer or in the cloud, from one web page.

  • New explainer

    SGLang

    SGLang is open-source server software that runs large language models on GPUs and answers many requests quickly, partly by reusing work shared between prompts.

  • New explainer

    Mastra

    Mastra is an open-source TypeScript toolkit for making AI agents, step-by-step workflows and the apps that use them.

  • New explainer

    Self-reflection

    Self-reflection is when an AI model checks its own first answer, writes feedback on it, and then tries again using that feedback.

  • New explainer

    llama.cpp

    llama.cpp is an open-source program that runs large language models on your own computer.

  • New explainer

    Hugging Face Transformers

    Transformers is a free, open-source Python library that lets you download a ready-trained AI model by name and run or train it with a short piece of code.

  • New explainer

    Self-hosting LLMs

    Self-hosting means running a language model on a computer you control, so your prompts go to your own machine instead of a company's service.

  • New explainer

    Smolagents

    Smolagents is a small open-source Python library from Hugging Face for building AI agents, including ones that act by writing short pieces of Python code.

  • New explainer

    Amazon SageMaker AI

    SageMaker AI, from AWS, lets you build, train and run machine learning models without managing your own servers.

  • New explainer

    Unsloth

    Unsloth is an open-source tool for training and running AI language models on your own computer, built to fine-tune them faster and with less graphics-card memory.

  • New explainer

    Semantic Kernel

    Semantic Kernel is an open-source Microsoft toolkit that connects AI models to your own code, so a model can ask your functions to do real work.

  • New explainer

    DSPy

    DSPy is a Python framework where you describe an AI job by naming its inputs and outputs, and it tunes the prompts for you.

  • New explainer

    Vercel AI SDK

    The Vercel AI SDK is an open-source TypeScript toolkit, free to use, that lets apps talk to AI models from many companies through the same code.

  • New explainer

    vLLM

    vLLM is open-source software that runs large language models on a server and answers many requests at once, using GPU memory carefully so fewer bytes go to waste.

27 Sept 2026

The catalogue was created, audited, and updated to prioritize current, useful AI terms.

  • Catalogue

    Whole site Created the catalogue, removed the low-value Code repos category, and added the current Gemma 4, Jev and Laya model records.

    The catalogue should help readers understand AI, not inflate its count with repository names.

  • Catalogue

    Whole site Removed superseded model versions and discontinued or redundant products, normalized current product names, and added seven current model and product references.

    The catalogue should lead with what readers encounter now while retaining only historically useful landmark terms.