Glossary
Every AI term, in plain words. 805 terms; search, pick a category, or browse A to Z.
| Term | In one line | Read |
|---|---|---|
| 3D Gaussian splatting | 3D Gaussian splatting rebuilds a scene from ordinary photos as many small, soft, coloured blobs that can be drawn quickly from new viewpoints. | Quick |
| 3D generation models | Three-dimensional generation models create digital shapes, scenes, assets, or representations from text, images, geometry, or other conditioning inputs. | Quick |
| A/B testing LLMs | A/B testing LLMs compares two model, prompt, or workflow variants on real traffic to measure differences in quality and product outcomes. | Quick |
| A2A protocol | A2A is an open protocol for AI agents: one agent can find another, hand it a task and track the result, even when different companies built them. | 4 min |
| Accuracy | Accuracy is the share of a model's predictions that were right, the correct answers divided by all answers. | 4 min |
| Activation function | An activation function is the non-linear step a neuron applies to its weighted sum, so a network of layers can learn curved patterns, not only straight lines. | 4 min |
| Activepieces | Activepieces is a workflow automation platform with visual flows, integrations and AI-oriented components that can be self-hosted. in production workflows. | Quick |
| Adam and AdamW | Adam is an optimizer that adjusts each weight's step size using running averages of recent gradients; AdamW is a version that handles weight decay separately and is widely used to train large models. | Quick |
| Adobe Firefly | Adobe Firefly is Adobe’s family of generative creative tools for images, video, audio and design, integrated across Adobe applications. | Quick |
| Adversarial examples | Adversarial examples are inputs deliberately altered to cause model errors, often through changes that appear small or irrelevant to people. | Quick |
| Agent evaluation | Agent evaluation tests an AI agent on set tasks, checking both its final result and the steps it took, over several tries. | 4 min |
| Agent memory | Agent memory is how an AI agent saves useful information outside the model and loads the right pieces back into its prompt later. | 5 min |
| Agent orchestration | Agent orchestration is how an app with several AI agents decides which agent works, in what order, and who picks what happens next. | 4 min |
| Agent planning | Agent planning is how an AI agent splits a big task into smaller steps first, then carries out those steps one by one. | 4 min |
| Agent skills | Agent skills are reusable packages of instructions, knowledge, or tool procedures that help an AI agent perform a defined class of tasks consistently. | Quick |
| Agent state | Agent state is the running record an AI agent keeps while it works, so a task can pause, resume, or recover from a crash. | 5 min |
| Agentic models | An agent plans, acts, observes the result, adjusts its approach, and repeats until the task is complete. | 3 min |
| Agentic RAG | Agentic RAG gives a model control over retrieval, allowing it to plan searches, choose sources, refine queries, and verify evidence before answering. | Quick |
| Agentic workflows | An agentic workflow runs a task as several language model and tool steps, joined by a path that your code fixes in advance. | 4 min |
| AGENTS.md | AGENTS.md is a convention for placing repository-specific instructions where coding agents can discover guidance on structure, commands, style, and verification. | Quick |
| AGI | AGI is a proposed category for AI at least as capable as a person across most thinking tasks. Definitions differ; a 2023 Google DeepMind framework marked the level matching many definitions as not yet achieved. | 4 min |
| Agno | Agno is open-source software for building AI agents and running them as a live service. You can build single agents, teams of agents and step-by-step workflows. | 4 min |
| AI accelerators | AI accelerators are chips built to speed up the maths inside AI models, which is largely multiply-add sums. | 5 min |
| AI agent with tools | An AI agent with tools lets a model choose defined functions, inspect their results and continue until it reaches a controlled stopping point. | Quick |
| AI agents | An AI agent is a language model that chooses its own next steps and tools while it works through a task. | 4 min |
| AI art | AI art is visual, musical, literary, or other creative work generated or transformed with artificial-intelligence tools, often under human direction. | Quick |
| AI engineering | AI engineering is the discipline of building dependable products around models, covering data, prompts, retrieval, evaluation, infrastructure, safety, and user experience. | Quick |
| AI ethics | AI ethics examines moral questions about designing and using AI, including fairness, autonomy, accountability, privacy, labor, power, and potential harm. | Quick |
| AI governance | AI governance comprises the policies, roles, processes, standards, and oversight used to direct and control AI development and deployment. | Quick |
| AI in education | AI in education supports tutoring, feedback, accessibility, assessment, content creation, and administration, while raising questions about accuracy, privacy, and academic integrity. | Quick |
| AI in finance | AI in finance supports forecasting, fraud detection, trading, risk assessment, customer service, compliance, and operational automation within regulated financial systems. | Quick |
| AI in healthcare | AI in healthcare supports tasks such as diagnosis, imaging, monitoring, administration, research, and treatment planning, generally requiring careful clinical validation and oversight. | Quick |
| AI incidents | AI incidents are events where an AI system causes, contributes to, or nearly causes harm, failure, misuse, or unexpected disruption. | Quick |
| AI model | An AI model is the part of an AI system that has learned from data. It is a structure plus a set of learned numbers that turns an input into a prediction or new content. | 4 min |
| AI safety | AI safety studies and applies methods to prevent, detect, and reduce harms arising from the development, deployment, or misuse of AI systems. | Quick |
| AI search | In consumer applications, AI search combines information retrieval with language models or other AI methods to interpret questions, rank evidence, and synthesize answers. | Quick |
| AI weather forecasting | AI weather forecasting uses learned models to predict atmospheric conditions from historical observations, simulations, and current measurements, complementing numerical forecasting methods. | Quick |
| AI winter | An AI winter is a period when disappointment with artificial intelligence leads to reduced funding, investment, public interest, and research activity. | Quick |
| Aider | Aider is an open-source AI pair-programming tool that edits code in your local git repository from the terminal. | Quick |
| Alexa+ | Alexa+ is Amazon’s generative assistant for conversation, smart-home control, entertainment and supported tasks across Echo devices and connected services. | Quick |
| AlexNet | AlexNet is an influential convolutional neural network that demonstrated the effectiveness of deep learning on large-scale image classification. | Quick |
| Algorithm | An algorithm is a precise, step-by-step procedure for getting a result. In AI, learning algorithms are the procedures that turn data into a trained model. | 4 min |
| Alignment | AI alignment is the effort to make systems reliably pursue intended goals and behave consistently with relevant human values, instructions, and constraints. | Quick |
| Allen Institute for AI | The Allen Institute for AI is a nonprofit research institute that publishes AI research, models, datasets and open software. | Quick |
| AlphaCode | AlphaCode is Google DeepMind research on generating competitive-programming solutions, using large language models and large-scale candidate filtering to solve coding challenges. | Quick |
| AlphaFold | AlphaFold is Google DeepMind’s system for predicting three-dimensional protein structures from amino-acid sequences, accelerating parts of biological research. | Quick |
| AlphaGo | AlphaGo is Google DeepMind’s Go-playing system, combining deep neural networks and tree search to defeat leading professional players. | Quick |
| AlphaZero | AlphaZero is a DeepMind reinforcement-learning system that learned chess, shogi, and Go from self-play without human game examples. | Quick |
| Amazon Bedrock | Amazon Bedrock is an AWS service that lets developers call AI models from several companies through one set of APIs, with extra tools for search and safety. | 5 min |
| Amazon Bedrock AgentCore | A set of AWS services, called Amazon Bedrock AgentCore, for running AI agents in the cloud, with hosting, memory, sign-in, tool access and monitoring. | 5 min |
| Amazon Nova | Amazon Nova is AWS’s family of foundation models for text, multimodal understanding, image generation, video generation, and related application workflows. | Quick |
| Amazon Q Business | Amazon Q Business is an enterprise assistant that searches connected organisational data, answers questions and supports authorised workplace tasks. | Quick |
| Amazon Q Developer | Amazon Q Developer assists with coding, cloud operations and AWS development through IDE, command-line and AWS service integrations. | Quick |
| Amazon SageMaker AI | SageMaker AI, from AWS, lets you build, train and run machine learning models without managing your own servers. | 4 min |
| AMD Instinct | AMD Instinct is AMD’s family of data-centre GPU accelerators for high-performance computing and machine-learning training or inference. in practical systems. | Quick |
| Anomaly detection | Anomaly detection finds the few items or events that do not fit the usual pattern in data, such as a fraudulent card payment or an odd burst of network traffic. | 4 min |
| Anthropic API | The Anthropic API gives developers hosted access to Claude models for text, vision, tool use and agentic application workflows. | Quick |
| AnythingLLM | AnythingLLM is a self-hosted AI workspace that combines model chat, document retrieval, agents and team-oriented knowledge features. under user control. | Quick |
| Apple Foundation Models | Apple Foundation Models are the on-device and server models underlying Apple Intelligence features and exposed to developers through Apple frameworks. | Quick |
| Apple Intelligence | Apple Intelligence is Apple’s system-level collection of generative features for writing, images, notifications, Siri and personal context across supported devices. | Quick |
| Apple M-series chips | Apple M-series chips are system-on-chip processors for Macs and selected Apple devices, combining CPU, GPU, media and neural-processing components. | Quick |
| Apple Neural Engine | Apple Neural Engine is dedicated hardware within Apple chips for accelerating machine-learning operations efficiently on supported devices. in practical systems. | Quick |
| Approximate nearest neighbor search | Approximate nearest neighbor search finds stored vectors that are very close to a query quickly, by checking only a promising part of the collection and accepting a few misses. | 4 min |
| ARC-AGI | ARC-AGI presents demonstration grid pairs to a system, then asks it to construct outputs for new test grids. | 3 min |
| Arize Phoenix | Arize Phoenix is an open-source AI observability and evaluation platform for tracing applications, inspecting retrieval, and analyzing model performance. | Quick |
| Artificial intelligence | An AI system infers from received inputs how to generate predictions, content, recommendations or decisions for explicit or implicit objectives. | 5 min |
| AssemblyAI | AssemblyAI provides hosted speech-recognition APIs with transcription and audio-intelligence features such as speaker labels, chapters and content analysis. | Quick |
| Attention | Attention lets a model build an output by assigning weights to available pieces of information and mixing their values. | 3 min |
| Autoencoders | An autoencoder learns an encoder that maps an input to a code and a decoder that reconstructs the input from that code. | 3 min |
| AutoGen | AutoGen is an open-source Microsoft framework for building apps where several AI agents talk to each other, and sometimes to people, to finish a task. | 4 min |
| AutoGPT | AutoGPT is an experimental agent project that helped popularise autonomous, tool-using task loops built around language models. in production workflows. | Quick |
| AutoML | AutoML automates parts of building a machine learning model, such as choosing the algorithm, the features and the settings. | Quick |
| Autonomous vehicles | Autonomous vehicles use sensors, maps, perception, prediction, planning, and control systems to navigate with reduced or no direct human driving. | Quick |
| Autoregressive models | An autoregressive model factorizes a sequence probability into conditional probabilities. | 3 min |
| AWS Inferentia | AWS Inferentia is a family of computer chips that Amazon designed to run trained AI models cheaply and quickly in its cloud. | 4 min |
| AWS Trainium | AWS Trainium is a computer chip Amazon designed for training and running AI models, rented through special Amazon EC2 cloud servers. | 4 min |
| Axolotl | Axolotl is an open-source tool for fine-tuning language models where you describe the model, data and training method in one configuration file instead of writing code. | Quick |
| Azure Machine Learning | Azure Machine Learning is Microsoft’s managed service for training, tracking, deploying and governing machine-learning models and pipelines. in production. | Quick |
| Azure OpenAI | Azure OpenAI lets companies use OpenAI's models through Microsoft's Azure cloud, with Azure billing, safety filters and data controls. | 5 min |
| Backpropagation | Backpropagation computes each parameter's gradient by applying the chain rule from the last layer back to the first. | 4 min |
| Base models | A base language model is a pretrained model before further instruction or conversational fine-tuning. | 3 min |
| Batch inference | Batch inference sends many AI requests as one job that runs in the background, trading an instant reply for a lower price. | 3 min |
| Batch normalization | Batch normalization rescales a layer's outputs using statistics from the current batch of examples, which helps many networks train faster and more steadily. | Quick |
| Batch size | Batch size is how many training examples are grouped before the weights change once. | 3 min |
| Bayesian inference | Bayesian inference treats an unknown number as uncertain, starts from a belief about it, and uses Bayes' theorem to update that belief as data arrives. | 5 min |
| Beam search | Beam search keeps a fixed number of high-scoring partial sequences, expands them, and retains the best candidates at each step. | 3 min |
| Benchmark contamination | Benchmark contamination happens when test questions, or close copies of them, end up in a model's training data, so its score can look better than its real skill. | 4 min |
| Benchmarks | A benchmark is a fixed set of test tasks with a scoring rule, so different AI models can be compared on the same exam. | 3 min |
| BERT | BERT is Google’s bidirectional transformer encoder, influential for understanding word meaning in context and fine-tuning on language classification tasks. | Quick |
| bfloat16 | Bfloat16 is a 16-bit floating-point format with a wide exponent range but reduced precision, commonly used for efficient neural-network computation. | Quick |
| BGE | BGE is BAAI’s family of open embedding and reranking models for semantic search, retrieval, and multilingual representation learning. | Quick |
| Bias in AI | Bias in AI is systematic skew in data, design, or outputs that can produce inaccurate or unfair results across groups or situations. | Quick |
| Bias–variance tradeoff | Bias is average error across training sets, while variance indicates sensitivity to varying training sets. | 4 min |
| BIG-bench | BIG-bench is a collaborative collection of diverse language-model tasks designed to probe capabilities, limitations, and behaviors beyond standard benchmarks. | Quick |
| bitsandbytes | bitsandbytes is an open-source library for 8-bit and 4-bit quantization, used to load and fine-tune large models with less GPU memory. | Quick |
| BLEU and ROUGE | BLEU and ROUGE score generated text by counting how many words and phrases overlap with reference texts; they are common in translation and summarisation. | Quick |
| BM25 | BM25 is a widely used keyword-ranking algorithm that scores documents using term frequency, term rarity, and document length. | Quick |
| Bolt.new | Bolt.new is StackBlitz’s browser-based AI app builder, combining code generation with a live development environment, package installation and deployment. | Quick |
| Braintrust | Braintrust is an AI evaluation platform for running experiments, managing datasets and prompts, tracing applications, and monitoring production quality. | Quick |
| Browser agents | Browser agents are AI systems that navigate websites, extract information, fill forms, and complete web tasks through browser controls or automation APIs. | Quick |
| Browser automation agent | A browser automation agent observes web pages and performs bounded actions such as clicking, typing and collecting results. | Quick |
| Browser Use | Browser Use is an open-source framework for connecting AI agents to web browsers so they can observe pages and perform actions. | Quick |
| Byte-pair encoding | Byte-pair encoding builds a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair. | 3 min |
| C4 | C4, or Colossal Clean Crawled Corpus, is a cleaned English web-text dataset derived from Common Crawl for language-model training. | Quick |
| Canva AI | Canva AI brings generative writing, image, design and editing features into Canva’s visual communication and template-based creation platform. | Quick |
| Catastrophic forgetting | Catastrophic forgetting occurs when learning new information substantially degrades a neural network's performance on knowledge or tasks learned earlier. | Quick |
| CatBoost | CatBoost is Yandex’s gradient-boosting library with built-in handling for categorical features and tools for tabular prediction problems. in practice. | Quick |
| Cerebras | Cerebras provides AI computing systems and hosted inference services built around its wafer-scale processors. in production. in production. | Quick |
| Chain of thought | Chain-of-thought prompting asks for or demonstrates intermediate reasoning text before a final answer, but that text is not a guaranteed view of hidden model internals. | 4 min |
| Character.AI | Character.AI is a conversational platform centred on user-created AI characters, role-play and interactive storytelling rather than one general-purpose assistant. | Quick |
| Chat templates | A chat template converts a structured list of messages into a formatted sequence. | 3 min |
| Chat with PDF | A Chat with PDF application extracts and indexes a PDF so users can ask questions and receive answers grounded in its pages. | Quick |
| Chat with your database | A Chat with your database application translates natural-language requests into controlled queries, then explains results returned from structured data. | Quick |
| Chatbots | Chatbots are software interfaces that conduct text or voice conversations using rules, retrieval, generative models, or combinations of these methods. | Quick |
| ChatGPT | ChatGPT is OpenAI’s conversational product for asking questions, creating things and completing work with models and tools. | 5 min |
| ChatGPT Work | ChatGPT Work is OpenAI’s agent for longer projects, able to work across connected apps and files and create finished documents, spreadsheets, presentations, reports, and sites. | Quick |
| Checkpoints | Checkpoints are saved snapshots of model parameters and training state, enabling recovery, comparison, sharing, or continued training from an earlier point. | Quick |
| Chinchilla scaling | Chinchilla scaling allocates a fixed training-compute budget by growing model parameters and training tokens in roughly equal proportions. | 3 min |
| Chroma | Chroma is a developer-oriented vector database for storing embeddings and adding semantic retrieval to AI applications. in production workflows. | Quick |
| Chunking | Chunking splits a long document into smaller pieces that can be embedded, searched, or fitted into a model prompt. | 3 min |
| Classification | Classification predicts which discrete category or categories an input belongs to, such as identifying an email as spam or legitimate. | Quick |
| Classifier-free guidance | Classifier-free guidance is a setting in diffusion image generators that controls how strongly the result follows the text prompt. | Quick |
| Claude | Claude is Anthropic’s family of conversational AI models, designed for writing, analysis, coding, multimodal understanding, and tool-assisted workflows. | Quick |
| Claude Agent SDK | Anthropic's Claude Agent SDK is a Python and TypeScript library for building agents that run on the same loop, tools and context handling as Claude Code. | 5 min |
| Claude app | Claude is Anthropic’s AI product for conversation, analysis and completing tasks with files, connected tools and editable outputs. | 5 min |
| Claude Code | Claude Code is Anthropic's coding assistant that reads a project, edits files and runs commands for you, asking before risky steps. | 4 min |
| Claude Fable 5.1 | Claude Fable 5.1 is Anthropic’s generally available model for ambitious coding, knowledge work, research, and long-running agents, with configurable effort and broad platform availability. | Quick |
| Claude Haiku | Claude Haiku refers to Anthropic’s fastest, smallest model tier, designed for responsive and cost-sensitive workloads such as classification and support automation. | Quick |
| Claude Opus 5.5 | Claude Opus 5.5 is Anthropic’s 2026 Opus model, positioned for demanding coding, research, and professional work with stronger efficiency than the previous Opus generation. | Quick |
| Cline | Cline is an open-source AI coding agent that works inside the VS Code editor and can edit files and run commands with your approval. | Quick |
| CLIP | CLIP is an OpenAI model that learns shared representations of images and text, enabling zero-shot image classification and cross-modal retrieval. | Quick |
| Closed models | A fully closed model keeps its weights and code proprietary for internal use. | 3 min |
| Cloudflare Workers AI | Cloudflare Workers AI lets developers run supported AI models from Cloudflare’s serverless platform close to applications and users. | Quick |
| Clustering | Clustering sorts data that has no labels into groups of items that resemble each other. | 4 min |
| COCO | COCO is a large image dataset annotated for object detection, segmentation, captioning, and keypoint estimation in everyday scenes. | Quick |
| Code generation | Code generation uses AI to produce software from natural-language instructions, partial code, examples, specifications, or surrounding repository context. | Quick |
| Code interpreters | A code interpreter lets an AI model write a small program, run it in a sealed-off sandbox, and use the result in its answer. | 4 min |
| Code models | Code Llama is a family of language models for code with infilling and instruction-following capabilities. | 3 min |
| Code review bot | A code review bot examines proposed changes for defects, style issues or policy violations and leaves review comments for developers. | Quick |
| Codestral | Codestral is Mistral AI’s code-focused model family, supporting code generation, completion, fill-in-the-middle editing, and programming assistance across common languages. | Quick |
| Coding assistant | A coding assistant uses repository context to explain code, propose changes, generate tests and help debug software. It remains subject to human review. | Quick |
| Cohere Embed | Cohere Embed is Cohere’s commercial embedding model family for semantic search, classification, clustering, and retrieval across multiple languages and content types. | Quick |
| Cohere Platform | Cohere’s platform provides APIs for enterprise language models, embeddings, reranking and retrieval-oriented generative applications. in production. in production. | Quick |
| Cohere Rerank | Cohere Rerank is a model service that reorders retrieved documents by their relevance to a query, improving search and RAG results. | Quick |
| ColBERT | ColBERT is a retrieval method that compares a query and a document word by word using token-level embeddings, sitting between fast vector search and slower reranking. | Quick |
| Collaborative filtering | Collaborative filtering recommends items by learning from what many people rated, clicked or watched, then predicting the gaps in your own history. | 4 min |
| ColPali | ColPali is a document retrieval model that searches page images directly, using a vision-language model with late-interaction matching. | Quick |
| ComfyUI | ComfyUI is a free, open-source app where you build AI image, video and audio generators by wiring boxes called nodes together. | 4 min |
| Command A | Command A is Cohere’s enterprise language model for retrieval-augmented generation, multilingual work, tool use, and business-focused agentic applications. | Quick |
| Common Crawl | Common Crawl is a nonprofit repository of web-page snapshots, widely used as raw training data for search, language, and web research. | Quick |
| Computer use | Computer use lets an AI interpret a graphical interface and take actions such as clicking, typing, scrolling, or reading screen content. | Quick |
| Computer vision | Computer vision enables machines to extract information from images and video, including objects, text, motion, depth, scenes, and spatial relationships. | Quick |
| Computer-use models | A computer-use model reads a screen and emits mouse or keyboard actions that an application can execute. | 4 min |
| Confusion matrix | A confusion matrix is a table that counts a classifier's predictions by what it predicted and what was really true, so you see which mistakes it makes. | 5 min |
| Constitutional AI | Constitutional AI trains and evaluates model behavior against an explicit written set of principles. | 4 min |
| Content moderation pipeline | A content moderation pipeline classifies submitted material against written policies, applies thresholds and sends ambiguous or high-risk cases to reviewers. | Quick |
| Content provenance | Content provenance records information about media's origin and editing history, helping people and systems assess how a digital artifact was created. | Quick |
| Context compaction | Context compaction swaps the older part of a long AI conversation for a short summary, so the chat can keep going inside the model's memory limit. | 4 min |
| Context engineering | Context engineering is choosing, trimming and updating the tokens a model sees at each step, so an AI agent works from a small, useful context. | 4 min |
| Context window | A context window is all the text a language model can reference while generating a response, including the response itself. | 3 min |
| Contextual retrieval | Contextual retrieval has a model write a short note that places each chunk in its document, then indexes the note with the chunk so search can find it. | 5 min |
| Continual learning | Continual learning trains a model on a stream of changing data while measuring whether later learning damages or helps earlier tasks. | 4 min |
| Continuous batching | Continuous batching dynamically adds and removes inference requests as sequences finish, improving accelerator utilization compared with waiting for an entire fixed batch. | Quick |
| ControlNet | ControlNet is an open neural-network approach and tooling ecosystem for guiding diffusion image generation with edges, poses, depth and other conditions. | Quick |
| Convolutional neural networks | A convolutional neural network uses sparse convolutions that reuse the same weights at multiple locations in an ordered grid. | 3 min |
| Copilot+ PCs | Copilot+ PCs are Windows computers that meet Microsoft’s hardware requirements for on-device AI features, including a sufficiently capable neural processor. | Quick |
| Copyright and AI | Copyright and AI concerns how protected works may be used in training and how authorship, ownership, licensing, and infringement apply to generated outputs. | Quick |
| Coqui TTS | Coqui TTS is a deep-learning toolkit for training and running text-to-speech, voice-cloning and vocoder models across many languages. | Quick |
| Cosine similarity | Cosine similarity scores how closely two vectors point the same way, ignoring how long they are. The score runs from −1 to 1. | 4 min |
| Cost optimization | AI cost optimization reduces spending through model selection, caching, batching, shorter contexts, efficient retrieval, usage controls, and infrastructure tuning. | Quick |
| CPUs | A CPU is the chip that runs a computer's software, carrying out instructions with just a few cores backed by lots of cache memory. | 5 min |
| Crawl4AI | Crawl4AI is a web crawler designed to extract configurable, structured and AI-ready content from sites in self-hosted Python workflows. | Quick |
| CrewAI | CrewAI is an open-source toolkit, written in Python, for building teams of AI agents, each with a role, that work through a list of tasks together. | 4 min |
| Cross-validation | Cross-validation repeatedly trains and evaluates a method on different data partitions, providing a more reliable performance estimate when data is limited. | Quick |
| Curriculum learning | Curriculum learning presents training examples in a deliberate order, often from easier to harder, to improve learning efficiency or final performance. | Quick |
| Cursor | Cursor is a coding tool with an AI agent that can read your project, edit files and run commands for you. | 4 min |
| Custom GPTs | Custom GPTs are user-configured ChatGPT experiences with tailored instructions, knowledge files, capabilities and optional actions for a particular purpose. | Quick |
| Customer support bot | A customer support bot answers common questions, retrieves policy information and hands uncertain or sensitive cases to people for review. | Quick |
| DALL·E | DALL·E is OpenAI’s family of generative image models, creating original images from natural-language descriptions and supporting image editing in some versions. | Quick |
| Data analysis agent | A data analysis agent inspects datasets, writes or runs analysis code and explains results while preserving traceable calculations. | Quick |
| Data and concept drift | Data drift and concept drift are changes after launch in the data a model sees, or in what the right answer is, that make its predictions worse over time. | Quick |
| Data augmentation | Data augmentation makes extra training examples by changing existing ones in small ways that keep the answer the same, such as flipping or cropping a photo. | 4 min |
| Data curation | Data curation selects, cleans, balances, documents, and organizes examples so a dataset better supports its intended training or evaluation purpose. | Quick |
| Data labeling | Data labeling attaches the answer a supervised model should learn to raw examples such as images, text, audio or table rows. | 3 min |
| Data leakage | Data leakage is when a model learns from information it will not have when making real predictions, so its test score looks better than it really is. | 4 min |
| Data parallelism | Data parallelism runs replicas of one model on different slices of a batch, then combines their gradients before a synchronized update. | 4 min |
| Data poisoning | Data poisoning deliberately corrupts training data to damage a model, introduce hidden behavior, or manipulate its future predictions. | Quick |
| Data privacy | Data privacy is about what an AI company may do with your chats, such as training models, and the settings that let you limit it. | 5 min |
| Databricks Mosaic AI | Databricks Mosaic AI brings model development, evaluation, retrieval, agents and governance into the Databricks data and lakehouse environment. | Quick |
| Decision trees | Decision trees make predictions through a sequence of feature-based splits, forming interpretable branches that end in class or value estimates. | Quick |
| Decoder-only models | A decoder-only model predicts each next token from the tokens that came before it. | 3 min |
| Deep learning | Deep learning is machine learning with neural networks that stack many layers, so each layer can build a more abstract picture of the data than the one before. | 4 min |
| Deep research agents | Deep research agents plan multi-step investigations, search and read many sources, synthesize evidence, and produce cited reports with limited supervision. | Quick |
| DeepEval | DeepEval is an open-source framework for testing language-model applications with configurable metrics, datasets, model-based judges, regression checks, and continuous-integration support. | Quick |
| Deepfakes | Deepfakes are synthetic or manipulated media that realistically depict people saying or doing things that did not occur. | Quick |
| Deepgram | Deepgram provides speech-to-text, text-to-speech and voice-agent APIs for developers building real-time audio applications. for practical use. for practical use. | Quick |
| DeepSeek API | The DeepSeek API provides hosted access to DeepSeek language and reasoning models through developer-compatible text interfaces. in production. | Quick |
| DeepSeek-R1 | DeepSeek-R1 is an open-weight reasoning model family from DeepSeek, trained to solve mathematics, coding, and other multi-step problems. | Quick |
| DeepSpeed | DeepSpeed is Microsoft's open-source library for training and running very large models across many GPUs, known for its ZeRO memory-saving technique. | Quick |
| Descript | Descript is an audio and video editor built around editable transcripts, with AI tools for cleanup, clips, captions and synthetic voice correction. | Quick |
| Devin | Devin is Cognition’s software-engineering agent for taking on scoped coding tasks, using development tools and returning changes for review. | Quick |
| Devstral | Devstral is Mistral AI’s model family specialized for software-engineering agents that inspect repositories, edit files, and solve coding tasks. | Quick |
| Differential privacy | Differential privacy adds carefully calibrated randomness so aggregate analysis reveals useful patterns while limiting what can be inferred about any individual record. | Quick |
| Diffusers | Diffusers is Hugging Face’s library for running and training diffusion models for images, video, audio and related generative tasks. | Quick |
| Diffusion models | A diffusion model learns to reverse a process that gradually adds noise to data. | 3 min |
| Dify | Dify is an open platform for building, testing and operating AI applications, agents and retrieval workflows through visual and API tools. | Quick |
| Dimensionality reduction | Dimensionality reduction represents data using fewer variables while preserving useful structure, aiding visualization, compression, denoising, or downstream modeling. | Quick |
| DINOv2 | DINOv2 is Meta’s self-supervised vision model, trained to produce general-purpose image features without relying on manually labeled datasets. | Quick |
| Discord bot | A Discord bot listens for commands or events and provides focused AI features within authorised servers and channels. | Quick |
| Distillation | Knowledge distillation trains a smaller student model to imitate a larger teacher's outputs or internal representations, often reducing deployment cost. | Quick |
| Docling | Docling parses formats such as PDF and Word into structured representations that preserve text, tables, layout and document hierarchy. | Quick |
| Document parsing | Document parsing extracts structure and content from files such as PDFs, forms, and slides so downstream systems can search or analyze them. | Quick |
| Document Q&A | A document Q&A application retrieves relevant sections from one or more files and answers questions using that supplied context. | Quick |
| Dolma | Dolma is Ai2’s open corpus of web text, books, code, papers, and other sources used to train OLMo models. | Quick |
| Dot product | The dot product multiplies two equal-length lists of numbers position by position and adds the results into one number set by the angle between them and their lengths. | 4 min |
| DPO | DPO teaches a language model to prefer the better of two answers by learning from the comparison itself, with no separate reward model trained first. | 4 min |
| Dropout | During training, dropout randomly replaces some input tensor elements with zero. | 3 min |
| Drug discovery | AI-assisted drug discovery uses computational models to identify targets, predict molecular properties, design candidates, and prioritize experiments during medicine development. | Quick |
| DSPy | DSPy is a Python framework where you describe an AI job by naming its inputs and outputs, and it tunes the prompts for you. | 4 min |
| DVC | DVC adds versioning and reproducible pipelines for datasets, models and machine-learning experiments alongside Git-managed code. in practice. in practice. | Quick |
| E5 | E5 is a family of text-embedding models trained for retrieval and semantic similarity using a unified text-to-vector approach. | Quick |
| Elasticsearch | Elasticsearch is a distributed search and analytics engine that supports keyword, vector and hybrid retrieval over indexed data. | Quick |
| EleutherAI | EleutherAI is a nonprofit open research collective known for language models, training datasets and evaluation tooling released for public use. | Quick |
| ElevenLabs | ElevenLabs is an AI audio platform for speech synthesis, voice cloning, dubbing, transcription and conversational voice applications in practical workflows. | Quick |
| Email assistant | An email assistant drafts, summarises, classifies or routes messages using the user’s instructions and authorised mailbox context. It remains subject to human review. | Quick |
| Embedding layer | An embedding layer turns integer indexes into dense vectors of fixed size. | 3 min |
| Embedding models | An embedding model maps an input to a numeric vector whose distance from another vector can measure relatedness. | 3 min |
| Embeddings | An embedding maps a discrete item such as a token ID to a dense vector of numbers. | 3 min |
| Emergent abilities | Emergent abilities are task capabilities that appear absent in smaller models but present at a larger scale. | 3 min |
| Encoder-decoder | An encoder-decoder model combines an encoder with an autoregressive decoder for sequence generation. | 3 min |
| Encoder-only models | An encoder-only model uses bidirectional self-attention to represent a supplied sequence. | 3 min |
| Ensemble learning | Ensemble learning combines the predictions of several models, usually by a vote or an average, so that their separate mistakes partly cancel out. | 4 min |
| Epoch | An epoch is one complete pass through every example used to train a model. | 4 min |
| EU AI Act | The EU AI Act is a European Union legal framework that regulates AI systems according to risk, with differing obligations for providers and deployers. | Quick |
| Evals | An eval checks an AI system's work. You feed it an input and use a set of rules to score what comes back. | 5 min |
| Evaluation harness | An evaluation harness runs models or prompts against a fixed test set, records outputs and calculates repeatable quality, safety or performance measures. | Quick |
| Existential risk | Existential risk from AI refers to scenarios where AI contributes to human extinction or permanently and drastically curtails humanity's future potential. | Quick |
| Expert systems | An expert system stores a specialist's know-how as if-then rules and uses a separate inference engine to apply those rules to a new case and explain its advice. | 5 min |
| F1 score | The F1 score rolls a classifier's precision and recall into one number from 0 to 1, and it stays low if either of them is low. | 4 min |
| FAISS | FAISS is Meta’s open library for efficient similarity search and clustering over dense vectors, including very large collections. | Quick |
| faster-whisper | faster-whisper is an open reimplementation of Whisper inference using CTranslate2 for faster, more memory-efficient speech transcription in practical workflows. | Quick |
| FastMCP | FastMCP is a Python framework for building Model Context Protocol servers and clients, so an AI app can reach your own tools and data through one standard interface. | Quick |
| Feature engineering | Feature engineering turns raw data, like prices and colour names, into the lists of numbers a model can learn from. | 4 min |
| Feature importance | Feature importance gives each input column a score for how much a trained model relies on it, so you can see what drives its predictions. | 4 min |
| Feature scaling | Feature scaling puts numeric variables on comparable ranges, limiting disproportionate influence from large-magnitude features in methods sensitive to distances or gradient sizes. | Quick |
| Features | Features are the input facts a machine learning model reads about each example, such as a car's mileage or colour, turned into a list of numbers. | 5 min |
| Federated learning | Federated learning trains a shared model across distributed devices or organizations while keeping raw training data at its original location. | Quick |
| Few-shot learning | Few-shot learning is the ability to perform a task from only a small number of labelled examples. | 3 min |
| Few-shot prompting | Few-shot prompting places a small set of example inputs and outputs in the prompt so the model can imitate the task on a new input. | 3 min |
| Fine-tuning | Fine-tuning takes a model that is already trained and keeps training it on a smaller set of examples for one task. | 4 min |
| Fine-tuning pipeline | A fine-tuning pipeline prepares examples, runs an adaptation job, records configuration and evaluates the resulting model against a held-out set. | Quick |
| FineWeb | FineWeb is Hugging Face’s large, open, filtered web-text dataset created from Common Crawl for training and researching language models. | Quick |
| Firecrawl | Firecrawl crawls websites and converts pages into cleaner Markdown or structured data for search, retrieval and agent applications. | Quick |
| Fireworks AI | Fireworks AI provides managed inference and model customisation for generative applications, with APIs optimised for speed and production operation. | Quick |
| FlashAttention | FlashAttention computes exact attention with tiling that reduces reads and writes between GPU memory levels. | 3 min |
| FLOPs | FLOPs count the small arithmetic steps, like one multiply or one add on decimal numbers, that a computer does to train an AI model. | 5 min |
| Flow matching | Flow matching trains a generative model to learn a smooth path that turns random noise into data, an approach closely related to diffusion. | Quick |
| Flowise | Flowise is a visual builder for creating model, retrieval and agent workflows from connected nodes and integrations. in production workflows. | Quick |
| FLUX | FLUX is Black Forest Labs’ family of generative image models, offered in open and hosted variants for text-to-image creation and editing. | Quick |
| Foundation models | A foundation model is trained on broad data and can be adapted for many downstream tasks. | 3 min |
| FP8 | FP8 refers to eight-bit floating-point formats that reduce memory and increase accelerator throughput, while requiring careful scaling to manage limited precision and range. | Quick |
| Fraud detection | AI fraud detection identifies suspicious transactions or behavior by learning patterns associated with legitimate activity, known abuse, and unusual events. | Quick |
| Frequency and presence penalties | Frequency and presence penalties are generation settings that make a model less likely to repeat tokens it has already used. | Quick |
| Frontier models | Frontier models are highly capable general-purpose AI models at or beyond the capabilities of the most advanced current models. | 3 min |
| Full Self-Driving (Supervised) | Full Self-Driving (Supervised) is Tesla’s driver-assistance software for navigation and vehicle control; the human driver must remain attentive and responsible. | Quick |
| Function calling | Function calling lets a model request a named function by returning structured arguments that an application, or the provider, then runs. | 4 min |
| Gamma | Gamma generates and edits presentations, documents and web-style pages from prompts using structured layouts rather than traditional slide-by-slide authoring. | Quick |
| GANs | A GAN trains a generator to make samples and a discriminator to distinguish generated samples from training data. | 3 min |
| Gemini | Gemini is Google’s family of multimodal AI models, spanning cloud and on-device systems for text, images, audio, video, code, and reasoning. | Quick |
| Gemini 3.5 | Gemini 3.5 is Google’s 2026 multimodal model generation, built for stronger reasoning, coding, agentic workflows, and understanding across text, images, audio, video, and tools. | Quick |
| Gemini API | The Gemini API gives developers hosted access to Google’s Gemini models for text, image, audio, video and tool-using applications. | Quick |
| Gemini app | Gemini is Google’s AI app: you ask questions, upload files for answers and summaries, and can have it research a topic across many sources. | 5 min |
| Gemini CLI | Gemini CLI is Google’s open-source terminal agent for coding, research and automation with Gemini models and tool integrations. | Quick |
| Gemini Enterprise | Gemini Enterprise is Google Cloud’s workplace AI platform for searching company information, using enterprise data and applications, and building governed agents for business workflows. | Quick |
| Gemini for Google Workspace | Gemini for Google Workspace adds writing, summarisation, analysis and meeting assistance across Gmail, Docs, Sheets, Slides, Drive and Meet. | Quick |
| Gemini Nano | Gemini Nano is Google’s compact model family for on-device AI, enabling selected generative features without sending every request to the cloud. | Quick |
| Gemini Robotics | Gemini Robotics is Google DeepMind’s model family for connecting multimodal reasoning with physical robot actions, spatial understanding, and instruction following. | Quick |
| Gemma | Gemma is Google’s family of lightweight open models, derived from Gemini research and intended for developers to run, fine-tune, and deploy themselves. | Quick |
| Gemma 4 | Gemma 4 is Google’s open-weight, natively multimodal model family, designed for advanced reasoning and agentic workflows across edge devices and personal computers. | Quick |
| Generative AI | Generative AI learns patterns from existing data, then uses those patterns to create new text, images, audio, video, code or other content. | 5 min |
| Generative engine optimization | Generative engine optimization adapts content so AI-powered answer systems can discover, understand, cite, or accurately represent it to users. | Quick |
| Genetic algorithms | A genetic algorithm searches for good answers the way breeding does. It keeps a population of candidates, mixes the fitter ones and repeats over many generations. | 4 min |
| Genkit | Genkit is Google’s open-source framework for building and observing AI applications with models, tools, retrieval, flows, and evaluation. | Quick |
| GGUF | GGUF is a binary format for storing quantized models and metadata, widely used by llama.cpp and related local-inference tools. | Quick |
| GitHub Copilot | GitHub Copilot is an AI helper for programmers that offers the next lines of code while you write, explains code when you ask, and can take on jobs you give it. | 4 min |
| Glean | Glean is an enterprise search and AI platform for finding organisational knowledge and building assistants or agents over connected company systems. | Quick |
| GLM-4.5 | GLM-4.5 is a Zhipu AI open-weight model generation designed for reasoning, coding, tool use, and agentic application workflows. | Quick |
| GloVe | GloVe is a word-embedding method that learns vector representations from global word co-occurrence statistics across a text corpus. | Quick |
| Google ADK | Google ADK is an open-source toolkit from Google for writing AI agents in code, giving them tools, a record of each chat and helper agents. | 5 min |
| Google AI Overviews | Google AI Overviews are generated summaries shown for some searches, combining information from web results with links for further reading. | Quick |
| Google AI Studio | Google AI Studio is a browser-based workspace for prototyping prompts, testing Gemini capabilities and obtaining API integration code. | Quick |
| Google Colab | Google Colab is a hosted Jupyter notebook environment with shareable documents and optional access to accelerated computing resources. | Quick |
| Google Flow | Google Flow is an AI filmmaking tool that combines Google’s video, image and language models for creating shots, scenes and cinematic sequences. | Quick |
| Google TPU | Google Tensor Processing Units are specialised accelerators designed for machine-learning workloads across Google’s cloud and internal infrastructure. in practical systems. | Quick |
| GPQA | GPQA is a 448-question multiple-choice benchmark written by experts in biology, physics and chemistry. | 3 min |
| GPT | GPT is OpenAI’s family of generative transformer models, designed to understand prompts and produce text, code, structured data, and other outputs. | Quick |
| GPT Image | GPT Image is OpenAI’s image generation family for creating and editing visuals from conversational instructions, including text rendering and iterative refinements. | Quick |
| GPT-2 | GPT-2 is an early OpenAI transformer language model that helped demonstrate coherent long-form text generation from a simple textual prompt. | Quick |
| GPT-3 | GPT-3 is OpenAI’s influential 175-billion-parameter language model, which demonstrated that large-scale pretraining can support many tasks through prompting alone. | Quick |
| GPT-4 | GPT-4 is an OpenAI large language model generation known for stronger reasoning and instruction following than the earlier GPT-3.5 family. | Quick |
| GPT-5 | GPT-5 is an OpenAI model generation built for general-purpose reasoning, coding, writing, analysis, and tool-assisted work across consumer and developer products. | Quick |
| GPT-6 Astra | GPT-6 Astra is OpenAI’s 2026 frontier model for advanced reasoning, coding, computer use, scientific work, and long-running agentic tasks across ChatGPT and the API. | Quick |
| gpt-oss | gpt-oss is OpenAI’s family of open-weight language models, intended for developers who need locally deployable reasoning models they can inspect and adapt. | Quick |
| GPUs | A GPU is a chip built to run many similar calculations at the same time, which suits the matrix maths inside neural networks. | 5 min |
| Gradient boosting | Gradient boosting builds one strong model from many small decision trees, added one at a time, each trained to fix the errors the trees before it still make. | 4 min |
| Gradient checkpointing | Gradient checkpointing stores fewer activations by recomputing them when needed, exchanging extra computation for lower memory use. | 4 min |
| Gradient descent | Gradient descent repeatedly measures how loss changes with each parameter and moves the parameters a small step toward lower loss. | 3 min |
| Gradio | Gradio is a Python library for quickly building web interfaces and demos around machine-learning models or functions. in practice. | Quick |
| Grammarly | Grammarly provides writing assistance for grammar, clarity, tone and rewriting across browser, desktop and workplace applications. in daily work. | Quick |
| Granite | Granite is IBM’s open model family for enterprise language, code, retrieval, safety, document understanding, and time-series applications across business domains. | Quick |
| Graph neural networks | A graph neural network updates each node's hidden state based on messages. | 3 min |
| GraphRAG | GraphRAG turns your documents into a map of people, places and links, groups and summarises it, then answers questions from that map. | 4 min |
| Greedy decoding | Greedy decoding generates a sequence by choosing the most likely next token at every step. | 3 min |
| Grok | Grok is xAI’s family of conversational models, built for general knowledge, reasoning, coding, multimodal understanding, and tool use. | Quick |
| Grok 4.7 | Grok 4.7 is xAI’s 2026 model release for reasoning, coding, information synthesis, and tool-using tasks across the Grok product and developer platform. | Quick |
| Grok app | Grok is the assistant from SpaceXAI (formerly xAI) for conversation, web and X search, file analysis, voice and media creation. | 5 min |
| Groq | Groq provides hosted model inference on its Language Processing Unit hardware, focusing on low-latency execution for supported workloads. | Quick |
| Ground truth | Ground truth is the best available account of what really happened, the reality that a model's predictions are checked against. | 5 min |
| Grounding | Grounding gives a model relevant source material and asks it to base its response on that evidence. | 3 min |
| Grouped-query attention | Grouped-query attention lets each group of query heads share one key head and one value head. | 3 min |
| GRPO | Group relative policy optimization trains a policy by comparing rewards among multiple responses to the same prompt, avoiding a separate learned value model. | Quick |
| GRU | A GRU is a recurrent unit with reset and update gates. | 3 min |
| GSM8K | GSM8K is a dataset of grade-school mathematics word problems used to evaluate multi-step arithmetic reasoning in language models. | Quick |
| Guardrails | Guardrails are checks placed around an AI model that inspect what goes in and what comes out, and can block, change or check inputs and replies that are unsafe, off-topic or break the app's rules. | 5 min |
| Guardrails AI | Guardrails AI is an open-source framework for checking and correcting language-model inputs and outputs against rules you define. | Quick |
| Guidance | Guidance is a language for controlling model generation with templates, constraints, tool calls, and programmatic branching around generated text. | Quick |
| Hailuo | Hailuo is MiniMax’s generative video product and model family for creating, extending, and transforming clips from text and image prompts. | Quick |
| Hallucination | A hallucination is generated content that sounds plausible but is false, unsupported by evidence, or inconsistent with the given input. | 3 min |
| Harvey | Harvey is a legal AI platform for research, drafting, analysis and professional workflows within law firms and legal departments. | Quick |
| Haystack | An open-source Python toolkit from deepset that builds AI apps (search, RAG, agents) by linking small parts into a pipeline. | 4 min |
| HBM memory | High-bandwidth memory stacks memory near processors to deliver very high data transfer rates, helping keep AI accelerators supplied with model data. | Quick |
| Helicone | Helicone is an observability gateway for language-model applications, providing request logging, cost tracking, caching, experiments, and production monitoring. | Quick |
| HellaSwag | HellaSwag tests commonsense reasoning by asking a model to select the most plausible continuation of a described everyday situation. | Quick |
| HeyGen | HeyGen creates avatar-led and translated videos from scripts, with synthetic presenters, voice tools and lip-synchronised dubbing. for creative production. | Quick |
| Hidden Markov models | A hidden Markov model infers a sequence of states you cannot see from a sequence of observations you can, using the odds of each state following another and of each state producing each observation. | 4 min |
| HNSW | HNSW finds the stored vectors closest to a query quickly by searching a stack of linked graphs, from a sparse top layer down to a dense bottom one. | 4 min |
| Hugging Face | Hugging Face runs the Hub, a website where people share AI models, datasets and small demo apps so others can find, download and reuse them. | 4 min |
| Hugging Face Accelerate | Hugging Face Accelerate simplifies running PyTorch training and inference across CPUs, GPUs, mixed precision and distributed environments. in practice. | Quick |
| Hugging Face Datasets | Hugging Face Datasets provides consistent tools for finding, loading, processing and streaming datasets used in machine-learning workflows. in practice. | Quick |
| Hugging Face Transformers | Transformers is a free, open-source Python library that lets you download a ready-trained AI model by name and run or train it with a short piece of code. | 4 min |
| Human evaluation | Human evaluation asks people to assess model outputs for qualities that automated metrics may miss, such as usefulness, clarity, correctness, or preference. | Quick |
| Human in the loop | With this setup, a person checks an AI system's work at chosen points and can approve it, change it, reject it or step in. | 5 min |
| HumanEval | HumanEval measures functional correctness for programs synthesized from docstrings. | 3 min |
| Humanity's Last Exam | Humanity's Last Exam is a broad benchmark of difficult expert-level questions intended to probe advanced academic reasoning across many disciplines. | Quick |
| HunyuanVideo | HunyuanVideo is Tencent’s generative video model family for text-to-video and image-conditioned video creation, with open model releases available. | Quick |
| Hybrid search | Hybrid search runs a keyword ranker and a vector ranker on one query, then merges the two lists into a single order. | 4 min |
| HyDE | HyDE has a language model write a made-up answer to your question, then searches for real documents that look like that answer. | 4 min |
| Hyperparameters | Hyperparameters are the settings people choose to control training, such as the learning rate, rather than values the model learns itself. | 5 min |
| Ideogram | Ideogram is a generative image product known for prompt-based visual creation and comparatively strong rendering of text inside images. | Quick |
| Image captioning app | An image captioning application analyses visual content and produces descriptive text for search, accessibility or downstream processing. It remains subject to human review. | Quick |
| Image classification | Image classification assigns one or more predefined category labels to an entire image based on its overall visual content. | Quick |
| Image editing models | An image editing model changes a picture you already have. | 4 min |
| Image generator app | An image generator application turns prompts and optional references into generated visuals with controls for style, size and iteration. | Quick |
| Image segmentation | Image segmentation assigns labels to individual pixels or regions, separating objects, surfaces, or semantic categories within an image. | Quick |
| Imagen | Imagen is Google’s family of text-to-image diffusion models, built to generate and edit high-quality images from natural-language instructions. | Quick |
| ImageNet | ImageNet is a large labeled image dataset and benchmark that helped drive modern deep-learning progress in visual recognition. | Quick |
| Imbalanced data | Imbalanced data is a dataset where one class is far rarer than another, such as a few fraud cases among many ordinary card payments. | 4 min |
| Imitation learning | Imitation learning trains an agent to act by copying examples of an expert's behaviour instead of learning only from rewards. | Quick |
| In-context learning | In-context learning is a model's ability to infer a task or pattern from instructions and examples in its prompt without updating parameters. | Quick |
| Indirect prompt injection | Indirect prompt injection is when an attacker hides instructions in content an AI app reads later, such as an email or web page, instead of typing them in. | 4 min |
| Inference | Inference is using a trained AI model on new input to get an answer, such as a label, a number or a written reply, without changing what the model learned. | 5 min |
| Inference optimization | Inference optimization is a set of methods, like caching, quantization and speculative decoding, that let a trained language model answer with less work or less memory. | 3 min |
| Inspect AI | Inspect AI is the UK AI Security Institute’s open-source framework for building reproducible model evaluations with tasks, tools, solvers, and scoring. | Quick |
| Instruct models | An instruct model is a pretrained base version further fine-tuned on instructions and conversational data. | 3 min |
| Instruction following | Instruction following is how well a model does what a written request asks, meeting each rule it sets, such as rules on content, style, format or numbers. | 3 min |
| Instruction tuning | Instruction tuning fine-tunes a pretrained model on many tasks written as instructions paired with desired responses. | 3 min |
| Instructor | Instructor is a library for extracting validated structured data from model responses using schemas, retries, and provider-specific integrations. | Quick |
| Interpretability | Interpretability seeks to understand how a model processes information and reaches outputs, using analyses of behavior, representations, parameters, or internal computations. | Quick |
| Invoice processing | An invoice processing workflow extracts supplier, line-item, tax and payment fields, validates them and routes exceptions for review. | Quick |
| Jailbreaks | A jailbreak is a prompt written to trick an AI model into ignoring its safety training and producing something it would normally refuse. | 4 min |
| Jan | Jan is an open desktop application for running local language models and connecting to compatible hosted providers through one interface. | Quick |
| JAX | JAX is a Python numerical-computing library that adds automatic differentiation, compilation, vectorization, and accelerator support to NumPy-style programs. | Quick |
| Jev | Jev is TypeSafe AI’s decision model, returning typed choices, scores, or probabilities for software automation instead of generating open-ended natural-language responses. | Quick |
| JSON mode | JSON mode is an API setting that asks a language model to reply in JSON, text a program can parse (read into data), and on some providers, such as OpenAI, enforces valid syntax, without fixing which keys or types appear. | 4 min |
| JSON Schema | JSON Schema is a standard vocabulary for describing and validating the structure, types, constraints, and required fields of JSON data. | Quick |
| Jules | Jules is Google’s asynchronous coding agent, designed to work on repository tasks in the background, propose changes, and return results for developer review. | Quick |
| Jupyter | Jupyter is an open interactive computing environment built around notebooks that combine executable code, narrative text, data and visual output. | Quick |
| K-means clustering | K-means splits unlabelled data into k groups by putting each point with its nearest centre, then moving each centre to the average of its group, over and over. | 4 min |
| K-nearest neighbors | k-nearest neighbours labels a new example by finding the k most similar stored examples and taking their majority vote, or their average for numbers. | 4 min |
| Kaggle | Kaggle is a data-science platform for datasets, notebooks, models, competitions and community learning with hosted compute. in production. | Quick |
| Keras | Keras is a high-level deep-learning API for defining and training neural networks with a concise interface across supported computation backends. | Quick |
| Kimi K2 | Kimi K2 is Moonshot AI’s open-weight mixture-of-experts language model, designed for coding, tool use, software development, and agentic application tasks. | Quick |
| Kiro | Kiro is AWS’s agentic development environment, combining code assistance with specification-driven planning, tasks and automated implementation workflows. in software teams. | Quick |
| Kling | Kling is Kuaishou’s generative video model family, producing and editing video from text or image prompts through a commercial platform. | Quick |
| Knowledge cutoff | A knowledge cutoff is the latest period a model can reliably know from its training alone. | 3 min |
| Knowledge graph builder | A knowledge graph builder extracts entities and relationships from source material, resolves duplicates and stores linked facts with provenance. | Quick |
| Knowledge graphs | A knowledge graph stores facts as labelled links between real-world things, so software can follow the links to answer questions. | 4 min |
| Kokoro | Kokoro is a lightweight open text-to-speech model designed to produce natural-sounding speech across multiple voices with relatively modest computational requirements. | Quick |
| KV cache | A KV cache lets a language model save work from earlier tokens, so each new token does not redo the same maths. | 3 min |
| Labels | A label is the answer attached to a training example, such as a category, a number or a ranking, that a supervised model learns to predict. | 5 min |
| LAION-5B | LAION-5B is a large open dataset of image-text pairs collected from the web and filtered using CLIP representations. | Quick |
| LanceDB | LanceDB is an embedded and serverless-oriented vector database built on the Lance columnar data format for multimodal AI data. | Quick |
| LangChain | LangChain is an open-source framework for building apps and agents on top of large language models, with one standard way to talk to many AI providers. | 5 min |
| Langflow | Langflow is a visual framework for composing, testing and serving AI workflows, agents and retrieval pipelines. in production workflows. | Quick |
| Langfuse | Langfuse is an open-source platform for tracing, evaluating, prompt-managing, experimenting with, and monitoring language-model applications in development and production. | Quick |
| LangGraph | LangGraph is an open-source framework for building AI agents as graphs, where steps share one state that can be saved, paused and resumed. | 5 min |
| LangSmith | LangSmith is LangChain’s platform for tracing, testing, evaluating, debugging, and monitoring language-model applications and agent workflows throughout development and production. | Quick |
| Language learning tutor | A language-learning tutor provides conversation practice, corrections, vocabulary support and level-appropriate exercises with immediate feedback. It remains subject to human review. | Quick |
| Latency | Request latency is the elapsed time between sending an inference request and receiving its final response. | 3 min |
| Laya | Laya is ConvAI Innovations’ open-weight decision-model family, designed to return typed classifications, scores, and calibrated probabilities rather than open-ended generated text. | Quick |
| Layer normalization | Layer normalization normalizes selected activations independently for each example in a batch. | 3 min |
| Lead qualification agent | A lead qualification agent evaluates prospective customers against explicit criteria, enriches records and routes qualified opportunities for human follow-up. | Quick |
| Learning rate | The learning rate multiplies a gradient, and that product is how far the weight moves. | 3 min |
| Learning-rate schedules | A learning-rate schedule changes how big each training step is over time, often starting with a short warmup and then lowering the rate. | Quick |
| Letta | Letta is a framework for building stateful agents with explicit memory management, tools and long-running interactions. in production workflows. | Quick |
| LibreChat | LibreChat is a self-hosted multi-provider chat interface with model switching, agents, tools, retrieval and user-management features. under user control. | Quick |
| LightGBM | LightGBM is Microsoft’s gradient-boosting library, optimised for efficient training on large tabular datasets using histogram-based tree methods. in practice. | Quick |
| Linear algebra for ML | The small part of linear algebra that models use every day, vectors, matrices, their products and their sizes, to hold data, run layers and learn. | 4 min |
| Linear attention | Linear attention rewrites attention with feature maps so key-value summaries are formed before they are combined with queries. | 3 min |
| Linear regression | Linear regression predicts a number by multiplying each feature by a learned weight, adding the results and adding a bias. | 5 min |
| LiteLLM | LiteLLM provides a unified API and proxy for calling many model providers, with routing, spend tracking, retries, and observability features. | Quick |
| Llama | Llama is Meta’s family of openly available large language models, used for research, fine-tuning, and self-hosted generative AI applications. | Quick |
| Llama 4 | Llama 4 is a generation of Meta’s open-weight model family, using mixture-of-experts designs and supporting multimodal inputs in selected variants. | Quick |
| Llama Guard | Llama Guard is a family of Meta safety models that classify prompts and responses as safe or unsafe. | Quick |
| LLaMA-Factory | LLaMA-Factory is an open-source toolkit, with a web interface and a command line, for fine-tuning many open language models. | Quick |
| llama.cpp | llama.cpp is an open-source program that runs large language models on your own computer. | 4 min |
| llamafile | llamafile packages model weights and a portable inference runtime into executable files intended to run across common operating systems. | Quick |
| LlamaIndex | LlamaIndex is an open-source toolkit that brings your own files and data to a language model when you ask a question, so apps and agents can answer from them. | 4 min |
| LlamaParse | LlamaParse is LlamaIndex's document parsing service that turns PDFs and other files into text ready for LLM apps. | Quick |
| LLM | An LLM predicts a token or sequence of tokens, sometimes many paragraphs long. | 3 min |
| LLM evaluation | LLM evaluation means testing an AI feature with structured tests, so you can check how accurate and reliable it is even though its answers vary. | 5 min |
| LLM gateways | An LLM gateway is one server that sits between your apps and AI model providers, adding limits, caching, fallbacks and logs to model calls. | 5 min |
| LLM observability | LLM observability means recording what an AI app does on each request, such as the prompt, each step, the time taken and the tokens used. | 4 min |
| LLM-as-a-judge | LLM-as-a-judge means asking a strong language model to grade answers to open-ended questions. | 3 min |
| LLMOps | The day-to-day work of running an app built on a language model, from managing prompts and testing answers to watching speed and cost. | 5 min |
| llms.txt | llms.txt is a proposed website file format that gives language models a concise, curated map of important documentation and resources. | Quick |
| LM Studio | LM Studio is a desktop application for discovering, downloading and running compatible language models locally through chat and developer endpoints. | Quick |
| LMArena | LMArena is a platform that compares language models through blinded user votes on responses to the same prompts. | Quick |
| LobeChat | LobeChat is an open self-hosted chat and agent interface supporting multiple model providers, plugins, knowledge bases and multimodal interaction. | Quick |
| Local LLM chat app | A local LLM chat app runs a compatible model on user-controlled hardware and presents it through a conversational interface. | Quick |
| LocalAI | LocalAI is a self-hosted API layer for running multiple local model types behind interfaces compatible with common hosted AI services. | Quick |
| Logistic regression | Logistic regression weighs and adds up an example's features, squeezes the total into a probability between 0 and 1, and uses a threshold to pick a class. | 4 min |
| Logits | Logits are a model's raw, unnormalized prediction scores. | 3 min |
| Long context | Transformer-XL reuses hidden states from previous segments instead of computing them from scratch for each new segment. | 4 min |
| LoRA | LoRA adapts a large model to a new task by freezing its weights and training two thin matrices whose product is added to them. | 5 min |
| Loss function | A loss function turns the difference between a model's prediction and its target into a number that training tries to reduce. | 3 min |
| Lost in the middle | Language models often use facts near the opening or closing of a long input better than facts buried in the middle. | 5 min |
| Lovable | Lovable turns natural-language product requests into editable full-stack web applications, with visual iteration, code ownership and integrated deployment workflows. | Quick |
| LSTM | An LSTM is a recurrent unit whose multiplicative gates regulate access to its memory path. | 3 min |
| Luma Dream Machine | Luma Dream Machine is a generative video service that creates and modifies clips from text prompts, images and visual direction. | Quick |
| Lyria | Lyria is Google DeepMind’s generative music model family, designed to create instrumental and vocal music from textual or musical guidance. | Quick |
| Machine learning | Machine learning trains software on data so it can find patterns and make useful predictions or generate content for new inputs. | 5 min |
| Machine translation | Machine translation automatically converts text or speech from one language into another while attempting to preserve meaning, tone, and context. | Quick |
| Magistral | Magistral is Mistral AI’s reasoning-model family, designed to work through multi-step mathematics, coding, analysis, and decision problems with deliberate inference. | Quick |
| Mamba | Mamba is a sequence architecture whose state space parameters depend on the input, letting each token control what information is kept or forgotten. | 4 min |
| Manus | Manus is a general AI agent product that works through multi-step tasks using web research, files, code and other tools in a managed environment. | Quick |
| Marker | Marker converts PDFs and other documents into Markdown, JSON or structured output for search, analysis and model-based workflows. | Quick |
| Markov chains | A Markov chain models transitions among states where the next state's probability depends on the current state under the standard Markov assumption. | Quick |
| Markov decision process | A Markov decision process describes a decision problem as states, actions, rewards and the chances of moving between states; it is the standard frame for reinforcement learning. | Quick |
| Masked language modeling | Masked language modeling trains a model to guess words hidden in a sentence from the words around them; it is the training task used by BERT. | Quick |
| Mastra | Mastra is an open-source TypeScript toolkit for making AI agents, step-by-step workflows and the apps that use them. | 3 min |
| MATH benchmark | MATH is a benchmark of competition-style mathematics problems that tests multi-step reasoning across algebra, geometry, calculus, probability, and related subjects. | Quick |
| Matrices | A matrix is a rectangular grid of numbers with a fixed number of rows and columns. Machine learning keeps data and layer weights in matrices and multiplies them. | 4 min |
| Max tokens | Max tokens sets an upper bound on how many tokens a model may generate for one response. | 3 min |
| MCP | MCP standardizes how AI applications connect to external context and tools. | 4 min |
| MCP server | An MCP server exposes tools, resources or prompts through the Model Context Protocol so compatible AI clients can discover and use them. | Quick |
| Mechanistic interpretability | Mechanistic interpretability investigates neural networks by identifying internal components, representations, and computations that causally produce particular observed behaviors. | Quick |
| Meeting summarizer | A meeting summarizer turns a transcript into concise decisions, discussion themes and follow-up actions that participants can verify. | Quick |
| Mem0 | Mem0 is a memory layer for AI applications that extracts, stores and retrieves useful information across user or agent interactions. | Quick |
| Meta AI | Meta AI is Meta’s assistant across its apps, the web and a standalone app for conversation, search, voice and image creation. | 5 min |
| Metadata filtering | Metadata filtering limits a vector search to records whose labels, such as category or year, match rules you set. | 4 min |
| Microsoft 365 Copilot | Microsoft 365 Copilot embeds AI assistance across Word, Excel, PowerPoint, Outlook, Teams and organisational data for drafting, analysis and collaboration. | Quick |
| Microsoft Agent Framework | Microsoft Agent Framework is Microsoft's open-source framework for building AI agents and multi-agent workflows, the successor to AutoGen and Semantic Kernel. | Quick |
| Microsoft Copilot | Microsoft Copilot is an AI helper from Microsoft that you can type to, talk to or show a picture, and it will reply, write something new or carry out a job for you. | 5 min |
| Microsoft Copilot Studio | Microsoft Copilot Studio is a managed environment for designing, connecting, governing and publishing agents across Microsoft and external systems. | Quick |
| Microsoft Foundry | Microsoft Foundry is Microsoft's Azure platform for building AI apps and agents, with a large model catalogue, tools, testing and safety controls in one place. | 5 min |
| Midjourney | Midjourney is a generative image service for prompt-driven visual creation and iterative editing through its web and community interfaces. | Quick |
| Milvus | Milvus is a distributed vector database built for large-scale similarity search across high-dimensional embedding collections. in production workflows. | Quick |
| MiniMax | MiniMax is a family of language and multimodal models from the company MiniMax, used for text, speech, music, image, and video applications. | Quick |
| Ministral | Ministral is Mistral AI’s compact model family, designed for edge, on-device, and latency-sensitive language applications with constrained computing resources. | Quick |
| Missing data | Missing data are empty cells in a dataset. You can drop them, fill them with reasoned guesses or use a model that accepts gaps, and why they are missing decides which is safe. | 5 min |
| Mistral La Plateforme | Mistral La Plateforme provides APIs, model deployment, fine-tuning and agent-building services around Mistral’s model family. in production. in production. | Quick |
| Mistral Large | Mistral Large is Mistral AI’s higher-capability commercial model tier for complex reasoning, multilingual work, code, and enterprise applications. | Quick |
| Mistral Medium | Mistral Medium is a mid-tier Mistral AI model positioned between small, efficient models and the company’s highest-capability offerings. | Quick |
| Mistral Small | Mistral Small is Mistral AI’s efficiency-focused model tier for responsive conversational, coding, and business automation workloads at production scale. | Quick |
| Mixed precision | Mixed precision trains with both 16-bit and 32-bit numbers so the run is faster and uses less memory. | 4 min |
| Mixtral | Mixtral is Mistral AI’s open mixture-of-experts model family, activating only part of the network for each token to improve efficiency. | Quick |
| Mixture of experts | A mixture-of-experts layer uses a learned gate to select a sparse combination of expert subnetworks for each input. | 3 min |
| MLC LLM | MLC LLM compiles and deploys language models across GPUs, CPUs, browsers and mobile devices using machine-learning compilation techniques. | Quick |
| MLflow | MLflow is an open platform for tracking experiments, packaging models, managing versions and supporting machine-learning deployment workflows in practical workflows. | Quick |
| MLOps | MLOps is a way of working that makes building and releasing machine learning models simpler and more automatic. | 5 min |
| MLX | MLX is Apple's open-source array framework for machine learning on Apple silicon, used to run and fine-tune models on Macs. | Quick |
| MMLU | MMLU is a 57-task test of a text model's multitask accuracy. | 3 min |
| MNIST | MNIST is a classic dataset of handwritten digit images, commonly used to teach and benchmark basic image-classification systems. | Quick |
| Model cards | Model cards are structured documents describing a model's intended uses, evaluation results, limitations, training context, and important ethical or safety considerations. | Quick |
| Model collapse | Model collapse is degradation that can occur when successive models train heavily on generated data, losing diversity or fidelity to the original distribution. | Quick |
| Model merging | Model merging combines parameters or learned changes from multiple related models to blend capabilities without conventional joint retraining. | Quick |
| Model parallelism | Model parallelism divides a model's computation or parameters across multiple devices when one device cannot efficiently hold or run it. | Quick |
| Model routing | Model routing sends each prompt to the model that suits it, so simple requests go to cheaper models and hard ones to stronger models. | 4 min |
| Model serving | Model serving means running a trained model on a server so apps can send it requests over the network and get answers back. | 5 min |
| Monte Carlo methods | Monte Carlo methods estimate a number by running many random simulations and averaging what comes out, instead of working the answer out exactly. | 4 min |
| Moshi | Moshi is Kyutai’s open speech-language model for real-time, full-duplex spoken conversation with low-latency audio input, understanding, generation, and output. | Quick |
| Movie Gen | Movie Gen is Meta research on generative media models for creating and editing video and producing synchronized audio from prompts. | Quick |
| Multi-agent systems | A multi-agent system is a group of AI agents, each a language model using tools in its own loop, that work together on one job. | 5 min |
| Multi-agent team | A multi-agent team divides work among specialised model-driven roles, with explicit handoffs, shared state and a final integration step. | Quick |
| Multi-armed bandits | A multi-armed bandit applies an action, observes its reward, then continues the process with another action. | 3 min |
| Multi-head attention | Multi-head attention consists of several attention layers running in parallel. | 3 min |
| Multilayer perceptron | A multilayer perceptron can fit a nonlinear model to training data. | 3 min |
| Multimodal models | A multimodal model can process more than one type of input, such as text, images, audio or video. | 4 min |
| Muse | Muse is Meta’s personal AI agent for carrying out tasks and longer-term goals across connected apps, using a dedicated secure virtual machine and asking for approval before sensitive actions. | Quick |
| Muse Spark 1.1 | Muse Spark 1.1 is Meta’s multimodal reasoning model for agentic tasks, with tool and computer use, coding, long-context work, and multi-agent orchestration through Meta AI and the Meta Model API. | Quick |
| Music generation models | A music generation model can generate music in the raw-audio domain. | 3 min |
| n8n | n8n is a workflow-automation platform with visual and code-based building blocks, including AI nodes, agents and integrations in practical workflows. | Quick |
| Naive Bayes | Naive Bayes sorts things into classes by multiplying how common each class is by how likely each feature is in that class, then picking the top score. | 4 min |
| Named entity recognition | Named entity recognition finds and categorizes mentions such as people, organizations, locations, dates, products, or quantities within text. | Quick |
| Nano Banana | Nano Banana is a widely used nickname associated with a Google Gemini image-generation and editing model, especially for conversational visual transformations. | Quick |
| Narrow AI | Narrow AI is designed for a limited task or domain, even when its performance there appears sophisticated or exceeds human ability. | Quick |
| Natural language processing | Natural language processing develops computational methods for understanding, generating, translating, searching, and analyzing human language in text or speech. | Quick |
| NeMo Guardrails | NeMo Guardrails is NVIDIA's open-source toolkit for adding programmable rules around a conversational AI app, such as blocking unsafe topics or keeping answers on subject. | Quick |
| Nemotron | Nemotron is NVIDIA’s family of language models and training resources, often used to create synthetic data and develop customized enterprise models. | Quick |
| Neural networks | A neural network is a model built from layers of simple units that each weigh their inputs, add them up and bend the result, so that together they can learn curved, complex patterns. | 4 min |
| Neural radiance fields | A neural radiance field maps a 3D position and viewing direction to density and view-dependent color. | 3 min |
| Next-token prediction | Next-token prediction estimates which token should follow the tokens already present in a sequence. | 3 min |
| Next.js AI chatbot | A Next.js AI chatbot combines a React interface, server routes, model streaming and storage into a deployable conversational web application. | Quick |
| NLLB | NLLB, or No Language Left Behind, is Meta’s machine-translation research program and model family covering many low-resource languages. | Quick |
| Nomic Embed | Nomic Embed is Nomic AI’s open text-embedding model family, designed for retrieval, clustering, classification, and long-context semantic representation across documents. | Quick |
| NotebookLM | NotebookLM is Google’s source-grounded research and learning tool for asking questions, creating summaries and generating audio or study material from selected sources. | Quick |
| Notion AI | Notion AI adds writing, summarisation, search and workflow assistance to Notion pages, databases and connected workspace knowledge in practical workflows. | Quick |
| Nous Research | Nous Research is an open AI research community and company known for publishing language models, datasets and training work. | Quick |
| NPUs | An NPU is a part of a chip built to speed up AI models at low power. | 5 min |
| NVIDIA Blackwell | NVIDIA Blackwell is a GPU architecture and computing platform designed for large-scale AI training and inference in data centres. | Quick |
| NVIDIA CUDA | CUDA is NVIDIA’s parallel computing platform and programming model for running general-purpose workloads, including AI, on NVIDIA GPUs. | Quick |
| NVIDIA H100 | NVIDIA H100 is a data-centre GPU based on the Hopper architecture, widely used for training and serving large AI models. | Quick |
| NVIDIA NIM | NVIDIA NIM packages optimised model inference as deployable microservices for NVIDIA-accelerated infrastructure and enterprise AI applications. in production. | Quick |
| o3 | o3 is an OpenAI reasoning model designed for complex analysis, coding, mathematics, visual reasoning, and tasks that benefit from deliberate computation. | Quick |
| Object detection | Object detection identifies and locates instances of specified object categories within an image or video, usually using bounding boxes. | Quick |
| OCR | Optical character recognition converts text visible in images or scanned documents into machine-readable characters for searching, editing, or analysis. | Quick |
| Ollama | Ollama is an open-source program that downloads open AI models and runs them on your own computer, with a simple command and a local API. | 4 min |
| OLMo | OLMo is Ai2’s fully open language-model family, publishing model weights, training data, code, and research artifacts for transparent study. | Quick |
| Omni models | An omni model processes several input and output modalities within one multimodal model rather than routing everything through a text-only core. | 4 min |
| On-device AI | On-device AI runs the model on your own phone, laptop or browser, so your data does not have to go to a server. | 4 min |
| On-device assistant | An on-device assistant runs some or all inference locally to reduce latency, limit data transfer and work with device context. | Quick |
| One-hot encoding | One-hot encoding turns a category, such as a colour or a type of transport, into a list of 0s with a single 1 marking which category it is. | 4 min |
| ONNX | ONNX is an open format for representing machine-learning models, enabling models to move between frameworks and inference runtimes. | Quick |
| ONNX Runtime | ONNX Runtime is an open inference engine for running models exported in the ONNX format across different hardware and operating systems. | Quick |
| Open WebUI | Open WebUI, a free chat app you host yourself, lets you talk to AI models running on your own computer or in the cloud, from one web page. | 4 min |
| Open-weight models | An open-weight model makes its trained weights publicly available for download. | 3 min |
| OpenAI Agents SDK | OpenAI's Agents SDK is a small open-source toolkit from OpenAI for building AI agents that use tools, pass work to each other and can record each run as a trace. | 5 min |
| OpenAI API | The OpenAI API gives developers hosted access to OpenAI models and tools for text, audio, vision, images and agentic applications. | Quick |
| OpenAI Codex | OpenAI Codex is an agentic software-development product that can inspect repositories, edit code, run commands and collaborate on engineering tasks. | Quick |
| OpenAI embeddings | OpenAI embeddings are API models that convert text into numerical vectors for semantic search, clustering, recommendations, classification, and retrieval systems. | Quick |
| OpenAI Evals | OpenAI Evals is an open-source framework and registry for evaluating model behavior on structured test cases and custom tasks. | Quick |
| OpenAI-compatible APIs | OpenAI-compatible APIs imitate common OpenAI request and response formats, letting existing clients connect to other model providers with fewer changes. | Quick |
| OpenAPI | OpenAPI is a machine-readable standard for describing HTTP APIs, including endpoints, inputs, outputs, authentication methods, errors, and reusable data schemas. | Quick |
| OpenCV | OpenCV is a computer-vision library for image and video processing, feature extraction, camera workflows and model integration. in practice. | Quick |
| OpenRouter | OpenRouter provides one API and routing layer for accessing models from multiple providers, with unified billing, usage and availability controls. | Quick |
| OpenSearch | OpenSearch is a search and analytics suite with keyword, vector and hybrid search capabilities derived from the Elasticsearch ecosystem. | Quick |
| OpenVINO | OpenVINO is Intel’s toolkit for optimising and running AI inference across supported Intel CPUs, GPUs and accelerators. in practice. | Quick |
| Opik | Opik is Comet’s open-source platform for tracing, evaluating, testing, debugging, and monitoring language-model applications and agent workflows across development and production. | Quick |
| Optuna | Optuna is a hyperparameter-optimisation framework that searches configuration spaces, prunes weak trials and records experiment results. in practice. | Quick |
| Otter.ai | Otter.ai records, transcribes and summarises meetings, with collaboration features for reviewing conversations, identifying speakers and tracking follow-up items. | Quick |
| Outliers | An outlier is a data point that sits an unusually long way from the rest. It may be a mistake to fix or a rare real event worth keeping. | 4 min |
| Outlines | Outlines is a library for constrained generation, guiding language models to produce outputs that follow regex patterns, grammars, or typed schemas. | Quick |
| Overfitting | Overfitting happens when a model learns its training examples so specifically that it performs poorly on new examples. | 3 min |
| PagedAttention | PagedAttention stores a model's attention memory in small fixed-size blocks, like pages in computer memory, so a server can fit and share more requests at once; it was introduced with vLLM. | Quick |
| Parameter-efficient fine-tuning | PEFT adapts a pretrained model by training a small set of added or selected parameters while keeping most base weights frozen. | 4 min |
| Parameters | Parameters are the model values that training can change to improve how inputs map to outputs. | 3 min |
| PEFT | PEFT is a free Hugging Face library that teaches a big AI model a new task by training a small set of extra weights instead of the whole model. | 4 min |
| Perceptron | A perceptron is a linear classifier. | 3 min |
| Perplexity | Perplexity is an AI search product that searches sources, synthesizes an answer and links the evidence used. | 5 min |
| Perplexity (metric) | Perplexity is the exponentiated average negative log-likelihood that a language model assigns to a token sequence. | 4 min |
| Perplexity Comet | Comet is Perplexity’s AI browser, combining conventional browsing with page-aware assistance, search and agent-like actions across open tabs and websites. | Quick |
| Personal knowledge base | A personal knowledge base indexes a user’s notes and files for private search, summarisation and question answering. It remains subject to human review. | Quick |
| Personalization | Personalization adapts content, recommendations, interfaces, or services to an individual's observed behavior, stated preferences, context, or predicted needs. | Quick |
| pgvector | pgvector is an open PostgreSQL extension that adds vector storage, distance operators and indexes for similarity search in practical workflows. | Quick |
| Phi | Phi is Microsoft’s family of small language models, designed to deliver useful reasoning and language capability with modest compute requirements. | Quick |
| Phi-4 | Phi-4 is a Microsoft small-language-model generation focused on efficient reasoning, mathematics, coding, instruction following, and selected multimodal application workloads. | Quick |
| Pinecone | Pinecone is a managed vector database for semantic search, recommendation and retrieval workloads in AI applications. in production. | Quick |
| Pipeline parallelism | Pipeline parallelism places consecutive model stages on different devices and overlaps their work using multiple batches moving through the pipeline. | Quick |
| Pixtral | Pixtral is Mistral AI’s vision-language model family, capable of understanding images alongside text in conversational and document workflows. | Quick |
| Podcast transcriber | A podcast transcriber converts long-form audio into timestamped text, often adding speaker labels, chapters or searchable excerpts. It remains subject to human review. | Quick |
| Positional encoding | Positional encoding adds token-order information to representations before attention reads them. | 3 min |
| PPO | Proximal policy optimization is a reinforcement-learning algorithm that limits the size of policy updates to improve training stability. | Quick |
| Precision and recall | Precision asks how many of the things a model flagged were right. Recall asks how many of the real cases it managed to find. | 5 min |
| Predictive maintenance | Predictive maintenance uses sensor data and models to estimate equipment failures or degradation, enabling service before costly breakdowns occur. | Quick |
| Prefill and decode | Prefill processes all prompt tokens in parallel to initialize model state; decode then generates new tokens sequentially using that cached state. | Quick |
| Pretraining | In one transfer-learning setup, a model is pretrained on a data-rich task before fine-tuning on a downstream task. | 3 min |
| Principal component analysis | PCA finds the directions in which data spreads out most, then keeps only the first few, so many features become a few new ones. | 4 min |
| Probability | A probability is a number from 0 to 1 that says how likely something is. Many machine learning problems need one as the answer. | 4 min |
| Prompt caching | Prompt caching saves the work a model did on the unchanged start of a prompt, so the next request with the same start is faster and cheaper. | 4 min |
| Prompt chaining | In prompt chaining, you give a model one small job at a time in a set order, and each answer becomes the starting material for the next job. | 5 min |
| Prompt engineering | Prompt engineering is writing and testing the instructions and examples you give a model so its answer meets a goal you can check. | 4 min |
| Prompt injection | Prompt injection is text that sneaks new instructions into what an AI model reads, so the model behaves in unintended ways. | 5 min |
| Prompt playground | A prompt playground lets teams compare instructions, models and settings on shared test cases before changes reach production. | Quick |
| Prompt templates | Prompt templates are reusable prompt structures with variable fields, helping applications apply consistent instructions and context across many requests. | Quick |
| Prompt versioning | Prompt versioning tracks changes to prompts alongside metadata and results, making experiments reproducible and production behavior easier to audit. | Quick |
| promptfoo | promptfoo is an open-source tool for testing prompts, models, and RAG systems with configurable assertions, comparisons, and security evaluations. | Quick |
| Protein structure prediction | Protein structure prediction estimates a protein's three-dimensional shape from its amino-acid sequence or related evidence, helping researchers investigate function and interactions. | Quick |
| Pruning | Pruning removes weights, units, or connections judged less important, shrinking computation or storage while aiming to preserve model quality. | Quick |
| Pydantic AI | Pydantic AI is a Python framework for building AI agents whose tool inputs and final answers are checked against types you define. | 4 min |
| PyTorch | PyTorch is a free Python library for building and training neural networks, with fast maths on GPUs and automatic gradients. | 4 min |
| PyTorch Geometric | PyTorch Geometric is a library of data structures, layers and utilities for graph neural networks built on PyTorch. | Quick |
| PyTorch Lightning | PyTorch Lightning is a framework that organizes PyTorch training code, reducing boilerplate while supporting distributed training, logging, and reproducible experiments. | Quick |
| Q-learning | Q-learning improves an estimate of each state-action value from sampled rewards and the best estimated value at the next state. | 3 min |
| Qdrant | Qdrant is a vector database and search engine with filtering, payload storage and APIs for production semantic retrieval. | Quick |
| QLoRA | QLoRA fine-tunes a large language model by storing its frozen weights in 4 bits and training only a small add-on that sits beside them. | 5 min |
| Qualcomm Snapdragon X | Qualcomm Snapdragon X is a family of Arm-based PC processors with integrated neural-processing hardware for Windows laptops. in practical systems. | Quick |
| Quantization | Quantization stores a model's numbers with fewer bits, such as 8-bit integers instead of 32-bit decimals, so it needs less memory. | 4 min |
| Query rewriting | Query rewriting transforms a user’s request into one or more clearer search queries, improving retrieval when the original wording is vague or conversational. | Quick |
| Question answering | Question-answering systems produce answers to natural-language questions using learned knowledge, provided context, retrieved sources, databases, tools, or combinations of these. | Quick |
| Qwen | Qwen is Alibaba Cloud’s family of language and multimodal models, covering general conversation, coding, vision, audio, and specialized tasks. | Quick |
| Qwen3 | Qwen3 is a generation of Alibaba’s open-weight model family, offering multiple sizes and modes for both direct responses and extended reasoning. | Quick |
| Qwen3-Coder | Qwen3-Coder is an Alibaba model family specialized for software development, repository-scale understanding, tool use, and agentic coding workflows. | Quick |
| RAG | RAG lets an AI model look up relevant documents first, then answer from what it found. | 4 min |
| RAG chatbot | A RAG chatbot retrieves relevant passages from a chosen knowledge source and gives them to a language model before it answers. | Quick |
| Ragas | Ragas is an open-source evaluation framework for retrieval-augmented generation, measuring retrieved-context relevance, answer faithfulness, and end-to-end response quality. | Quick |
| RAGFlow | RAGFlow is an open-source retrieval-augmented generation engine that focuses on reading complex documents, such as tables and scanned pages, before answering questions about them. | Quick |
| Random forest | A random forest combines many decision trees trained on varied samples and features, reducing individual-tree instability through aggregated predictions. | Quick |
| Rate limits | A rate limit caps how much you can use an AI service in a set stretch of time, such as how many requests you send each minute. | 4 min |
| Ray | Ray is a distributed computing framework for scaling Python and machine-learning workloads from one machine to a cluster. | Quick |
| Ray-Ban Meta glasses | Ray-Ban Meta glasses combine cameras, microphones, speakers, and Meta AI so wearers can capture media, ask questions, and use voice-controlled assistance. | Quick |
| ReAct | ReAct is a prompting method in which a language model alternates written thoughts with tool actions, reading each result before it picks the next step. | 5 min |
| Reasoning models | Reasoning models use intermediate reasoning tokens before producing a final response. | 4 min |
| Recommendation engine | A recommendation engine ranks items using behavioural, content or business signals and evaluates whether its suggestions are useful. | Quick |
| Recommender systems | A recommender system picks the few items, out of a catalogue too big to browse, that each person sees first, by predicting what that person is likely to want. | 4 min |
| Recurrent neural networks | At each time step, a recurrent neural network updates its hidden state from the previous hidden state and the current input. | 3 min |
| Red teaming | Red teaming means deliberately testing an AI system the way an attacker would, to find its weak spots. | 4 min |
| Regression | Regression predicts a number, such as a house price, a tree's life span or a rainfall total, from an example's features. | 3 min |
| Regularization | Regularization discourages overly complex model behavior, often by penalizing large weights or adding noise, to improve performance on unseen data. | Quick |
| Reinforcement learning | Reinforcement learning trains a program, called an agent, to choose actions by trial and error, guided by a number called a reward that says how well it is doing. | 5 min |
| Reinforcement learning with verifiable rewards | Reinforcement learning with verifiable rewards trains models on tasks whose answers can be checked automatically, such as by tests, proofs, or exact calculations. | Quick |
| ReLU | ReLU is an activation function that replaces every negative input with zero and leaves every positive input unchanged. | 3 min |
| Replicate | Replicate is a hosted platform and API for running published machine-learning models without managing their serving infrastructure directly. | Quick |
| Replit Agent | Replit Agent builds, tests and deploys applications inside Replit from natural-language requests, while exposing the project for further editing. | Quick |
| Reranker models | A reranker sorts existing text candidates by semantic relevance to a specified query. | 3 min |
| Reranking | Reranking takes the short list a first search returned and scores each item again with a slower, more careful model. | 4 min |
| Research agent | A research agent breaks a question into searches, gathers evidence, compares sources and produces a cited synthesis for review. | Quick |
| Residual connections | A residual connection combines a residual mapping F(x) with the block input x as F(x) + x. | 4 min |
| Residual networks | A residual block adds a learned residual function F(x) to a shortcut carrying the block input x. | 3 min |
| Resume screener | A resume screener compares application material with explicit job criteria and presents evidence for human review rather than making an opaque hiring decision. | Quick |
| Reward hacking | Reward hacking occurs when an AI system finds an unintended way to maximize its measured reward without achieving the outcome designers actually wanted. | Quick |
| Reward models | Reward models assign scores to candidate behaviors or outputs, approximating preferences or objectives that reinforcement learning can optimize. | Quick |
| RLAIF | RLAIF runs the usual RLHF training loop, but another AI model supplies some of the preference labels that people would normally give. | 4 min |
| RLHF | RLHF uses human comparisons to learn a reward signal, then optimizes a model to produce responses that score better under that signal. | 4 min |
| RoBERTa | RoBERTa is a robustly optimized BERT-style encoder that changed training choices and data scale to improve language-understanding performance. | Quick |
| Robotics | Robotics combines sensing, computation, planning, mechanical design, and physical control to build machines that act in the real world. | Quick |
| Robotics foundation models | A robotics foundation model is pretrained across many tasks, environments or robot bodies so it can serve as a starting point for new robot policies. | 4 min |
| ROC curve | An ROC curve shows, for every threshold a classifier could use, how many real positives it catches against how many false alarms it raises. | 4 min |
| Rotary position embeddings | RoPE encodes absolute position with a rotation matrix and adds explicit relative-position dependence to self-attention. | 3 min |
| Runway | Runway is a creative AI platform for generating and editing video, images and other media through models and production tools. | Quick |
| RWKV | RWKV is an open neural architecture combining transformer-like training with recurrent inference, allowing language models to process sequences without a key-value cache. | Quick |
| Safetensors | Safetensors is a secure, fast tensor-storage format designed to load and share model weights without executing arbitrary serialized program code. | Quick |
| Sales email writer | A sales email writer drafts personalised outreach from approved product facts and recipient context, leaving claims and final sending under human control. | Quick |
| Salesforce Agentforce | Salesforce Agentforce is a platform for building and deploying AI agents that use Salesforce data, workflows and authorised business actions. | Quick |
| SAM 2 | SAM 2 is Meta’s Segment Anything model for images and video, supporting prompted object segmentation and tracking across video frames. | Quick |
| Sandboxing agents | Sandboxing an agent means running its commands inside a fence, set up by the operating system, that limits which files and websites they can reach. | 4 min |
| Scalable oversight | Scalable oversight develops ways for people to supervise systems on tasks too numerous or difficult for unaided humans to evaluate directly. | Quick |
| Scaling laws | Scaling laws are empirical relationships describing how model performance changes predictably with factors such as parameter count, training data, and computation. | Quick |
| scikit-learn | scikit-learn is an open Python library for classical machine learning, offering consistent tools for preprocessing, modelling, evaluation and pipelines. | Quick |
| Seedance | Seedance is ByteDance’s generative video model family for producing coherent, multi-shot video sequences from text or image instructions and references. | Quick |
| Seedream | Seedream is ByteDance’s image-generation model family, designed for prompt-based image creation and editing with strong text and visual fidelity. | Quick |
| Segment Anything | Segment Anything is Meta’s foundation model for image segmentation, producing object masks from points, boxes, or other visual prompts. | Quick |
| Self-attention | Self-attention updates each position by comparing it with positions in the same sequence and mixing their information. | 3 min |
| Self-consistency | Self-consistency generates several independent reasoning paths for the same problem and selects the answer that appears most consistently across them. | Quick |
| Self-hosting LLMs | Self-hosting means running a language model on a computer you control, so your prompts go to your own machine instead of a company's service. | 4 min |
| Self-reflection | Self-reflection is when an AI model checks its own first answer, writes feedback on it, and then tries again using that feedback. | 3 min |
| Self-supervised learning | Self-supervised learning trains a model on raw data by hiding part of each example and asking the model to predict it, so the data supplies its own labels. | 5 min |
| Semantic caching | Semantic caching reuses earlier responses when a new request is sufficiently similar in meaning, reducing latency and model cost. | Quick |
| Semantic Kernel | Semantic Kernel is an open-source Microsoft toolkit that connects AI models to your own code, so a model can ask your functions to do real work. | 4 min |
| Semantic search | Semantic search finds text by meaning, so a passage can match a question even when the two share no words. | 3 min |
| Semantic search engine | A semantic search engine embeds queries and indexed content so it can retrieve material by meaning, not only matching words. | Quick |
| Semi-supervised learning | Semi-supervised learning trains a model on a few labelled examples plus many unlabelled ones, so the unlabelled data can help when answers are scarce. | 5 min |
| Sentence Transformers | Sentence Transformers is a Python library for creating and using text or image embeddings for similarity, retrieval and clustering. | Quick |
| Sentence-BERT | Sentence-BERT adapts BERT into a model that produces useful sentence embeddings, making semantic similarity and retrieval substantially more efficient. | Quick |
| Sentiment analysis | Sentiment analysis estimates attitudes or emotional polarity in text, such as positive, negative, or neutral, often for a specific target. | Quick |
| Sentiment dashboard | A sentiment dashboard classifies feedback, aggregates trends and preserves links to source material so teams can inspect what drove the summary. | Quick |
| SEO content assistant | An SEO content assistant organises research, search intent and page structure into drafts that still require factual and editorial review. | Quick |
| SGLang | SGLang is open-source server software that runs large language models on GPUs and answers many requests quickly, partly by reusing work shared between prompts. | 4 min |
| SigLIP | SigLIP is Google’s vision-language representation model, aligning images and text with a sigmoid-based training objective for retrieval and classification. | Quick |
| Siri | Siri is Apple’s voice assistant for device controls, personal requests and questions across Apple hardware and connected services. | Quick |
| Slack bot | A Slack bot receives messages or events, applies defined logic or model calls and responds through authorised workspace actions. | Quick |
| Small language models | Small-model research includes sub-billion-parameter language models for mobile deployment. | 3 min |
| Smolagents | Smolagents is a small open-source Python library from Hugging Face for building AI agents, including ones that act by writing short pieces of Python code. | 4 min |
| Social media assistant | A social media assistant adapts approved messages into platform-specific drafts, schedules and response suggestions for human review. It remains subject to human review. | Quick |
| Softmax | Softmax determines a probability for each possible class, and those probabilities add up to exactly 1. | 3 min |
| Sora | Sora is OpenAI’s generative video model, creating and transforming video from text, images, or existing footage through a prompt-driven workflow. | Quick |
| spaCy | spaCy is a production-oriented Python library for natural-language processing, including tokenisation, tagging, parsing and entity recognition. in practice. | Quick |
| Sparse attention | Sparse attention computes selected query-key connections instead of filling the entire attention matrix. | 3 min |
| Sparsity | Sparsity means that many possible model values or computation paths are zero, absent, or inactive for a given input. | 4 min |
| Spec-driven development | Spec-driven development begins with an explicit description of behavior, constraints, and acceptance criteria that guides human and AI implementation work. | Quick |
| Speculative decoding | Speculative decoding uses a target model to evaluate guesses in parallel. | 4 min |
| Speech-to-text app | A speech-to-text application accepts audio, transcribes speech and presents time-aligned text for review or further processing. It remains subject to human review. | Quick |
| Speech-to-text models | A speech-to-text model can transcribe a speech utterance into written characters. | 3 min |
| Spreadsheet assistant | A spreadsheet assistant helps create formulas, clean tables, analyse values and explain calculations within bounded workbook context. It remains subject to human review. | Quick |
| Spring AI | Spring AI brings model clients, vector stores, tools, structured outputs, and retrieval patterns into the Spring application ecosystem. | Quick |
| SQuAD | SQuAD is a widely used reading-comprehension dataset whose human-written questions are answered from passages drawn from English Wikipedia articles. | Quick |
| Stable Audio | Stable Audio is Stability AI’s generative audio model family for creating music, sound effects, and other audio from text prompts. | Quick |
| Stable Diffusion | Stable Diffusion is Stability AI’s open latent-diffusion model family for generating and editing images from text and image prompts. | Quick |
| Stable Diffusion 3 | Stable Diffusion 3 is a Stability AI image-model generation using a multimodal diffusion-transformer architecture to improve prompt adherence and typography. | Quick |
| Stagehand | Stagehand is an open-source framework from Browserbase for building AI agents that control a web browser, mixing plain-language instructions with regular browser automation code. | Quick |
| StarCoder | StarCoder is the BigCode collaboration’s open code-model family, trained on permissively licensed source code for generation and completion. | Quick |
| State space models | A continuous-time linear state space model uses x-dot = Ax + Bu for state evolution and y = Cx + Du for readout. | 4 min |
| Statistics for ML | Statistics describes data and measures how far to trust a model's score, since every test set is only a sample of the cases the model will meet. | 4 min |
| Stochastic gradient descent | Stochastic gradient descent steps the weights using a slope estimated from one randomly picked example. | 3 min |
| Stop sequences | A stop sequence is a configured character string that stops output generation. | 3 min |
| Strands Agents | Strands Agents is a free AWS toolkit, released as open source under Apache 2.0, for making AI agents with Python or TypeScript, where the model itself plans the steps and picks the tools. | 5 min |
| Streaming responses | Streaming responses deliver a model’s output incrementally as it is generated, improving perceived speed and enabling responsive conversational interfaces. | Quick |
| Streamlit | Streamlit is a Python framework for turning data scripts and model workflows into interactive web applications. in practice. | Quick |
| Streamlit chatbot | A Streamlit chatbot uses Streamlit’s Python interface components to wrap a model call in a simple interactive web application. | Quick |
| Structured data extraction | A structured data extraction workflow converts unstructured documents into a defined schema, then validates required fields and uncertain values. | Quick |
| Structured outputs | Structured outputs constrain a model response to a declared schema so software receives expected fields and types instead of free-form prose. | 3 min |
| Study tutor | A study tutor explains material, asks diagnostic questions and adjusts practice to the learner while avoiding simply completing assessed work. | Quick |
| Sub-agents | A sub-agent is a helper AI that a main agent sends off to do one side task in its own workspace, then report back a short result. | 3 min |
| Summarization | Summarization is when a model turns a long text into a shorter version that keeps the important information. | 3 min |
| Suno | Suno generates complete songs from text descriptions or lyrics, combining composition, vocals and production in a consumer music-creation interface. | Quick |
| Superintelligence | Superintelligence is a hypothetical form of intelligence that substantially exceeds the best human abilities across most cognitively demanding fields. | Quick |
| Supervised fine-tuning | Supervised fine-tuning keeps training an already-trained model on example questions paired with good answers, so it learns to answer that way. | 4 min |
| Supervised learning | Supervised learning trains a model on examples that already carry the right answer, called a label, so it can predict that answer for new examples that do not. | 4 min |
| Support vector machines | A support vector machine separates two classes with the boundary that leaves the widest possible gap to the nearest training points of each class. | 5 min |
| SWE-bench | SWE-bench asks a system to generate a repository patch for a real GitHub issue. | 3 min |
| Sycophancy | Sycophancy is when an assistant favors matching a user's beliefs over a truthful response. | 3 min |
| Symbolic AI | Symbolic AI writes knowledge down as readable symbols and if-then rules, then reaches answers by applying logic and search to them. | 4 min |
| Synthesia | Synthesia creates presenter-led videos from scripts using synthetic avatars, voices, templates and enterprise publishing controls. for creative production. | Quick |
| Synthetic data | Synthetic data is artificial data built from seed data so some patterns of that seed can be reused. | 4 min |
| Synthetic data generator | A synthetic data generator creates artificial examples under defined constraints, then checks coverage, realism, privacy and downstream usefulness. | Quick |
| System prompt | A system prompt is high-priority context that sets an AI assistant's role, boundaries, style, and response rules before the user's request is handled. | 3 min |
| t-SNE | t-SNE is a nonlinear visualization method that places similar high-dimensional examples nearby in two or three dimensions, but global distances may be misleading. | Quick |
| T5 | T5 is Google’s text-to-text transformer framework, recasting many language tasks so both inputs and outputs use a unified text format. | Quick |
| Temperature | Temperature enters softmax by dividing each logit by T. | 3 min |
| Tensor parallelism | Tensor parallelism splits individual tensor operations and parameter matrices across devices, allowing them to collaborate within the same model layer. | Quick |
| TensorFlow | TensorFlow is Google’s open-source machine-learning framework for building, training, and deploying models across servers, browsers, mobile devices, and specialized hardware. | Quick |
| TensorRT-LLM | TensorRT-LLM is NVIDIA's open-source library for compiling and running large language models efficiently on NVIDIA GPUs, aimed at fast, low-cost serving in production. | Quick |
| Test set | A test set is a reserved group of labeled examples used for a final check of a model on data it did not train on. | 3 min |
| Test-time compute | Test-time compute lets a language model improve its output by using more computation at test time. | 3 min |
| Text classification | Text classification assigns predefined labels to documents or passages, supporting tasks such as topic detection, moderation, routing, and intent recognition. | Quick |
| Text Generation Inference | Text Generation Inference (TGI) is Hugging Face's open-source toolkit for deploying and serving large language models behind an API, with batching and streaming built in. | Quick |
| Text-to-image models | A text-to-image model can use text as a conditioning input to an image generator. | 3 min |
| Text-to-speech models | A text-to-speech model synthesizes speech directly from text. | 3 min |
| Text-to-SQL | Text-to-SQL systems translate natural-language questions into database queries, often adding database schema context, validation, permissions, and execution safeguards. | Quick |
| Text-to-video models | Given a text prompt, a text-to-video model generates a video. | 3 min |
| The Stack | The Stack is BigCode’s large source-code dataset, assembled for training code models with tools for filtering and attribution. | Quick |
| Throughput | Throughput is the amount of completed inference work divided by the time spent measuring it. | 3 min |
| Time series forecasting | Time series forecasting predicts future values of something measured over time, such as sales or electricity demand, from the patterns in its past. | 5 min |
| Time to first token | Time to first token is how long a model takes to start its reply after receiving a request, a key measure of how responsive an AI app feels. | Quick |
| timm | timm is a PyTorch library of image models, pretrained weights, training utilities and reproducible computer-vision recipes. in practice. | Quick |
| Together AI | Together AI provides hosted inference, fine-tuning and GPU infrastructure for open and custom generative AI models. in production. | Quick |
| Token costs | AI model APIs charge by the token, with separate prices for the text you send in and the text the model writes back. | 4 min |
| Tokenization | Tokenization turns text into a sequence of vocabulary IDs that a model can process, then reverses generated IDs back into readable text. | 3 min |
| Tokens | Tokens are chunks of text processed by text-generation and embedding models. | 3 min |
| Tokens per second | Tokens per second reports how many tokens a server returns in a set amount of time. | 3 min |
| Tool use | Tool use is a loop where a model proposes a call, a program outside the model runs it, and the result is written back. | 5 min |
| Top-k sampling | Top-k sampling keeps the k highest-probability next-token candidates, rescales their probabilities and samples one of them. | 3 min |
| Top-p sampling | Top-p keeps the smallest set of most probable tokens whose probabilities add up to at least p. | 3 min |
| TPUs | TPUs are chips Google designed to speed up the maths of machine learning, rented out through Google Cloud. | 5 min |
| Training | Training is the repeated loop that nudges a model's internal numbers so its predictions get closer to the right answers in its examples. | 5 min |
| Training data | Training data is the set of examples a machine learning model studies to learn its patterns, kept apart from the examples used to test it. | 5 min |
| Transfer learning | Transfer learning reuses representations learned on a source task to improve learning on a different target task. | 3 min |
| Transformers | A transformer is a neural network that uses attention to decide which parts of a sequence matter to one another. | 3 min |
| Translation app | A translation application converts text or speech between languages while preserving meaning, terminology and formatting where possible. It remains subject to human review. | Quick |
| Tree of thoughts | Tree of Thoughts is a prompting and search approach that explores several candidate reasoning paths, evaluates them, and selects promising branches. | Quick |
| Triton Inference Server | NVIDIA Triton Inference Server serves models from multiple frameworks with batching, model management and HTTP or gRPC endpoints. | Quick |
| TRL | TRL is Hugging Face’s library for post-training language models with supervised fine-tuning, preference optimisation and reinforcement-learning methods. for practical reuse. | Quick |
| TruLens | TruLens is an open-source framework for evaluating and tracing language-model applications, especially retrieval quality, groundedness, and feedback signals. | Quick |
| TruthfulQA | TruthfulQA evaluates whether language models avoid producing common human misconceptions when answering questions designed to elicit plausible falsehoods. | Quick |
| Turing test | The Turing test asks whether a judge, chatting by text with a hidden person and a hidden machine, can reliably tell which one is the machine. | 5 min |
| txtai | txtai is an open-source Python framework for semantic search, embeddings databases, retrieval workflows, agents, and language-model pipelines over unstructured data. | Quick |
| U-Net | U-Net combines upsampled decoder features with matching high-resolution features from a contracting path. | 4 min |
| Udio | Udio is a generative music service for creating and extending songs from prompts, lyrics and uploaded audio within an editing workflow. | Quick |
| Ultralytics | Ultralytics provides computer-vision tooling best known for training, evaluating and deploying YOLO object-detection and segmentation models. in practice. | Quick |
| UMAP | UMAP reduces high-dimensional data using a graph of local relationships, often producing useful visualizations whose apparent spacing should not be treated as exact geometry. | Quick |
| Underfitting | Underfitting occurs when a model is too simple or insufficiently trained to capture important patterns, producing weak results even on training data. | Quick |
| Unsloth | Unsloth is an open-source tool for training and running AI language models on your own computer, built to fine-tune them faster and with less graphics-card memory. | 4 min |
| Unstructured | Unstructured partitions PDFs, office files, HTML and other documents into typed elements for downstream search, retrieval and AI pipelines. | Quick |
| Unsupervised learning | Unsupervised learning trains a model on examples that have no answers attached, so it finds structure on its own, such as groups of similar items or points that do not fit. | 5 min |
| V-JEPA | V-JEPA is Meta research on video representation learning by predicting abstract visual features rather than reconstructing every image pixel. | Quick |
| v0 | v0 is Vercel’s AI interface and application builder for generating web UI, code and deployable projects through conversation. | Quick |
| Validation set | A validation set is held-out data used during development to compare model choices without training on the same examples. | 3 min |
| Vanishing and exploding gradients | Vanishing and exploding gradients happen when the error signal shrinks toward zero or grows very large as it passes back through many layers, which makes deep networks hard to train. | Quick |
| Variational autoencoders | A variational autoencoder pairs a deep latent-variable model with a corresponding approximate inference model. | 4 min |
| Vector databases | A vector database stores vectors and returns the saved records closest to a query vector, often with a score and any extra fields you ask for. | 4 min |
| Vectors | A vector is an ordered list of numbers that you can also picture as an arrow in space. Machine learning stores examples, words and images this way. | 4 min |
| Veo | Veo is Google’s family of generative video models, producing video from text or image prompts with controls for style, composition, and motion. | Quick |
| Veo 3 | Veo 3 is a generation of Google’s video model family that creates prompted video and can generate synchronized audio alongside visuals. | Quick |
| Vercel AI SDK | The Vercel AI SDK is an open-source TypeScript toolkit, free to use, that lets apps talk to AI models from many companies through the same code. | 4 min |
| Vertex AI | Vertex AI is Google Cloud’s platform for choosing AI models, tuning them on your own examples and running models and agents for real users. | 5 min |
| Vespa | Vespa is a serving engine for large-scale search, recommendation and ranking that combines text, structured fields and vector retrieval. | Quick |
| Vibe coding | Vibe coding is an informal development style in which a person describes desired software behavior while an AI generates and revises much of the code. | Quick |
| Video summarizer | A video summarizer combines transcript and visual analysis to identify key moments, topics and a shorter account of the recording. | Quick |
| Virtual assistants | Virtual assistants help users complete tasks through conversational commands, often combining language understanding with search, applications, devices, or external services. | Quick |
| Vision transformers | A vision transformer represents an image as a sequence of patches and processes that sequence with a transformer. | 3 min |
| Vision-language models | A vision-language model can take visual data and text as input and produce text as output. | 4 min |
| Vision-language-action models | A vision-language-action model maps visual observations and a language instruction to robot actions. | 4 min |
| vLLM | vLLM is open-source software that runs large language models on a server and answers many requests at once, using GPU memory carefully so fewer bytes go to waste. | 4 min |
| Vocabulary | A model vocabulary supplies the mapping used to convert text into an ID sequence and back. | 3 min |
| Voice agents | A voice agent is an AI app you talk to out loud; it listens, works out what you want, can use tools, and answers in speech. | 5 min |
| Voice assistant | A voice assistant combines speech recognition, a reasoning or dialogue layer and speech synthesis for real-time spoken interaction. | Quick |
| Voice cloning | Voice cloning uses a speech model to produce new speech that sounds like a particular person, usually learned from a short recording of their voice. | Quick |
| Voyage embeddings | Voyage embeddings are commercial embedding models optimized for retrieval across general text, code, finance, law, and other specialized domains. | Quick |
| Wan | Wan is Alibaba’s generative video model family, designed to create and edit video from text, images, and other conditioning inputs. | Quick |
| Watermarking | AI watermarking embeds detectable signals in generated content or model outputs to help identify origin, though robustness and reliability vary by method. | Quick |
| Waymo | Waymo operates autonomous ride-hailing vehicles using cameras, lidar, radar, maps and driving software within selected service areas. in practical systems. | Quick |
| Weaviate | Weaviate is a vector database for semantic and hybrid search, with schema, filtering and optional integrated model capabilities. | Quick |
| Web scraping agent | A web scraping agent navigates selected sites, extracts defined fields and returns structured data while respecting access and rate limits. | Quick |
| WebGPU | WebGPU is a web standard that exposes modern GPU computation and graphics capabilities to browsers, including acceleration for local AI inference. | Quick |
| Weights | Trainable weights are model parameters that can be learned from data. | 3 min |
| Weights & Biases | Weights & Biases supports experiment tracking, model and application evaluation, artifact management and collaboration across machine-learning teams. in production. | Quick |
| Whisper | Whisper is OpenAI’s open-source speech-recognition model, trained for multilingual transcription, translation into English, and robust audio understanding across varied recordings. | Quick |
| Word2Vec | Word2Vec is a technique for learning vector representations of words from nearby context, making semantic relationships measurable through geometry. | Quick |
| Workflow automation with LLMs | Workflow automation with LLMs adds language understanding or generation to a deterministic process while keeping triggers, permissions and actions explicit. | Quick |
| World models | A world model predicts how an environment may change after an action, so an agent can evaluate possible futures before acting. | 4 min |
| xAI API | The xAI API gives developers hosted access to Grok models for text, vision, tool use and structured application workflows. | Quick |
| XGBoost | XGBoost is an open gradient-boosted tree library designed for efficient, accurate supervised learning on structured data. in practice. | Quick |
| YOLO | YOLO, or You Only Look Once, is a family of real-time computer-vision models that detect and classify objects in a single pass. | Quick |
| Zapier Agents | Zapier Agents lets users configure AI agents that work with Zapier-connected applications to research information and perform approved actions. | Quick |
| Zero-shot learning | Zero-shot learning applies a model to a task without task-specific examples, relying on prior training and a natural-language instruction or label description. | Quick |