223 terms
The glossary.
Every term defined in plain English, with an example. If a definition here needs another definition to make sense, we’ve failed and would like to know.
223 terms
Ablation Study
A research method where components of an AI system are systematically removed or modified to understand their individual contribution to overall performance. Ablation studies help researchers identify which design decisions matter most.
Accuracy
The proportion of all predictions that are correct, calculated as the number of correct predictions divided by the total number of predictions. While intuitive, accuracy can be misleading on imbalanced datasets.
Activation Function
A mathematical function applied to the output of each neuron that introduces non-linearity into the network. Without activation functions, a neural network would be equivalent to a linear model regardless of its depth.
Active Learning
A machine learning approach where the model selects the most informative unlabeled examples for human annotation, minimizing the amount of labeled data needed. Active learning reduces annotation costs by focusing human effort on the most useful examples.
Adam Optimizer
Adaptive Moment Estimation, a popular optimization algorithm that maintains per-parameter learning rates adapted based on first and second moments of gradients. Adam combines the benefits of AdaGrad and RMSProp and works well across a wide range of problems.
Adapter
Small trainable modules inserted between layers of a frozen pre-trained model for parameter-efficient fine-tuning. Adapters allow task-specific customization while keeping the vast majority of model parameters unchanged.
Adversarial Attack
Deliberately crafted inputs designed to cause an AI model to make incorrect predictions or produce undesirable outputs. Adversarial attacks exploit model vulnerabilities through subtle perturbations that are often imperceptible to humans.
Agentic AI
AI systems designed to act with agency, taking initiative to accomplish goals through planning, tool use, and iterative reasoning. Agentic AI represents a shift from passive response generation to proactive task completion.
Agentic Workflow
An AI-powered automation pipeline where one or more AI agents plan, execute, and iterate on tasks with minimal human intervention. Agentic workflows combine tool use, decision-making, and error handling to accomplish complex multi-step processes.
AI Agent
An AI system that can autonomously plan and execute multi-step tasks to achieve goals, using tools and making decisions based on observations. AI agents go beyond simple Q&A to take actions, iterate, and adapt their approach.
AI Alignment
The challenge of ensuring AI systems' goals and behaviors are aligned with human values and intentions. Alignment research aims to build AI that reliably does what its users want while avoiding unintended negative consequences.
AI Chip
Specialized semiconductor chips designed or optimized for AI workloads, including training and inference. AI chips include GPUs, TPUs, NPUs, and custom ASICs that offer higher performance and energy efficiency than general-purpose processors for AI tasks.
AI Ethics
The study of moral principles and values that should guide the development and deployment of AI systems. AI ethics addresses issues like bias, fairness, transparency, accountability, privacy, and the societal impact of AI technology.
AI Governance
The frameworks, policies, and institutions that manage the development, deployment, and use of AI technology. AI governance spans organizational policies, industry standards, and government regulations.
AI Safety
The research field focused on ensuring AI systems behave as intended and do not cause unintended harm. AI safety encompasses technical research on alignment, robustness, and interpretability, as well as policy and governance frameworks.
API
Application Programming Interface, a set of protocols and tools for building and integrating software applications. In AI, APIs provide programmatic access to model capabilities, allowing developers to send inputs and receive outputs over HTTP.
Attention Mechanism
A component in neural networks that allows the model to dynamically focus on the most relevant parts of the input when producing each part of the output. Attention computes weighted sums of input representations, where weights reflect relevance.
Attention Sink
A phenomenon where transformer models allocate disproportionate attention to the first few tokens of a sequence regardless of their semantic importance. Understanding attention sinks has led to optimizations for streaming and long-context inference.
AUC-ROC
Area Under the Receiver Operating Characteristic Curve, a metric that measures a classifier's ability to distinguish between classes across all classification thresholds. AUC-ROC ranges from 0.5 (random) to 1.0 (perfect separation).
Autoencoder
A neural network that learns to compress data into a lower-dimensional representation (encoding) and then reconstruct it (decoding). Autoencoders are used for dimensionality reduction, denoising, and feature learning.
AutoML
Automated Machine Learning, tools and techniques that automate the process of applying machine learning, including feature engineering, model selection, and hyperparameter tuning. AutoML makes ML accessible to non-experts and speeds up model development.
Autonomous Agent
An AI system capable of operating independently over extended periods, making decisions and taking actions without continuous human guidance. Autonomous agents can set sub-goals, recover from errors, and manage complex multi-step workflows.
Autoregressive Model
A model that generates output one element at a time, conditioning each prediction on all previously generated elements. Most language models are autoregressive, predicting the next token based on the sequence of tokens already produced.
Backpropagation
The algorithm used to compute gradients of the loss function with respect to each parameter in a neural network by propagating errors backward through the layers. Backpropagation enables efficient training of deep networks using gradient descent.
Batch Normalization
A technique that normalizes the inputs to each layer of a neural network across a mini-batch during training. Batch normalization stabilizes and accelerates training by reducing internal covariate shift, allowing higher learning rates.
Batch Processing
Running multiple AI inference requests together as a group to maximize hardware utilization and throughput. Batch processing is more efficient than processing individual requests sequentially, reducing per-request cost.
Beam Search
A search algorithm for text generation that maintains multiple candidate sequences (beams) simultaneously and selects the most probable complete sequence. Beam search produces more coherent text than greedy decoding but is slower.
Benchmark
A standardized test or dataset used to measure and compare the performance of AI models on specific tasks. Benchmarks provide objective metrics for tracking progress and comparing different approaches.
BERT
Bidirectional Encoder Representations from Transformers, a model developed by Google that reads text in both directions simultaneously. BERT revolutionized NLP by introducing effective pre-training for understanding tasks like classification, question answering, and named entity recognition.
Bias (Neural Network)
A learnable parameter added to the weighted sum of inputs in a neuron, allowing the model to shift the activation function. Bias terms provide additional flexibility, enabling the model to fit data that doesn't pass through the origin.
BLEU Score
Bilingual Evaluation Understudy, a metric for evaluating the quality of machine-translated text by comparing it to human reference translations. BLEU measures n-gram overlap between the generated and reference texts on a 0-100 scale.
BPE
Byte Pair Encoding, a subword tokenization algorithm that iteratively merges the most frequent pairs of characters or character sequences. BPE builds a vocabulary of subword units that efficiently represent both common words and rare terms.
Carbon Footprint AI
The environmental impact of training and running AI models, measured in carbon dioxide equivalent emissions. The carbon footprint includes energy consumed by GPUs during training and inference, as well as cooling and infrastructure.
Catastrophic Forgetting
The tendency of neural networks to completely lose previously learned information when trained on new data. Catastrophic forgetting is a major challenge in continual learning, where models need to acquire new knowledge without losing old skills.
Chain of Thought
A prompting technique that encourages language models to show their reasoning step by step before arriving at a final answer. Chain-of-thought prompting significantly improves performance on mathematical, logical, and multi-step reasoning tasks.
Chunking
The process of splitting large documents into smaller, semantically coherent segments for storage and retrieval in RAG systems. Effective chunking balances between preserving context and maintaining relevance for search queries.
Classification
A supervised learning task where the model assigns inputs to predefined categories or classes. Classification can be binary (two classes) or multi-class (many categories) and is used in spam detection, medical diagnosis, and sentiment analysis.
CLIP
Contrastive Language-Image Pre-training, a model by OpenAI that learns visual concepts from natural language descriptions. CLIP can match images and text in a shared embedding space, enabling zero-shot image classification without task-specific training data.
Clustering
An unsupervised learning task that groups similar data points together without predefined labels. Clustering algorithms discover natural groupings in data based on similarity metrics.
CNN
Convolutional Neural Network, an architecture that uses learnable filters to detect spatial patterns in data, primarily used for image processing. CNNs automatically learn hierarchical features from local patterns to complex structures.
Compute
The computational resources (processing power, memory, time) required to train and run AI models. Compute is measured in FLOPs and is a primary bottleneck and cost driver in AI development, with demand growing exponentially.
Computer Vision
A field of AI that enables computers to interpret and understand visual information from the world, including images and videos. Computer vision tasks include object detection, image classification, facial recognition, and scene understanding.
Confabulation
Another term for AI hallucination, referring to when models generate plausible-sounding but fabricated information with apparent confidence. Confabulation highlights the parallel with the psychological phenomenon where humans create false memories.
Constitutional AI
An approach developed by Anthropic where AI models are trained to follow a set of explicit principles (a constitution) that guide their behavior. The model critiques and revises its own outputs based on these principles during training.
Context Stuffing
The practice of filling a model's context window with relevant documents, data, or instructions to provide maximum information for generating a response. Context stuffing leverages large context windows to give models comprehensive reference material.
Context Window
The maximum number of tokens a language model can process in a single input-output exchange. A larger context window allows the model to consider more information when generating responses, enabling tasks like analyzing long documents.
Contrastive Learning
A self-supervised learning approach that trains models by pulling similar examples closer together and pushing dissimilar examples apart in embedding space. Contrastive learning creates powerful representations without labeled data.
ControlNet
A neural network architecture that adds spatial conditioning controls to diffusion models, allowing precise control over generated images. ControlNet can use edge maps, depth maps, poses, and other structural inputs to guide image generation.
Cross-Attention
An attention mechanism where queries come from one sequence while keys and values come from a different sequence. Cross-attention enables models to relate information across different inputs, such as an image and text in multimodal models.
Cross-Entropy
A loss function commonly used for classification tasks that measures the difference between predicted probability distributions and actual labels. Cross-entropy loss penalizes confident wrong predictions more heavily than uncertain ones.
Cross-Validation
A technique for evaluating model performance by splitting data into multiple subsets, training on some and testing on others in rotation. Cross-validation provides a more reliable estimate of model performance than a single train-test split.
Curriculum Learning
A training strategy that presents data to the model in a meaningful order, typically starting with easier examples and gradually increasing difficulty. Curriculum learning can improve training efficiency and final model performance.
Data Augmentation
Techniques for artificially expanding a training dataset by creating modified versions of existing data. Augmentation improves model robustness and generalization by introducing variation without collecting new data.
Data Leakage
When information from outside the training data inadvertently influences model evaluation, leading to overly optimistic performance estimates. Data leakage can occur when test data overlaps with training data or when future information is used in historical predictions.
Data Pipeline
An automated workflow for collecting, processing, transforming, and delivering data for model training or inference. Data pipelines ensure consistent data quality and enable reproducible experiments at scale.
Debiasing
Techniques for identifying and reducing unwanted biases in AI models and training data. Debiasing aims to ensure AI systems treat all demographic groups fairly and do not perpetuate or amplify societal prejudices.
Deep Learning
A subset of machine learning that uses neural networks with many layers (deep architectures) to learn hierarchical representations of data. Deep learning has driven breakthroughs in computer vision, NLP, and speech recognition.
Differential Privacy
A mathematical framework for providing provable privacy guarantees when analyzing or training on sensitive data. Differential privacy adds calibrated noise to data or computations, ensuring that no individual's data significantly affects the output.
Diffusion Model
A generative model that learns to create data by reversing a gradual noising process. During training, noise is progressively added to data; the model learns to denoise step by step. At generation time, it starts from pure noise and iteratively refines it into coherent output.
Dimensionality Reduction
Techniques for reducing the number of features in a dataset while preserving important information. Dimensionality reduction helps with visualization, noise removal, and improving model efficiency on high-dimensional data.
Distillation
A technique where a smaller student model is trained to replicate the behavior of a larger teacher model. Distillation transfers the teacher's knowledge into a compact model that is faster and cheaper to run while retaining much of the performance.
DPO
Direct Preference Optimization, a simpler alternative to RLHF that directly optimizes a language model using preference data without needing a separate reward model. DPO achieves comparable results to RLHF with less computational complexity and greater training stability.
Dropout
A regularization technique where random neurons are temporarily disabled during training, forcing the network to learn redundant representations. This prevents the network from becoming overly dependent on any single neuron and improves generalization.
Edge AI
Running AI models directly on local devices (phones, IoT devices, vehicles) rather than in the cloud. Edge AI reduces latency, improves privacy, and enables offline operation, but is constrained by the device's limited compute and memory.
Embedding
A dense numerical vector representation of data (text, images, audio) in a high-dimensional space where semantically similar items are positioned close together. Embeddings capture meaning and relationships, enabling similarity search and clustering.
Embodied AI
AI systems that interact with the physical world through robotic bodies or physical interfaces. Embodied AI combines perception, reasoning, and physical action to perform real-world tasks like manipulation, navigation, and assembly.
Emergent Abilities
Capabilities that appear in larger models but are absent in smaller ones, seemingly arising spontaneously at certain scale thresholds. Emergent abilities include tasks like multi-step reasoning, arithmetic, and code generation that smaller models cannot perform.
Encoder-Decoder
A neural network architecture with two main components: an encoder that processes and compresses the input into a representation, and a decoder that generates output from that representation. This pattern is fundamental to translation, summarization, and other transformation tasks.
Epoch
One complete pass through the entire training dataset during model training. Models typically require multiple epochs to converge, with each epoch refining the learned patterns. Too few epochs lead to underfitting, while too many can cause overfitting.
EU AI Act
The European Union's comprehensive regulatory framework for artificial intelligence, categorizing AI systems by risk level and imposing requirements accordingly. The EU AI Act is the world's first major AI-specific regulation.
Eval
Short for evaluation, the systematic process of measuring AI model performance across various dimensions including accuracy, safety, speed, and user satisfaction. Evals are critical for model development, comparison, and deployment decisions.
Explainable AI
AI systems and techniques designed to make model decisions understandable to humans. Explainable AI helps users, developers, and regulators understand why a model made a particular prediction or decision.
Extended Thinking
A capability where AI models generate internal reasoning chains before producing a final response, spending more compute on harder problems. Extended thinking enables models to tackle complex multi-step problems that quick responses would get wrong.
F1 Score
The harmonic mean of precision and recall, providing a single metric that balances both measures. F1 score ranges from 0 to 1 and is particularly useful for evaluating models on imbalanced datasets where accuracy can be misleading.
Feature Engineering
The process of selecting, transforming, and creating input variables (features) from raw data to improve model performance. Good feature engineering can significantly boost model accuracy and is often more impactful than model architecture changes.
Federated Learning
A training approach where models learn from data distributed across many devices without the data ever leaving those devices. Each device trains a local model, and only the model updates are shared, preserving data privacy.
Few-shot Learning
The ability of a model to perform a task after being shown only a few examples in the prompt. Few-shot learning leverages the model's pre-trained knowledge to generalize from minimal demonstrations without additional training.
Fine-tuning
The process of further training a pre-trained model on a specific dataset to specialize it for a particular task or domain. Fine-tuning adjusts the model's weights using task-specific data, typically requiring far less data and compute than training from scratch.
Fine-tuning vs RAG
A common architectural decision in AI application design. Fine-tuning modifies model weights for specialized behavior, while RAG retrieves external information at inference time. The choice depends on whether the need is for behavioral change (fine-tuning) or knowledge augmentation (RAG).
Flash Attention
An optimized attention computation algorithm that reduces memory usage from quadratic to near-linear by processing attention in blocks. Flash Attention uses tiling and kernel fusion to minimize GPU memory reads and writes, significantly speeding up training and inference.
FLOPS
Floating Point Operations Per Second, a measure of computing performance. In AI, total FLOPS measures the compute used during training (e.g., 10^25 FLOPS for GPT-4), while FLOPS/s measures hardware throughput.
Foundation Model
A large-scale AI model trained on broad data that can be adapted to a wide range of downstream tasks. Foundation models serve as a base that can be fine-tuned or prompted for specific applications without training from scratch.
Frontier Model
The most capable and advanced AI models at the cutting edge of performance, typically requiring enormous compute resources to train. Frontier models push the boundaries of what AI can achieve and often exhibit emergent capabilities not seen in smaller models.
Function Calling
A capability where language models can identify when to use external tools and generate structured function calls with appropriate arguments. Function calling enables models to interact with databases, APIs, and other systems to accomplish real-world tasks.
GAN
Generative Adversarial Network, an architecture consisting of two neural networks: a generator that creates synthetic data and a discriminator that tries to distinguish real from generated data. The two networks train adversarially, pushing each other to improve.
Generative AI
AI systems capable of creating new content including text, images, audio, video, and code. Generative AI models learn the underlying patterns and structure of training data, then produce novel outputs that follow similar patterns.
GGUF
A file format for storing quantized machine learning models, designed for efficient local inference using the llama.cpp library. GGUF replaced the older GGML format and supports metadata, multiple tensor types, and various quantization levels.
GPT
Generative Pre-trained Transformer, a family of autoregressive language models developed by OpenAI. GPT models are pre-trained on large corpora of text using unsupervised learning, then fine-tuned for specific tasks. The architecture uses decoder-only transformers.
GPT Architecture
A decoder-only transformer architecture that generates text autoregressively, predicting one token at a time from left to right. The GPT architecture uses masked self-attention to prevent the model from looking ahead, and scales effectively to billions of parameters.
GPU
Graphics Processing Unit, specialized processors originally designed for rendering graphics that excel at the parallel matrix computations required for deep learning. GPUs are the primary hardware for training and running AI models.
Gradient Descent
An optimization algorithm that iteratively adjusts model parameters in the direction that most reduces the loss function. Gradient descent computes the gradient (slope) of the loss with respect to each parameter and updates them proportionally.
Grounding
Techniques for connecting AI model outputs to verifiable, real-world information sources to reduce hallucinations. Grounding ensures model responses are anchored in factual data rather than relying solely on patterns learned during training.
Guardrails
Safety mechanisms built into AI systems to prevent harmful, inappropriate, or off-topic outputs. Guardrails include content filters, output validators, topic restrictions, and behavioral constraints that keep the model within desired boundaries.
Hallucination
When an AI model generates confident but factually incorrect, fabricated, or nonsensical information. Hallucinations occur because language models produce statistically plausible text rather than verified facts, and they cannot distinguish between what they know and what they are generating.
HellaSwag
A benchmark for evaluating commonsense reasoning by asking models to select the most plausible continuation of a scenario. HellaSwag questions are designed to be easy for humans but challenging for AI models.
HumanEval
A benchmark for evaluating AI code generation consisting of 164 Python programming problems with unit tests. HumanEval measures a model's ability to generate functionally correct code from docstrings.
Hybrid Search
A search approach that combines traditional keyword-based search with semantic vector search to leverage the strengths of both methods. Hybrid search provides better recall and precision than either method alone.
Hyperparameter Tuning
The process of optimizing the configuration settings that govern the training process, such as learning rate, batch size, and number of layers. Unlike model parameters, hyperparameters are set before training and significantly impact model performance.
Image Classification
A computer vision task that assigns one or more labels to an entire image based on its content. Image classification is one of the foundational tasks in computer vision and has seen dramatic improvement with deep learning.
In-Context Learning
The ability of language models to learn new tasks from examples provided within the prompt, without any parameter updates. The model adapts its behavior based on the demonstrations given in the conversation context.
Inference
The process of using a trained model to make predictions or generate outputs on new, unseen data. Inference is the production use of a model, as opposed to training, and its speed and cost are critical factors for deployment.
Inference Cost
The computational expense of running a trained model to generate predictions or outputs, typically measured in dollars per million tokens. Inference cost depends on model size, hardware, and optimization techniques, and is a major factor in AI deployment economics.
Instruction Tuning
A fine-tuning process where a model is trained on a dataset of instruction-response pairs to improve its ability to follow human instructions. Instruction tuning transforms a base language model into a more helpful and controllable assistant.
Interpretability
The degree to which humans can understand and reason about an AI model's internal workings and decision-making process. Interpretability research develops tools for understanding what models learn, how they process information, and why they produce specific outputs.
Jailbreaking
Techniques used to circumvent an AI model's safety restrictions and content policies to produce outputs the model was designed to refuse. Jailbreaking exploits weaknesses in the model's training through creative prompting strategies.
Knowledge Distillation
A specific form of model compression where a smaller student model is trained to match the output probability distribution of a larger teacher model. Knowledge distillation transfers the teacher's dark knowledge, including information about incorrect class probabilities.
Knowledge Graph
A structured representation of real-world entities and their relationships, stored as a network of nodes and edges. Knowledge graphs help AI systems reason about connections between concepts and provide structured context for generation.
KV Cache
Key-Value Cache, an optimization that stores previously computed attention key and value tensors to avoid redundant computation during autoregressive generation. KV caching significantly speeds up inference but increases memory usage proportional to sequence length.
Large Language Model
A neural network trained on massive text datasets to understand and generate human language. LLMs use billions of parameters to capture statistical patterns in language, enabling them to perform tasks like translation, summarization, and code generation.
Latency
The time delay between sending a request to an AI model and receiving the first response token. Low latency is critical for interactive applications like chatbots and real-time translation, measured in milliseconds.
Layer Normalization
A normalization technique that normalizes the activations across features within a single training example, independent of other examples in the batch. Layer norm is the standard normalization in transformers, providing stable training dynamics.
Leaderboard
A ranking system that compares AI model performance on standardized benchmarks. Leaderboards like the Open LLM Leaderboard track how different models score across multiple evaluation metrics.
Learning Rate
A hyperparameter that controls the step size of parameter updates during training. A higher learning rate enables faster training but risks overshooting optimal values, while a lower rate is more precise but slower to converge.
LoRA
Low-Rank Adaptation, a parameter-efficient fine-tuning method that adds small trainable matrices to frozen model weights. LoRA drastically reduces the number of trainable parameters and memory needed for fine-tuning, making it practical to customize large models.
Loss Function
A mathematical function that measures the difference between a model's predictions and the actual target values. The loss function provides the signal that guides training by quantifying how wrong the model is on each example.
LSTM
Long Short-Term Memory, a recurrent neural network architecture designed to learn long-range dependencies in sequential data. LSTMs use gating mechanisms to selectively remember or forget information, addressing the vanishing gradient problem of simple RNNs.
Machine Learning
A branch of artificial intelligence where systems learn patterns from data rather than being explicitly programmed. ML algorithms improve their performance on tasks through experience, using statistical methods to find patterns in training data.
Mamba
A selective state space model architecture that offers an alternative to transformers for sequence modeling. Mamba achieves linear-time inference scaling with sequence length while matching transformer performance on many tasks.
Memory (AI)
Mechanisms that allow AI systems to retain and recall information across interactions, enabling personalization and continuity. AI memory can be short-term (within a conversation) or long-term (persisting across sessions).
Mixture of Agents
A system architecture where multiple LLMs collaborate by generating and refining responses iteratively. Each agent may have different strengths, and their outputs are aggregated or refined to produce a higher-quality final response.
Mixture of Depths
An architecture optimization where not all tokens are processed by every transformer layer. Mixture of Depths routes tokens through different numbers of layers based on their complexity, reducing compute for simpler tokens.
Mixture of Experts
An architecture where multiple specialized sub-networks (experts) exist within a model, with a routing mechanism that activates only a subset of experts for each input. This allows models to have enormous total parameter counts while only using a fraction of compute per inference.
MMLU
Massive Multitask Language Understanding, a benchmark testing AI knowledge across 57 academic subjects from elementary to professional level. MMLU evaluates a model's breadth and depth of knowledge in areas like medicine, law, mathematics, and history.
Model Card
A documentation framework for AI models that describes their intended use, performance characteristics, limitations, and ethical considerations. Model cards promote transparency by providing standardized information about what a model can and cannot do.
Model Context Protocol
An open standard developed by Anthropic for connecting AI models to external data sources and tools. MCP provides a universal protocol for AI systems to access context from various sources through standardized server connections.
Model Merging
Techniques for combining the weights of multiple trained models into a single model that inherits capabilities from all source models. Model merging enables creating new models without additional training, by averaging or interpolating parameters.
Multi-Agent System
A system where multiple AI agents work together, each with specialized roles or capabilities, coordinating to solve complex problems. Multi-agent systems can tackle tasks too complex for a single agent by dividing work and sharing information.
Multi-Head Attention
An extension of attention where multiple attention operations run in parallel, each with different learned projections. This allows the model to simultaneously capture different types of relationships (syntactic, semantic, positional) from the input.
Multimodal AI
AI systems that can process and generate multiple types of data, such as text, images, audio, and video, within a single model. Multimodal models can understand relationships across different data types and translate between modalities.
NeRF
Neural Radiance Fields, a technique that uses neural networks to reconstruct 3D scenes from a collection of 2D images. NeRFs create photorealistic 3D representations that can be rendered from any viewpoint.
Neural Network
A computing system inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers. Neural networks learn patterns from data by adjusting connection weights through training, enabling them to make predictions and decisions.
Neural Processing Unit
A specialized hardware accelerator designed specifically for running neural network computations efficiently on edge devices. NPUs enable on-device AI processing in smartphones, laptops, and IoT devices with low power consumption.
NLP
Natural Language Processing, a field of AI focused on enabling computers to understand, interpret, and generate human language. NLP encompasses tasks like translation, sentiment analysis, named entity recognition, and question answering.
Normalization
Techniques for scaling and shifting data or intermediate representations to improve training stability and speed. Common forms include batch normalization, layer normalization, and RMS normalization, each operating across different dimensions.
NVIDIA CUDA
A parallel computing platform and programming model by NVIDIA that allows developers to use GPUs for general-purpose computing. CUDA is the dominant software ecosystem for AI training and inference, giving NVIDIA a strong competitive moat.
Object Detection
A computer vision task that identifies and locates objects in images or video by drawing bounding boxes around them and classifying what they are. Object detection combines classification with spatial localization.
OCR
Optical Character Recognition, technology that converts images of text (from scanned documents, photos, or handwriting) into machine-readable text. Modern OCR uses deep learning to achieve high accuracy even on distorted or handwritten text.
ONNX
Open Neural Network Exchange, an open format for representing machine learning models that enables interoperability between different ML frameworks. ONNX allows models trained in PyTorch to run efficiently in TensorFlow or custom inference engines.
Open Source AI
AI models and tools released with publicly available weights, code, and documentation that anyone can use, modify, and distribute. Open source AI democratizes access to powerful models and enables community-driven improvement and research.
Optimizer
An algorithm that determines how model parameters are updated during training based on the computed gradients. Optimizers implement strategies beyond basic gradient descent to improve convergence speed and stability.
Overfitting
When a model learns the training data too well, including its noise and outliers, resulting in poor performance on new, unseen data. Overfitting occurs when the model memorizes specific examples rather than learning generalizable patterns.
Parameter Count
The total number of learnable values (weights and biases) in a neural network. Parameter count is often used as a rough proxy for model capability, though architecture and training data quality also matter significantly.
PCA
Principal Component Analysis, a statistical technique that identifies the directions of maximum variance in high-dimensional data and projects it onto a lower-dimensional space. PCA is one of the most widely used dimensionality reduction methods.
PEFT
Parameter-Efficient Fine-Tuning, a family of techniques that adapt large models by training only a small subset of parameters. PEFT methods include LoRA, adapters, and prefix tuning, making fine-tuning accessible with limited compute resources.
Perplexity (metric)
A measure of how well a language model predicts a sample of text, calculated as the exponentiated average negative log-likelihood. Lower perplexity indicates the model is less surprised by the text, suggesting better language understanding.
Positional Encoding
A method for injecting information about token position into transformer models, which otherwise have no inherent sense of order. Positional encodings can be fixed sinusoidal functions, learned embeddings, or relative position representations like RoPE.
PPO
Proximal Policy Optimization, a reinforcement learning algorithm widely used in RLHF training of language models. PPO updates the policy in small, controlled steps to improve stability, preventing the model from changing too dramatically in any single update.
Precision (ML)
The proportion of positive predictions that are actually correct, measuring how many of the model's positive identifications were right. High precision means few false positives.
Prefix Tuning
A parameter-efficient fine-tuning method that prepends trainable continuous vectors (prefixes) to the input of each transformer layer. Only the prefix parameters are updated during training, leaving the original model weights frozen.
Prompt Caching
An optimization technique that stores and reuses computed key-value pairs from previously processed prompt prefixes. Prompt caching avoids redundant computation when the same system prompt or context is used across multiple requests.
Prompt Engineering
The practice of crafting and optimizing input prompts to elicit desired outputs from language models. Effective prompt engineering involves structuring instructions, providing examples, and using techniques like chain-of-thought reasoning.
Prompt Injection
An attack where malicious instructions are inserted into an AI system's input to override its original instructions or manipulate its behavior. Prompt injection can occur directly through user input or indirectly through data the model processes.
QLoRA
Quantized Low-Rank Adaptation, combining 4-bit quantization with LoRA to enable fine-tuning of large models on a single consumer GPU. QLoRA loads the base model in 4-bit precision while training LoRA adapters in higher precision.
Quantization
A technique for reducing model size and inference cost by using lower-precision number formats for weights and activations. Quantization can shrink a model from 16-bit to 4-bit precision with minimal quality loss, enabling deployment on consumer hardware.
RAG
Retrieval-Augmented Generation, a technique that enhances language model outputs by retrieving relevant information from external knowledge sources before generating a response. RAG reduces hallucinations and keeps responses current without retraining the model.
Reasoning Model
AI models specifically designed or trained to perform step-by-step logical reasoning before answering. Reasoning models use techniques like chain-of-thought and extended thinking to solve complex problems more reliably.
Recall (ML)
The proportion of actual positive cases that the model correctly identifies, measuring how many real positives the model catches. High recall means few false negatives.
Red Teaming
The practice of deliberately attempting to find vulnerabilities, biases, and failure modes in AI systems by adopting an adversarial perspective. Red teams test models with edge cases, manipulative prompts, and creative attacks to identify weaknesses before deployment.
Regression
A supervised learning task where the model predicts a continuous numerical value rather than a category. Regression is used for forecasting, pricing, and any prediction involving quantities.
Regularization
Techniques used to prevent overfitting by adding constraints or penalties to the training process. Regularization encourages the model to learn simpler, more generalizable patterns rather than memorizing training data.
Reinforcement Learning
A machine learning paradigm where an agent learns to make decisions by taking actions in an environment and receiving rewards or penalties. The agent optimizes its behavior to maximize cumulative reward over time through trial and error.
Reinforcement Learning from Human Feedback
A training paradigm that uses human preferences to train AI systems to produce outputs humans find more helpful, accurate, and safe. RLHF involves collecting comparison data from human raters, training a reward model, then optimizing the AI policy using RL.
ReLU
Rectified Linear Unit, the most widely used activation function in deep learning that outputs the input directly if positive, or zero if negative. ReLU is computationally efficient and helps mitigate the vanishing gradient problem in deep networks.
Reranking
A technique that takes an initial set of retrieved results and reorders them using a more sophisticated model to improve relevance. Reranking is used as a second stage in retrieval pipelines after fast but less precise initial retrieval.
Residual Connection
A skip connection that adds the input of a layer directly to its output, allowing gradients to flow through the network without degradation. Residual connections enable training of very deep networks and are a key component of transformers.
Responsible AI
The practice of designing, developing, and deploying AI systems in a manner that is ethical, transparent, fair, and accountable. Responsible AI frameworks help organizations identify and mitigate risks throughout the AI lifecycle.
Retrieval
The process of finding and fetching relevant information from a knowledge base or document store to provide context for AI model responses. Retrieval systems use techniques like semantic search and keyword matching to find pertinent data.
Retrieval-Augmented Generation
The full name for RAG, a technique that augments language model generation with information retrieved from external knowledge sources. RAG systems typically use embedding-based retrieval to find relevant documents before generating a response.
Reward Model
A model trained to predict human preferences, used to provide feedback signals for reinforcement learning. Reward models learn from human comparison data to score outputs, enabling automated training without continuous human involvement.
RLHF
Reinforcement Learning from Human Feedback, a training technique where human evaluators rank model outputs to train a reward model, which then guides the language model to generate more helpful, harmless, and honest responses through reinforcement learning.
RoPE
Rotary Position Embedding, a method for encoding token position information in transformers using rotation matrices. RoPE naturally encodes relative positions, enabling better length generalization and is used in most modern LLMs including Llama and Mistral.
SAM
Segment Anything Model, developed by Meta AI, is a foundation model for image segmentation that can identify and separate objects in any image with minimal prompting. SAM was trained on 11 million images with over 1 billion masks.
Scaling Laws
Empirical relationships showing that model performance improves predictably with increases in model size, dataset size, and compute. Scaling laws help researchers predict the capabilities of larger models before training them, guiding resource allocation.
SDK
Software Development Kit, a collection of tools, libraries, and documentation that helps developers integrate AI capabilities into their applications. SDKs provide language-specific wrappers around APIs, simplifying common tasks.
Self-Attention
A mechanism where each element in a sequence attends to all other elements in the same sequence to compute a representation. Self-attention allows the model to capture relationships between words regardless of their distance in the text.
Self-Supervised Learning
A learning paradigm where the model generates its own labels from unlabeled data, typically by predicting missing or corrupted parts of the input. Self-supervised learning enables pre-training on massive unlabeled datasets before fine-tuning on smaller labeled datasets.
Semantic Search
Search that understands the meaning and intent behind queries rather than just matching keywords. Semantic search uses embeddings to find conceptually similar content, even when different words are used to express the same idea.
Semantic Segmentation
A computer vision task that assigns a class label to every pixel in an image, creating a detailed map of what each region represents. Semantic segmentation provides more granular understanding than object detection.
SentencePiece
An unsupervised text tokenizer and detokenizer by Google that treats the input as a raw stream of characters, making it language-agnostic. SentencePiece can use BPE or unigram algorithms and is used by many multilingual models.
Seq2Seq
Sequence-to-Sequence, an architecture that maps an input sequence to an output sequence of potentially different length. Seq2Seq models typically use an encoder to process the input and a decoder to generate the output.
Sigmoid
An activation function that maps any input to a value between 0 and 1, following an S-shaped curve. Sigmoid is commonly used for binary classification outputs and gating mechanisms in neural networks like LSTMs.
Softmax
An activation function that converts a vector of raw scores into a probability distribution, where all values sum to 1. Softmax is typically used in the final layer of classification models and in attention mechanisms.
Sparse Attention
An attention variant where each token attends to only a subset of other tokens rather than all tokens in the sequence. Sparse attention reduces the quadratic computational cost of standard attention, enabling processing of much longer sequences.
Speculative Decoding
An inference optimization where a smaller draft model quickly generates candidate tokens, and the larger target model verifies them in parallel. This can speed up generation by 2-3x since verification is faster than sequential generation.
Speech-to-Text
AI technology that converts spoken audio into written text, also known as automatic speech recognition (ASR). Modern speech-to-text systems achieve near-human accuracy across many languages and accents.
Stable Diffusion
An open-source text-to-image diffusion model developed by Stability AI that generates detailed images from text prompts. It operates in a compressed latent space for efficiency and has spawned a large ecosystem of community extensions and fine-tuned variants.
Structured Output
The ability of language models to generate outputs in specific formats like JSON, XML, or predefined schemas. Structured output ensures model responses can be reliably parsed and integrated into software systems.
Supervised Learning
A machine learning approach where models learn from labeled training data, mapping inputs to known correct outputs. The model adjusts its parameters to minimize the difference between its predictions and the true labels.
Sycophancy
A behavior pattern where AI models agree with users even when the user is incorrect, prioritizing user satisfaction over accuracy. Sycophancy results from RLHF training where models learn that agreeable responses receive higher ratings.
Synthetic Data
Artificially generated data that mimics the statistical properties of real data, used for training and testing AI models. Synthetic data can address privacy concerns, data scarcity, and bias while reducing the cost of data collection.
System Prompt
A special instruction given to a language model that sets its behavior, role, and constraints for a conversation. System prompts establish the model's persona, capabilities, and limitations before user interaction begins.
T5
Text-to-Text Transfer Transformer, a model by Google that frames all NLP tasks as text-to-text problems. T5 converts inputs and outputs into text strings, allowing a single model architecture to handle translation, summarization, classification, and more.
Temperature
A parameter that controls the randomness of a language model's output by scaling the probability distribution over tokens. Lower temperature (e.g., 0.1) makes outputs more deterministic and focused, while higher temperature (e.g., 1.0) increases creativity and diversity.
Tensor Core
Specialized processing units within NVIDIA GPUs designed to accelerate matrix multiply-and-accumulate operations, the core computation in deep learning. Tensor Cores provide massive speedups for AI training and inference workloads.
Text-to-3D
AI technology that generates three-dimensional models and scenes from text descriptions. Text-to-3D enables rapid prototyping and content creation for games, VR, architecture, and product design.
Text-to-Image
AI technology that generates images from natural language descriptions, translating text prompts into visual content. Text-to-image models have advanced rapidly from simple compositions to photorealistic and artistically sophisticated outputs.
Text-to-Speech
AI technology that converts written text into natural-sounding spoken audio. Modern TTS systems use neural networks to generate speech that is nearly indistinguishable from human voice, with control over emotion, pace, and style.
Text-to-Video
AI technology that generates video content from text descriptions, creating moving scenes with characters, actions, and environments. Text-to-video represents one of the most complex generative AI challenges due to temporal coherence requirements.
Throughput
The number of tokens or requests a model can process per unit of time, measuring the system's overall processing capacity. High throughput is important for serving many users simultaneously and batch processing large datasets.
Token
The basic unit of text that language models process. Tokens can be words, parts of words, or individual characters depending on the tokenizer used. Most models process text by converting it into a sequence of token IDs.
Token Limit
The maximum number of tokens allowed in a single API request or response, set by the model provider. Token limits constrain both input context and output length, and exceeding them results in truncation or errors.
Tokenizer
A component that converts raw text into a sequence of tokens that a language model can process. Tokenizers split text into subword units using learned vocabularies, balancing between character-level and word-level representations.
Tokenomics
The pricing and economic considerations around token usage in AI APIs. Token costs vary by model size and provider, with input tokens typically cheaper than output tokens, and pricing directly impacts the viability of AI applications.
Tool Use
The ability of AI models to interact with external tools, APIs, and systems to accomplish tasks beyond pure text generation. Tool use extends model capabilities to include web search, code execution, file manipulation, and database queries.
Top-P Sampling
A text generation strategy that samples from the smallest set of tokens whose cumulative probability exceeds a threshold P. Also called nucleus sampling, top-P dynamically adjusts the number of candidate tokens based on the probability distribution.
TPU
Tensor Processing Unit, custom AI accelerator chips designed by Google specifically for machine learning workloads. TPUs are optimized for tensor operations and are used to train and serve Google's AI models including Gemini.
Training Cost
The total computational expense of training an AI model, including GPU hours, electricity, and cooling. Training frontier models can cost tens to hundreds of millions of dollars and require thousands of GPUs running for months.
Transfer Learning
A technique where knowledge gained from training on one task is applied to a different but related task. Transfer learning enables models to achieve strong performance on new tasks with limited data by leveraging patterns learned from large-scale pre-training.
Transformer
A neural network architecture introduced in the 2017 paper 'Attention Is All You Need' that processes sequential data using self-attention mechanisms instead of recurrence. Transformers enable massive parallelization during training, making them the foundation of modern LLMs and vision models.
Transformer XL
An extension of the transformer architecture that introduces recurrence to capture longer-range dependencies. Transformer XL uses segment-level recurrence with a novel positional encoding scheme, allowing it to learn dependencies beyond a fixed context length.
TruthfulQA
A benchmark that measures whether AI models generate truthful answers to questions that commonly elicit misconceptions or falsehoods in humans. TruthfulQA tests the model's resistance to repeating popular but incorrect beliefs.
Underfitting
When a model is too simple to capture the underlying patterns in the data, resulting in poor performance on both training and test data. Underfitting indicates the model lacks sufficient capacity or has not been trained long enough.
Unsupervised Learning
A machine learning approach where models find patterns and structure in data without labeled examples. The model discovers hidden relationships, clusters, or representations on its own from the raw input data.
VAE
Variational Autoencoder, a generative model that learns a compressed latent representation of data and can generate new samples by decoding points from this latent space. VAEs combine autoencoders with probabilistic modeling to enable controlled generation.
Vector Database
A specialized database designed to store, index, and query high-dimensional vector embeddings efficiently. Vector databases enable fast similarity search over millions of embeddings, powering RAG systems, recommendation engines, and semantic search.
Vision Transformer
A transformer architecture adapted for image processing that divides images into patches and processes them as a sequence, similar to how text transformers process tokens. ViTs have matched or exceeded CNNs on many computer vision benchmarks.
Voice Cloning
AI technology that replicates a specific person's voice from audio samples, enabling generation of new speech in that voice. Voice cloning raises important ethical and security considerations around consent and deepfakes.
Watermarking (AI)
Techniques for embedding imperceptible markers in AI-generated content to enable detection and attribution. AI watermarking helps distinguish AI-generated text, images, and audio from human-created content.
Weight Decay
A regularization technique that adds a penalty proportional to the magnitude of model weights to the loss function, encouraging smaller weight values. Weight decay helps prevent overfitting by limiting the model's capacity to memorize training data.
Weights
The numerical values in a neural network that are adjusted during training to minimize prediction errors. Weights determine how strongly each neuron connection influences the next layer, and a trained model's weights encode all its learned knowledge.
Whisper
An automatic speech recognition model by OpenAI trained on 680,000 hours of multilingual audio data. Whisper can transcribe speech in multiple languages and translate non-English speech into English with high accuracy.
World Model
An AI model that learns an internal representation of how the world works, enabling it to predict outcomes, simulate scenarios, and plan actions. World models are considered a key step toward more general artificial intelligence.
Zero-shot Learning
The ability of a model to perform a task without any task-specific examples, relying entirely on its pre-trained knowledge and the task description. Zero-shot performance has improved dramatically with larger, more capable models.