AI SYSTEMS ENGINEERING · LEARNING NOTES
Understand the systems
that run models.
Models, hardware and workloads.
Build understanding step by step, from shared concepts to inference, training and RL.
Four stages, starting with small examples
Shared Concepts
From model computation to GPU execution
Inference
Serving requests efficiently with fixed weights
Training
Pretraining and SFT, from one GPU to multiple nodes
RL
Connecting generation, evaluation and learning
This site groups pretraining and SFT under training, with RL as a separate learning stage. In common terminology, SFT is also part of post-training.
Available articles
Shared Concepts
Models
2026-09-09The Structure of an LLM: From the Embedding Layer to the LM Head
Follow text through token IDs, vectors, and vocabulary scores, and explore the roles of the embedding layer, decoder blocks, and LM Head.
2026-09-09The Flow Through a Decoder Block: Residual Connections and RMSNorm
Explore a decoder block around its residual stream, then follow vector addition and RMSNorm across each token’s d components.
2026-09-09Attention and MLP: Information Across Tokens and Transformations Within a Token
Compare Attention and MLP, then explore matrix multiplication and the computations inside basic and gated MLPs.
2026-09-09Attention Projections: Q, K, V and Multiple Heads
Explore Q, K and V projections, the organization of multiple heads, and how the output projection combines head results for the same token.
2026-09-09Core Attention: Combining Information Across Tokens
Follow Core Attention within each head through QKᵀ, scaling, a causal mask, Softmax, and a weighted sum of Value vectors.
2026-09-10RoPE: Incorporating Token Positions into Attention
Explore how RoPE rotates the components of Q and K by token position and how relative positions affect Attention scores.
2026-09-10MoE: Choosing Which MLPs to Use for Each Token
Compare dense MLPs and MoE, then follow routing, dispatch, expert computation, combination, and the shared expert path.
2026-09-10Revisiting the Flow Through the Model
Connect the components from Embedding to the LM Head, review Attention and MLP or MoE, and identify where information is combined across tokens.
2026-09-27MQA and GQA: Sharing KV Across Queries
Use MHA as a baseline to compare KV sharing in MQA and GQA, and examine how they reduce storage while preserving per-query outputs and what trade-offs follow.
2026-09-27MLA storage: Representing KV with a small latent vector
Explore how a joint latent vector represents head-specific keys and values, why a positional key is stored separately, what is shared across heads, and how KV cache storage grows.
2026-09-27MLA Computation: Attention Without Expanding KV
Follow how the current query reads a latent KV cache, and see why moving the key projection to the query and the value projection after the weighted sum preserves the result.
2026-09-27Local and Sparse Attention: Reading Fewer Token Positions
Start with sliding window attention as a form of sparse attention, trace fixed connection patterns and information flow through layers, and distinguish the read range from KV cache storage.
2026-09-27Sparse Attention Indexers: Choosing Positions by Content
Trace how an index query and keys select KV positions, how selection changes during generation, and how a sliding window combines with selection, while distinguishing reading cost from cache storage.
2026-09-27Token-Axis Compression: Reading Summaries of Multiple KV Positions
Trace how content and score projections build summary vectors, why completed summaries are read alongside recent per-token KV, and how compression combines with an indexer.
2026-09-27Linear Attention: Accumulating KV in a Fixed-Size State
Understand why softmax is replaced with independent feature maps, then follow outer-product writes, Query reads, and normalization by the sum of scores.
2026-09-28The Delta Rule: Revising State Associations Toward a New Value
Start from the purpose of writing KV associations, compare additive accumulation with the delta rule, and distinguish checking the old value with a Key from reading an output with a Query.
2026-09-28GDN and KDA: Retaining State and Applying Delta Corrections
Distinguish retention of the old state from correction toward a new Value, then follow GDN updates through KDA’s component-wise retention and input-dependent control values.
2026-09-28SSMs: Carrying Context Through State and New Inputs
Compare state updates in Linear Attention and SSMs, then use small numerical examples to trace retention, writing, reading, and the lingering influence of past inputs.
2026-09-28Mamba: Choosing What to Remember Based on the Input
Trace selective SSM coefficients for retention, writing, and reading, locate that computation within a Mamba block, and compare it with delta corrections in GDN and KDA.
2026-09-28Hybrid Models: Using State and Attention Together
Explore why models combine fixed-size state with position-indexed KV, the quality and storage trade-offs, and how one token passes through both types of layers.
2026-09-28HC and mHC: Widening and Connecting Residual Streams
Trace reading, bypass mixing, and writing from the same input, then examine why mHC constrains repeated mixing and what the wider residual path costs.
2026-09-28Gated Residual: Reading Components and Writing to Originals
Follow one numerical example through branch normalization, componentwise read gates, and branchwise scalar writes in Gated Residual.
2026-09-28Attention Residuals: Selecting Outputs from Earlier Layers
Calculate a weighted read over earlier layer outputs, derive its weights with a learned query, and examine the storage and selection trade-off of Block Attention Residuals.
2026-09-29Cross-Layer KV Sharing: Reusing Keys and Values from Earlier Layers
Compare KV storage across six layers, follow a new token as its Keys and Values are reused by the next layer, and examine the trade-off between storage and computation.
2026-09-29YOCO: Separating Layers That Produce and Read Shared KV
Follow shared KV from early to late layers and trace new-token processing to understand why prefill can skip past computations unnecessary for the first output.
2026-09-29CED: Continuing Context with a Causal Encoder and Decoder
Compare generation paths and the sources of global and local KV, then connect sparse selection, three reuse modes, sequence compression, and prefill replay.
Hardware
2026-09-10CPU and GPU: Two Devices That Execute Model Computation
Explore why GPUs are used, how CPU and GPU designs differ, how host and device work together, and why compute units need a supply of data.
2026-09-10GPU Architecture: Compute Units and Memory
Explore SMs, CUDA Cores, Tensor Cores, and GPU memory, then connect capacity, bandwidth, and access latency to running a model.
2026-09-11Parallelism in Model Operations: Element-wise Operations, Reductions, and Matrix Multiplication
Compare element-wise operations, reductions, and matrix multiplication through independent output elements and the inputs and computation needed for each result.
2026-09-11Parallel Execution on GPUs: From Threads to Warp Scheduling
Start with a common computation procedure and indexed executions, then explore threads, blocks, grids, SM assignment, warp scheduling, and latency hiding.
2026-09-17How the CPU and GPU Execute Work Together
Distinguish CPU requests from GPU execution, then use buffers, streams, and events to explain data preparation, completion waits, and execution order.
2026-09-12Starting GPU Optimization: Arithmetic Intensity and Data Movement
Connect memory and compute bottlenecks through arithmetic intensity and Roofline, then explore reduced data movement, concurrent progress, and element-wise kernel fusion.
2026-09-12Optimizing Matrix Multiplication: Input Reuse and Tiling
Explore input reuse in shared memory and registers, the relationship between tile size and resource usage, and how data movement can overlap with computation.
2026-09-12Why Attention Is Difficult to Optimize
Examine the cost of storing and rereading large score and probability matrices, and why a subset of scores cannot determine final Softmax probabilities.
2026-09-13Processing Scores in Chunks with Online Softmax
Explore why stable softmax subtracts the maximum, how the maximum and exponential sum work together, and how online softmax updates them with each new chunk of scores.
2026-09-14Reducing Memory Traffic with Output Accumulation in FlashAttention
Explore joint accumulation of the Value weighted sum and exponential sum, then follow FlashAttention-2 tiles to complete the output and reduce memory traffic from large intermediate matrices.
2026-09-14From Operation Optimization to Whole-Model Performance
Explore how improving one operation affects whole-model execution time, why optimization requires repeated bottleneck measurements, and when performance goals or memory capacity call for more resources.
2026-09-14Scaling Across Multiple GPUs
Explore compute and memory across GPUs, distinguish capacity, latency, and throughput goals, and examine how communication and synchronization-related waiting affect total execution time.
2026-09-17Transferring Data Between GPUs
Prepare communication participants and buffers, connect CPU send/receive submission to GPU execution, and explain overlapping independent computation and safe buffer reuse.
2026-09-17SMs and Copy Engines in GPU Communication
Distinguish the devices and connections used for GPU communication, and examine how resource contention with computation affects overall completion time.
2026-09-17GPU Communication Across Servers and the CPU's Role
Explore transfer paths that reduce intermediate copies through CPU memory and the CPU role in helping network communication progress.
2026-09-18The Basic Operations of Collective Communication
Compare Send/Recv with collectives, then follow the inputs and outputs of Broadcast, Scatter, Gather, and Reduce.
2026-09-18Collective Communication: Combinations and Extensions
Understand All-Gather, All-Reduce, Reduce-Scatter, and All-to-All through the data and output placement needed by the next computation.
2026-09-18How Collectives Move Data: Ring and Tree
Follow Ring and Tree implementations of the same All-Reduce result, and relate communication time to steps, data volume, and physical links.
2026-09-18DP: Replicate the Model, Partition the Inputs
Replicate a model across GPUs, distribute inputs, and distinguish inference request distribution from training gradient synchronization.
2026-09-18TP: Split One Operation Across GPUs
Partition weight matrices by columns and rows, then connect these partitions in FFN and attention to avoid intermediate communication.
2026-09-18SP: Partition Activations Alongside TP
Reduce activation duplication between TP regions, and alternate between token and feature partitioning to compute a Transformer layer.
2026-09-18CP: Split Long Contexts Across GPUs
Compare CP with SP, then follow how Ring and Ulysses connect information across GPUs to complete attention.
2026-09-18PP: Partition and Execute Model Layers
Place model layers across GPUs, then split the same batch into microbatches to overlap computation for different inputs.
2026-09-18EP: Partition Experts and Route Tokens
Use DP for attention and EP for experts on the same GPUs, then follow tokens as they fan out to experts and return to their original positions.
2026-09-18How Should We Arrange Multiple GPUs?
First partition the model so the workload can run, then replicate that configuration and adjust GPU placement for latency and throughput.
Inference
Workloads
2026-09-19Inference and the KV Cache
Follow repeated token generation to see why computation repeats, how the KV cache reuses it, and where output-token selection differs from KV computation.
2026-09-19Prefill and Decode
Compare Prefill, which processes input tokens together, with Decode, which processes one new token at a time, and examine how token count and weight reuse change GPU bottlenecks.
2026-09-19Batching and Scheduling
Explore the limits of static batching, how batch membership changes between executions, how to check computation and KV space requirements, and why unused reservations make requests wait.
2026-09-19KV Cache Management and PagedAttention
Allocate KV blocks as needed and compute attention over scattered KV. Learn how block sharing and copy-on-write reduce storage when generating multiple responses from the same input.
2026-09-19When KV Capacity Runs Out: Pausing and Resuming Requests
When KV space runs out during generation, pause some requests to reclaim blocks, then resume them from retained token history or KV. Explore how this process delays output.
2026-09-19Inference Metrics: Latency and Throughput
Measure one request's waits and overall server throughput, examine output gaps and request latencies hidden by averages, then connect them to goodput: throughput that meets latency targets.
RL
How Token Generation Becomes an Action in Reinforcement Learning
Compare choosing a move in a maze with an LLM choosing its next token, then follow how a response made of many tokens receives a reward.
2026-09-25How Rewards Change Token Generation Probabilities
Connect advantages computed from rewards with token log-probabilities, then explore probability ratios and clipping for reusing data, and KL regularization for staying close to a reference policy.
2026-09-25Computing Advantages with Groups and Critics
Start with GRPO comparisons within a prompt, examine critic predictions and advantages, then connect them to single-rollout learning in SAO and FlashREINFORCE.
2026-09-25Learning from a Teacher’s Probabilities: OPD
Explore on-policy distillation, which reads a teacher’s token probabilities in student-generated contexts and turns the differences into token-level learning signals, and how Miles combines those signals.
2026-09-27Combining RL and OPD in a Training Strategy
Starting with DeepSeek V4’s specialist training and OPD integration, we examine how to prepare a starting policy for RL and how MiMo MOPD² chooses teachers and contexts.
2026-09-25How an LLM RL System Fits Together
Follow generated experience to the learner and updated weights back to the inference engine to understand the roles, data, and GPU placement in an LLM RL system.
2026-09-25Carrying Generation Records into Training: TITO and R3
Preserve actual token IDs with TITO and reuse the MoE experts selected during generation with R3.
2026-09-25Using Quantization in RL
See how quantization reduces generation costs and changes training and inference computation, through shared quantization and BF16 training with FP8 inference.
2026-09-25Why Generation and Training Probabilities Differ
Explore how to reduce differences caused by computation order, then use probability ratios to account for remaining engine and model-version differences.
2026-09-25Asynchronous RL and Stale Data
See how published rollout versions advance in asynchronous RL, and when to exclude experience that ages during generation or while waiting in a buffer.
2026-09-25Inference Optimization for RL Rollouts
Explore where and under what conditions shared-prefix KV reuse, batching variable-length requests, speculative decoding, and MTP reduce RL generation costs.
2026-09-25Scheduling the Whole Agent Program
Track the state of programs that connect model calls and tool execution, and examine how execution, pausing, and reassignment balance KV retention against recomputation.
2026-09-25Delivering New Weights to the Inference Engine
We examine how to reconcile the learner's and inference engine's weight partitions, choose a transfer path, and resume generation with a consistent new policy.
2026-09-28LLM RL in Review: From Token Choices to the System Loop
Follow the key figures from thirteen articles in a 30-minute walkthrough, from learning that changes token probabilities to the system that moves experience and weights.