Shared Concepts · Models · 2026-09-27
MQA and GQA: Sharing KV Across Queries
Use MHA as a baseline to compare KV sharing in MQA and GQA, and examine how they reduce storage while preserving per-query outputs and what trade-offs follow.
In Q·K·V projections and Core Attention, we followed how each head gathers information using its own query, key, and value. In that arrangement, increasing the number of query heads also increases the number of corresponding key and value heads.
When generating tokens one at a time, we can store keys and values from earlier positions for reuse. As the context grows, however, there are more values to store. Could multiple queries use the same keys and values to reduce storage while preserving the computation that gathers information for each query?
This article first compares the sharing patterns of three structures. We then trace how one token leads to computation across multiple heads in MHA, before extending the example to MQA, where all query heads share KV, and GQA, where sharing occurs within groups. Finally, we examine exactly what becomes smaller, by how much, and the resulting trade-offs.
Connecting queries to KV
MHA (Multi-Head Attention) uses a separate key head and value head for each query head. MQA (Multi-Query Attention) shares one key head and one value head across multiple query heads. GQA (Grouped-Query Attention) divides query heads into groups and shares keys and values within each group.
We will use KV to refer to keys and values together, and one KV head to mean a corresponding pair of one key head and one value head. The following figure compares four, two, and one KV heads while keeping four query heads.

The vectors through are the query vectors computed by the current token’s four query heads. They are not queries from four different tokens. Each KV set below them represents the keys and values those queries will access. The arrows show which KV set each query uses.
In MHA, the mapping is one-to-one: uses , uses , and so on. In GQA, shown in the middle, and share , while and share . In MQA, all four queries use the same . All three structures have four query heads; what changes is the number of KV heads and the scope of sharing. This comparison redraws the GQA paper’s architecture diagram with four query heads.
MHA: separate KV for each head
Consider processing the current position in one layer. Positions , , and precede it, so there are four token positions to attend to, including the current position. The number after denotes a position in the context, rather than a token type or a head number.
For a small example, let the model dimension be 8, the number of query heads be 4, and the dimension of each head be 2. Keys and values also have two components per head. We will keep these conditions throughout the figures. These are illustrative dimensions, not measurements from an actual model.

In the first panel of Figure 1, each block from to represents one input vector at that position. It is the vector entering this layer, not a single vector component. We multiply the entire eight-component vector at the current position by three different weight matrices, , , and , to produce Q, K, and V.
In this MHA example, each projection produces eight components. Grouping them into four heads with two components each gives , , and . Each small cell below represents one vector component, and each row represents one head. For example, , , and are all vectors for Head 1 at the current position. We do not first split the input vector into four pieces and send one piece to each head. Each head projects the entire input using different weights.
Suppose we have kept the keys and values for through before computing the current position. This storage of keys and values from processed positions is the KV cache. Adding the newly computed and to Head 1’s cache leaves and containing four positions, from through . The other heads add their current-position values in the same way.
Here, lowercase and denote vectors at the current position, while uppercase and denote matrices collecting multiple positions. Each matrix is four positions × two components per head, or . Moving down one row means moving to the next token position; switching from to means using another head’s representation.
The third panel expands Head 1’s computation. Taking the dot product of the two-component with each of the four rows of produces four position scores. Applying the scaling and softmax described in Core Attention produces four weights that sum to 1. The current position can attend to every position from through itself.
Using those weights to take a weighted sum of the four rows of gives the two-component output . Here, the weights are the proportions determined by this query for incorporating the value at each position. They are distinct from the learned parameters , , and used to produce Q, K, and V.
The other three heads similarly compute , , and using their own queries and KV. Concatenating the four outputs gives components, which the output projection combines. Labels such as and in the figure represent the two components of each output vector.
Keys and values from past positions are reused when computing the next position, but past queries do not need to be stored in this KV cache. The new position’s query reads the stored keys and values. We can now consider reducing the number of KV sets stored in this computation.
MQA: sharing KV across all queries
In MHA, each head produced its own keys and values at every position. MQA produces one key vector and one value vector per position and shares them across all query heads. Queries are still produced separately for each head.
At the current position , the Q projection still produces four queries with two components each. The K and V projections, however, each produce only two components. With a model dimension of 8, this configuration has an matrix and matrices and .

The left side of Figure 2 contains one K matrix and one V matrix, each of size . Having one KV head does not mean storing only one token position. The shared K and V still contain values for all four positions, from through .
On the right, computes scores using the four rows of K, applies softmax, and takes a weighted sum of V to produce . The query uses the same K and V, but its different query vector can produce different position scores and weighting proportions. For example, one query may give more weight to the value at , while another gives more weight to the value at .
Thus, sharing the same KV does not force the outputs to be identical. Each of the four queries performs its own attention computation and produces its own output. The four Attention boxes distinguish these computation relationships; they do not require four separate kernel launches. As in MHA, the outputs are concatenated and passed to the output projection.
Storage changes substantially. MHA stores four sets of K and V, while MQA stores just one. In this example, that is components. By sharing KV across heads, MQA reduces the amount of KV to store and read during token-by-token generation. MQA paper
GQA: sharing KV within groups
GQA defines the scope of KV sharing by group. If we divide four queries into two groups, the first contains and , and the second contains and . Each group uses a separate K and V.

In Figure 3, and both read , while and both read . Each group’s KV contains all four positions. The groups divide query heads, not the token positions stored in the cache.
At the current position, the Q projection still produces four queries with two components each. The K and V projections each produce two heads with two components each. In this example, and are each , and each KV head adds its newly computed values to its corresponding cache.
Within a group, sharing works as in MQA. The queries and read the same KV but gather values using their own computed proportions. The second group follows the same process. There are still four outputs, through , but only two KV sets to store. Total storage is components.
Let denote the number of query heads and the number of KV heads. In the usual configuration with equally sized groups, query heads share each KV head. In this example, that is . Setting gives MQA’s sharing pattern; setting gives MHA’s. GQA adjusts the degree of sharing between these two endpoints. GQA paper, Section 2.2
KV head count and storage
We can now compare the three KV caches using equally sized cells. Each cell represents one stored component. All three configurations have four token positions and two components per head for both K and V.

In MHA, each of the four KV heads has eight key components and eight value components, giving 64 components in total. GQA has two KV heads and 32 components; MQA has one KV head and 16. This reduction comes from storing fewer separate KV heads, while preserving the number of token positions and components per head.
For one request in one layer, we can generalize this as follows. Here, is the number of token positions in the cache, is the number of KV heads, and is the dimension per head. We assume equal head dimensions for K and V.
The leading 2 accounts for storing both K and V. Substituting the MHA example gives . For GQA and MQA, only changes, to 2 and 1 respectively. With the other conditions and storage size per component held equal, GQA’s KV storage relative to MHA is . In this example, GQA uses half as much storage and MQA uses one quarter.
This expression counts the KV values themselves. For an entire model, storage must be summed across layers; for multiple requests, it must also be summed across requests. Actual memory allocation may include additional space depending on the implementation.
Even with sharing, each query computes scores against the keys at the positions it attends to and takes a weighted sum of their values. Keeping the number of query heads and the context length unchanged leaves this core attention arithmetic at the same scale. Meanwhile, the output size of KV projections and the amount of KV to store decrease. An implementation that reuses shared KV effectively can also reduce the amount of data fetched from memory.
Consequently, reducing the cache to one quarter of its size does not necessarily reduce total execution time to one quarter. Other operations, model-weight reads, and how the kernel reuses KV also affect runtime. The figures 64, 32, and 16 in this article compare stored component counts.
Choosing a sharing scope
MHA can learn separate key and value representations for each query head. MQA requires all queries to use the same key and value representations; GQA applies that constraint within each group. Broader KV sharing reduces storage, but also reduces the freedom to produce different keys and values for different heads.
For this reason, arbitrarily combining KV at runtime in a model trained with MHA does not preserve its outputs. MQA and GQA are model structures with different KV projections and head connections. The GQA paper presents a method that converts the KV projections of an existing MHA checkpoint and then performs additional training. In its experiments, it reports GQA quality close to MHA and speed close to MQA. Those results apply to the paper’s models and training and execution conditions. GQA paper, Sections 2–3
When reading an actual model configuration, therefore, check both the query-head count and the KV-head count. Their ratio helps explain the sharing scope and KV storage, while quality and speed must be evaluated with the model trained for that structure and its execution environment.
We have examined how to change KV sharing while preserving per-query computation and outputs. The next article moves beyond the scope of sharing to consider making the stored KV representation itself smaller.