← Learning path

RL · 2026-09-25

Asynchronous RL and Stale Data

See how published rollout versions advance in asynchronous RL, and when to exclude experience that ages during generation or while waiting in a buffer.

The previous article examined how to correct for differences between the probabilities of the policy that generated experience and the policy being trained. This time, we will follow the execution that lets those differences grow. If inference keeps producing the next experience while the learner updates the model using ready experience, the two jobs can spend less time waiting for each other. But even newly completed experience may already have been generated with older weights.

Asynchronous RL overlaps generation and training while managing experience that becomes stale along the way. We will first see how the two jobs overlap, then distinguish training steps from the versions installed on the inference engines. We will follow how experience ages during generation and in the buffer, and examine the order in which Miles’s default buffer consumes groups and the points at which it excludes them.

Overlapping generation and training

Allocating separate GPUs to inference and training does not automatically make the two jobs run concurrently. If training starts only after generation finishes, and the next generation starts only after training and weight synchronization finish, one side waits while the other works.

Figure 1 compares synchronous and asynchronous execution on separate GPUs. G1 and G2 are each a collection of completed groups needed for one round of training. The bar lengths illustrate execution order; they are not measured performance results.

Synchronous execution finishes G1 generation, G1 training, and synchronization before generating G2. Asynchronous execution generates G2 while training on G1, briefly pausing generation for synchronization. Training still waits when data is not ready.

In synchronous execution, the system generates G1, trains on it, installs the new weights on inference, and then generates G2. In asynchronous execution, G2 is generated while the learner trains on G1. The training loop that takes ready experience from the buffer progresses independently of the generation loop that produces and inserts the next experience.

Miles’s fully async configuration still briefly pauses generation during weight sync. The learner also waits if G2 is not ready when training on G1 finishes. Asynchrony reduces waiting by overlapping work that can proceed concurrently; it does not make each individual computation faster. Miles’s fully async schedule

The conditions required for training data remain unchanged. A configuration such as GRPO that compares several responses to the same prompt uses groups whose required samples and evaluation results are ready. Continuing to generate other groups does not mean arbitrarily taking one response from an unfinished group for training.

Training steps and published rollout versions

We need to distinguish two records of model progress. A training step records the learner’s progress as it updates weights. A published rollout version in this article records the installation of new weights on the inference engines. Publishing here means updating the generation model inside the RL system, rather than launching a service for external users.

Figure 2 assumes one update per training step and weight synchronization every three steps. We call the state that installs the weights from step 9 on inference v3.

The learner advances from step 9 through steps 10, 11, and 12. Inference keeps v3, the weights from step 9, until synchronization at step 12 publishes v4. At step 11, v3 experience has rollout-version lag zero, although the learner has updated twice.

The learner changes its weights at steps 10 and 11. But without another sync, inference continues generating with v3, the weights from step 9. At step 12, synchronization installs the weights from that step and advances the published version to v4. Training has advanced three times, while the published version has advanced once.

What should we call the training model when it uses experience generated by v3 at step 11? In this example, we can say “the learner has the weights from step 11, and its last published rollout version is v3.” Calling the learner’s weights simply v3 would hide the two updates.

The staleness used by Miles’s fully async buffer is this rollout-version gap:

Current version installed on inference − oldest generation version anywhere in the group

At step 11, a group generated entirely by v3 therefore has a gap of 3 − 3 = 0. This metric can be zero even though the learner’s weights have changed twice. Miles’s documentation also defines the metric in published rollout versions and distinguishes it from training-step counts when weights are not synchronized every step. Miles’s staleness metric

This definition describes this particular Miles buffer. We should not assume that every system’s staleness uses the same counter. Nor is a version gap a distance between probability distributions. A gap of zero can still accompany different current learner weights, and the engine computation differences from the previous article can remain.

Experience can age during generation

Experience does not begin aging only after generation finishes. While a long response or a multi-turn tool interaction is still running, weights learned from other experience may be installed on inference. Another publication can occur after the experience completes and waits in the buffer.

Figure 3 follows one experience that starts with v3. Generation briefly pauses to install v4, then resumes while retaining the prefix.

One experience starts with v3 and continues with v4 after synchronization. At completion, the gap between the oldest generation version v3 and current v4 is one. If v5 is published during buffer waiting, the gap becomes two at dequeue.

The earlier tokens were generated by v3 and the later tokens by v4. Continuing with v4 does not turn the prefix into experience generated by v4. The oldest generation version remains v3, so the gap at completion is 4 − 3 = 1.

An execution can therefore include multiple generation versions within one response. Miles’s default retract mode returns in-flight generation to the waiting queue for synchronization, recomputes the KV cache with the new weights, and resumes generation. The prefix’s generation history remains distinct from the tokens generated afterward. This does not mean every experience necessarily spans multiple versions. Pausing and resuming generation in Miles

The completed experience now enters the buffer. Its stored generation history remains unchanged if v5 is published while it waits. At dequeue, the gap is 5 − 3 = 2. The version gap has grown once during generation and once more while waiting in the buffer.

Miles’s default buffer actually checks the entire group, rather than an individual experience. If no other sample in the group containing the illustrated experience has a version older than v3, the group’s oldest version is also v3. If any other sample includes v2, the calculation uses v2. This is a conservative rule that never treats a group as fresher than its oldest generation record. Miles v0.1 §2.2.2

Checking the buffer’s entry and exit

Miles’s default buffer uses FIFO: it takes out the group that entered first. This differs from the order in which generation requests started. Even if A started first, the buffer takes B first if the groups became ready and entered in the order B, C, A. Buffer order determines which waiting group is checked first.

Figure 4 separates the three exclusion reasons described in the Miles v0.1 paper by their check points. The first two are checked when a group arrives; staleness is checked when the group is taken out for training.

The buffer checks aborted or failed generation and the user filter at entry. Accepted groups admitted in B,C,A order are dequeued in that order. With current v5 and maximum lag one, B with oldest generation version v3 is excluded, while C with v4 enters the training batch.

The first reason is a group whose generation could not finish. For example, an agent run may exceed its collection timeout and return an aborted status. This does not mean the buffer takes out and discards a group that is still generating normally. It checks a result that generation has given up on and excludes it from training at entry.

The second is a group rejected by a user filter. For example, if every response in a group receives the same reward and all relative advantages are zero, a filter can be configured to exclude the group. The training configuration determines which results to exclude.

The third is a group exceeding the allowed staleness. Because this condition can change after admission, it is checked at dequeue. In the figure, the current published version is v5 and the maximum gap is one. If B at the front has oldest generation version v3, it is excluded because 5 − 3 = 2. If the next group C has v4, it passes with 5 − 4 = 1 and enters the training batch. The learner waits longer if it has not collected enough usable groups. Miles v0.1 exclusion conditions, Table 1, Miles’s default buffer implementation

FIFO therefore means checking groups in admission order; it does not mean that every admitted group is trained on. A group that was fresh at entry can be excluded at consumption. These three reasons follow the paper’s scope. They are not an exhaustive list of every validity check in the latest implementation or of additional filters applied after assembling a training batch.

Balancing waiting and discarded work

In operation, buffer capacity, allowed staleness, and the weight-sync interval need to be considered together. They determine how much to store, what can be consumed, and how often to update the generation model.

A larger buffer can absorb more temporary differences between generation and training rates. It also allows experience to wait longer. When Miles’s buffer fills up, inserting another group waits for space. This restriction on production according to consumption speed is called backpressure.

Reducing the allowed staleness excludes older experience more aggressively. However, more generation work may then fail to contribute to training, or the learner may wait for fresh data. Miles configures this limit through --max-weight-staleness; leaving it unset disables this filter.

More frequent synchronization lets subsequent generation catch up with the learner’s latest weights sooner. It can also increase weight-transfer and generation-pause costs. Less frequent synchronization lets generation continue with older weights for longer. The same rollout-version gap of one can then contain a different number of training updates, so staleness numbers should not be compared unchanged after modifying the sync interval. Miles’s --update-weights-interval sets the publication interval in training-loop iterations; the number of optimizer updates within each iteration must also be checked.

The system can also choose whether to regenerate an excluded task. Miles’s retry returns the original prompts of aborted or overly stale groups to the data source to generate fresh experience. It does not put the same old records straight back into the buffer. Groups rejected by the user filter are excluded from this retry path. Miles’s buffer controls

Looking at how often the buffer empties, how many groups are excluded as stale, and the version gaps of groups actually consumed reveals the cost of reducing waiting. The next article will examine inference optimizations that reduce the cost of generation itself.

Back to contents ↑