RL · 2026-09-25
Carrying Generation Records into Training: TITO and R3
Preserve actual token IDs with TITO and reuse the MoE experts selected during generation with R3.
The previous article traced experience from the inference engine to the training engine. The learner recomputes that record to update the model. Is matching the human-readable response enough to train on the original experience?
The same text can have different token IDs, and the same tokens can take different expert paths in an MoE model. This article examines two ways to carry generation records into training. TITO preserves actual token IDs. R3 transfers the selected expert IDs so training uses the same paths.
Converting to Text Loses Token Boundaries
A model selects token IDs one at a time. For illustration, suppose token 101 represents “A,” 202 represents “B,” and 303 represents “AB.” These are teaching examples, not measurements from a particular tokenizer.
If the model generates 101 followed by 202, the readable result is “AB.” But when that text is tokenized again, the tokenizer may represent “AB” as the single token 303. The visible characters match, but the two choices the model actually made have not been preserved. The text does not retain the original token boundaries.

Training on 303 does not train the original choices of 101 and 202. Their two recorded generation log probabilities cannot simply be paired with the new single token. If the generated response becomes input to a later call, that input token sequence changes too.
Re-tokenization does not always change the IDs. The point is that matching text alone does not guarantee recovery of the original sequence. Applying a message template or reconstructing tool-call content can also change the text itself.
TITO Passes the Actual Token IDs
TITO (Token-In, Token-Out) manages inputs and generated outputs as token IDs. As on the right of Figure 1, the training record uses the IDs actually consumed and generated by inference, rather than reconstructing them from text. Readable output for the user can be produced separately.
The same principle applies across turns. Preserve the existing input and response tokens, tokenize new user messages or tool results with the required message delimiters, and append them. Previously generated responses need not be reconstructed from text. Keep their generation log probabilities aligned in the same order. Miles v0.1 §2.4, Miles TITO implementation
The policy loss evaluates a generated token in its context. TITO preserves the objects of that calculation: the actual input tokens and selected tokens.
Record Expert Choices in MoE Models
An MoE layer runs selected expert networks. The router chooses experts; an expert is a computation module that executes. Experts are components inside the model, not separately invoked chatbots.
Suppose inference selects E1 and E2 for a token. Small numerical differences can make training select E1 and E3 even with matching inputs and weights, changing the path used to compute the training result.
R3 (Rollout Routing Replay) records generation-time expert IDs and reuses those choices in training. Figure 2 transfers [1, 2] with the experience so training runs E1 and E2. R3 §3–4

Choices vary by token position and MoE layer, so records must identify both. One expert list for an entire response is insufficient.
Train the Same Experts with Current Weights
R3 reuses which experts were selected, not their old output values.
Training executes E1 and E2 with current weights and computes current router mixing weights, then evaluates the loss and backpropagates. Expert choices are replayed while expert and router parameters remain trainable. R3 §4.1
In the experience-transfer path from article 6, TITO carries actual token IDs, and R3 adds expert IDs for each token and layer. Training uses these records to reproduce generation-time inputs and expert choices.