What happens to an RL rollout when the sampling policy receives an update mid-sentence?
In a reinforcement learning setting for language models, the inference engine generates answers to prompts, while the trainer updates the weights based on their score. In a naive implementation, the engine generates a full batch, sends it to the trainer, and waits for the updated weights before resuming generation. A natural asynchronous alternative is to keep generating the answers while the trainer trains, and push new weights to the engine whenever they are ready. But what should happen to the responses which are mid-generation?
PipelineRL popularised regime D, reporting “only slightly higher divergence” on stale cache, while AReaL chose to interrupt the generation and recompute KVs (E).
On a host with two H100 SXM GPUs, Qwen3-0.6B, and vLLM 0.30, asynchronous regimes are almost indistinguishable on the GSM8K dataset, with stale cache regime (D) nominally edging ahead of B and E in accuracy, as it doesn’t need to wait for in-flight answers or recompute KV, and completes more optimiser steps throughout. However, for GSM8K, answers are generally very short and barely any of them got caught mid-generation, with <1% of tokens consumed by the trainer coming from a stale cache for regime D.
For the MATH dataset, the answers are generally longer, and 38% of tokens the trainer consumed came from stale cache. Here, we can see clear separation from the drain regime (B), with in-flight regimes D and E converging meaningfully faster. There is no measurable difference in resulting accuracy between stale cache (D) and the recomputed one (E).
In a separate teacher-forcing study on GSM8K, comparing stale cache generation against a recomputed KV cache shows that stale cache undoes over 80% of the updated policy effect for the first 8 tokens, and ~49% after token 32, where the share undone is , pooled over positions. However, weight updates have minimal effect (5e-5 nats per token per optimiser step), and the stale-cache regime is an off-policy importance sampler using generation-time log-probabilities. Recomputing the first 14 of 28 layers removes ~10% of the stale-cache divergence, while recomputing the first 21 of 28 removes ~83%.
GSM8K results computed across nine seeds; MATH used three. Code is available at zhebrak/half-written.
Related work and further reading
- Piché et al., PipelineRL, 2025
- Fu et al., AReaL, 2025
- Mistral AI, Magistral, 2025
- Khatri et al., The Art of Scaling Reinforcement Learning Compute for LLMs, 2025
- Yao et al., Your Efficient RL Framework Secretly Brings You Off-Policy RL Training, 2025
- Hugging Face, Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries, 2026
- vLLM, Async Reinforcement Learning
- Liu et al., DroidSpeak, 2024