Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
Abstract
Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory. We therefore improve the two components of training that shape these fixed points: the depth prior and input injection. Fixed-depth training breaks KV sharing, and Huginn's broad depth prior supports sharing but dilutes supervision at the target depth more than sharing requires; we learn the prior from prediction feedback, with an entropy term that keeps it broad. Existing injection schemes let the state's component along the input amplify or cancel the injection; we remove this component with orthogonal injection. From 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale relative to Huginn's prior and existing injection schemes, respectively. At 1.6B, the learned prior with a 3x smaller KV cache matches the downstream average of fixed-depth training with the full cache.
Community
Scaling up a model has meant paying twice, in compute and in memory.We show that looped models can pay in compute alone, using the loop's fixed point as a shortcut.
A 1.6B looped model runs twelve blocks deep on four blocks of memory. On the same tokens, it beats a four-block Transformer by 17.5% and comes within 2.4% of a twelve-block one needing 3x memory. It can also train, pre-fill, and post-train without paying for every loop.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory (2026)
- GrowMTP: Can RL Grow Its Own Draft Head? (2026)
- The Surprising Effectiveness of Shared Memory in Looped Transformers (2026)
- Reasoning on the Simplex: Geometric Fixed-Point Models (2026)
- ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation (2026)
- GradLev: Token-Parallel Test-Time Training Via Costate Prediction (2026)
- Loop Dropout: Regularizing Shared Updates in Looped Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 24
IFM/LoopedLM-P2-distilled-s-random-init
Datasets citing this paper 0
No dataset linking this paper