Definition
Recurrent-depth transformers, also called looped transformers, reuse the same transformer block for multiple sequential iterations. Instead of assigning each reasoning step to a different layer with different parameters, the model applies shared layers repeatedly, giving inference-time recurrence a role in the effective reasoning depth.
Key Points
- In controlled synthetic multi-hop tasks, recurrence enabled systematic generalization that vanilla transformers failed to achieve.
- Increasing the number of training-time iterations increased the learnable recursion depth without proportionally adding parameters.
- Increasing inference-time iterations enabled depth extrapolation, allowing models to solve compositions deeper than those seen during training.
- Dynamic recurrence performed better than a fixed iteration budget when the training data was sufficiently complex.
- Excessive recurrence caused overthinking: confidence and accuracy rose until a peak, then degraded as iterations continued.
Training Dynamics
The paper reports a three-stage grokking dynamic: memorization, in-distribution generalization, and systematic generalization. The final stage appears only after near-perfect in-distribution performance, suggesting that learning a reusable composition rule can take substantially longer than fitting the observed examples.
Adaptive Halting
Adaptive halting can allocate more iterations to harder inputs and fewer to easier ones. The paper reports that using only output-distribution change can halt too early; combining a KL-divergence threshold with an entropy threshold produced a more reliable stopping signal in its setting.
Related Concepts
- implicit-reasoning — The reasoning setting used to test whether parametric knowledge can be composed without explicit chain-of-thought.
- knowledge-graph — The paper represents its controlled multi-hop task as a directed graph of entities and relations.
- agent-memory-systems — Recurrent computation offers a model-internal alternative to repeatedly retrieving intermediate reasoning state.
Implications
The result points to a design trade-off: shared recurrence can make extra inference compute useful for deeper composition, but a stopping policy is necessary because more compute can eventually make predictions worse. This is promising architectural evidence, not yet proof that the same behavior transfers to frontier LLMs.
Open Questions
- Do recurrent-depth transformers retain these generalization benefits on natural-language and heterogeneous pretraining data?
- Can adaptive halting be learned robustly instead of relying on task-specific thresholds?
- How does recurrence compare with external retrieval or explicit chain-of-thought when both receive the same inference-time compute?