Definition
Implicit reasoning is the ability to combine stored facts or rules inside a single forward pass without exposing an explicit chain-of-thought. In the paper’s controlled task, the model receives a starting entity and a sequence of relations, then must infer the final entity by composing several atomic facts.
Key Points
- The task separates ordinary in-distribution generalization from two harder tests: systematic generalization and depth extrapolation.
- Systematic generalization means composing atomic facts in combinations that were never used together during training.
- Depth extrapolation means applying a learned composition rule to more hops than the model saw during training.
- Vanilla transformers struggled with both challenges in the reported experiments, while recurrent-depth-transformers handled them through shared iterative computation.
- The experiments use synthetic directed knowledge graphs, which make the composition problem measurable but limit direct claims about real-world language understanding.
Generalization and Grokking
The reported training trajectory moves from memorizing observed answers, to generalizing within the training distribution, and only later to composing unfamiliar facts. The paper interprets this as a three-stage grokking process rather than as one smooth improvement in accuracy.
Related Concepts
- recurrent-depth-transformers — architecture tested as a way to support implicit multi-hop composition.
- knowledge-graph — the controlled task stores atomic facts as graph edges and multi-hop answers as inferred paths.
- rag — external retrieval is a contrasting way to supply and compose information rather than relying only on parametric knowledge.
Implications
Implicit reasoning exposes a gap between knowing individual facts and composing them in unfamiliar arrangements. For system design, better factual recall should not be treated as evidence of better compositional reasoning; the two capabilities need separate tests.
Open Questions
- How much of the observed behavior depends on the synthetic task format?
- Would explicit intermediate state improve reliability, or merely make the same computation visible?
- What evaluation best distinguishes genuine composition from shortcuts in larger language models?