Ghosts vs. Animals (Jagged Intelligence)
Overview (confidence: medium)
Andrej Karpathy frames LLM intelligence as a fundamentally different kind of mind than animal intelligence, not a lesser or immature version of it. In “The Space of Minds,” he argues the space of possible intelligences is vast, and animal intelligence — the only kind we had ever encountered before LLMs — is a single point in it, shaped by a very specific optimization pressure: innate embodiment, self-preservation, natural-selection-driven drives for power/status/dominance/reproduction, deep social cognition (theory of mind, coalitions), and exploration/exploitation heuristics like curiosity and fun.
LLM intelligence is shaped by an entirely different optimization pressure: statistical imitation of human text (the bulk of “supervision bits”), reinforcement learning against task rewards (an “innate urge” to guess at the underlying task to collect reward), and at-scale A/B-testing for daily-active-user engagement (producing sycophancy — “craving an upvote from the average user”). Because the computational substrate, learning algorithm, and — most importantly — the optimization objective all differ from biological evolution, Karpathy argues LLMs should be understood as “ghosts”: statistical distillations of humanity’s text, not as immature or alien animals. He calls this humanity’s “first contact” with non-animal intelligence.
The Bitter Lesson Debate (confidence: medium)
In “Animals vs Ghosts,” written after listening to a Dwarkesh Patel podcast with Richard Sutton (author of “The Bitter Lesson,” a foundational text in frontier LLM circles about the primacy of compute-scalable methods), Karpathy engages Sutton’s counterintuitive claim: LLMs are not actually “bitter lesson pilled.” Sutton’s objection is that LLMs train on giant, finite, human-generated datasets — biased and eventually exhausted — rather than learning purely through environment interaction the way his “classicist” vision (a Turing-style “child machine”) would require. Sutton points out that supervised finetuning, where actions are directly imitated from demonstrations, has no clean analogue in the animal kingdom.
Karpathy’s counter: animals are not a clean “learning from scratch” counterexample either. A newborn zebra runs within minutes of birth — a complex sensorimotor feat achievable only because of a powerful genetic initialization encoded via evolution’s “outer loop” optimization, not because it is doing reinforcement learning at random from a blank slate. In this framing, pretraining on internet text is LLMs’ crude analogue to evolution: a way to gather enough soft constraint over billions of parameters to avoid starting from scratch, later refined by finetuning/RL — much like an animal’s genetically pre-initialized brain is later shaped by lived experience. Karpathy’s summary line: “today’s frontier LLM research is not about building animals, it is about summoning ghosts,” and he offers the open-ended analogy “ghosts:animals :: planes:birds” — useful and world-altering, but not necessarily convergent with biological intelligence.
A technical aside from the same post, relevant to agent-memory design in this vault: Karpathy notes that in-context learning is a form of test-time adaptation that does not involve weight training, and that recent interest in memory (explicitly citing files like CLAUDE.md) uses text/context — not weights — as the substrate for this kind of test-time learning.
Named as a 2025 Paradigm Shift (confidence: medium)
In his 2025 year-in-review, Karpathy names “Ghosts vs. Animals / Jagged Intelligence” as one of the year’s defining conceptual shifts — the point where he (and, in his view, the rest of the industry) began internalizing the shape of LLM intelligence intuitively rather than reflexively applying an animal-intelligence lens to it. He directly connects this framing to his loss of trust in benchmarks (see verifiability): jagged capability profiles mean an LLM can be simultaneously a “genius polymath” on verifiable tasks and easily jailbroken or confused on tasks resembling human common sense.
Implications: Treating LLM failures/successes through an animal-intelligence intuition (e.g. “it should generalize like a smart person would”) systematically mispredicts behavior. The ghosts framing suggests evaluating and designing around LLMs’ actual optimization history (imitation + verifiable-reward RL + engagement-tuning) rather than analogizing to biological cognition — with direct design consequences for agent memory (context-as-substrate, per the CLAUDE.md aside above) and for how much weight to put on any single benchmark score.
Related
- verifiability
- technology-diffusion-llms
- agent-memory-systems
- 2025-llm-year-in-review
- andrej-karpathy
Open Questions
- Sutton’s “Bitter Lesson” itself is referenced repeatedly but has not been captured as raw evidence in this vault — currently represented only secondhand through Karpathy’s summary of it.
- Will ghosts and animals “converge” (per Karpathy’s speculation) as finetuning techniques improve, or diverge permanently? Presented as genuinely open by the source.
Sources
^[raw/feeds/https-karpathy-bearblog-dev-the-space-of-minds-0ee8301f13fcc405.md] ^[raw/feeds/https-karpathy-bearblog-dev-animals-vs-ghosts-80f0901d56fc251c.md]