2025 LLM Year in Review (Karpathy)

Overview (confidence: medium)

Andrej Karpathy’s own retrospective lists six “paradigm changes” he found personally notable in 2025. Three of them — RLVR, jagged intelligence, and (implicitly) the diffusion pattern from “Power to the people” — are covered in dedicated pages this synthesis links out to: verifiability and ghosts-vs-animals and technology-diffusion-llms. This page anchors the full six-item retrospective and covers the three additional 2025 shifts that don’t yet have their own dedicated page.

The Six Paradigm Changes (confidence: medium)

  1. RLVR (Reinforcement Learning from Verifiable Rewards) — the new major training stage added after pretraining/SFT/RLHF; see verifiability for full detail.
  2. Ghosts vs. Animals / Jagged Intelligence — the industry’s growing intuition that LLM intelligence is not an immature version of animal intelligence; see ghosts-vs-animals for full detail.
  3. Cursor / a new layer of “LLM apps” — Cursor’s rise revealed a distinct application layer that bundles and orchestrates LLM calls for a specific vertical: doing context engineering, orchestrating multiple LLM calls into cost/performance-balanced DAGs, providing a vertical-specific GUI, and offering an “autonomy slider.” Karpathy frames the open question as whether LLM labs will absorb this layer or whether “LLM apps” retain durable value by supplying private data, sensors, actuators, and feedback loops that labs don’t have.
  4. Claude Code / AI that lives on your computer — the first convincing demonstration of an LLM agent looping tool-use and reasoning for extended problem solving, notable specifically because it runs locally, with the user’s own environment, data, and context, rather than in a cloud container. Karpathy credits Anthropic with getting the precedence right (local-first, minimal CLI) versus OpenAI’s early cloud-container-first Codex/agent efforts, framing Claude Code as “a little spirit/ghost that lives on your computer” — a distinct interaction paradigm from “a website you go to.”
  5. Vibe coding — 2025 crossed the capability threshold for building real software via natural language alone; Karpathy coined the term “vibe coding” in an offhand tweet earlier in the year, “totally oblivious to how far it would go.” He connects this directly back to technology-diffusion-llms: vibe coding is a concrete instance of regular people (not just professionals) gaining a capability, while also letting trained professionals ship far more software than they otherwise would (his own examples: nanochat’s custom Rust BPE tokenizer, menugen, llm-council, reader3, the HN time-capsule project).
  6. Nano Banana / an early “LLM GUI” — Google’s Gemini “Nano Banana” model is framed as an early hint of the graphical interface layer LLMs will eventually need, analogous to how personal computing needed the GUI on top of the command line. Karpathy’s argument: text is the favored representation for computers (and LLMs) but people prefer visual/spatial information, so LLMs should output images, infographics, slides, and interactive artifacts — not just chat text — and Nano Banana’s notable feature is that this capability is jointly encoded with text generation and world knowledge in one model, not bolted on separately.

Closing Stance (confidence: medium)

Karpathy’s overall take: 2025’s LLMs are “simultaneously a lot smarter and a lot dumber” than he expected going in, current capability is nowhere near 10% exploited by the industry, and he expects both continued rapid progress and a large amount of remaining unsolved work to coexist — a view he also expressed on the Dwarkesh podcast earlier in the year.

Implications: This retrospective is useful as a single index into Karpathy’s 2025 conceptual output — four of its six items already exist as independently-sourced concept pages in this vault; the other two (the “LLM app layer” via Cursor, and Nano Banana as an early “LLM GUI”) are candidates for their own dedicated pages if a future source discusses either in more depth than a one-paragraph mention here.

Open Questions

  • Cursor / “LLM app layer” and Nano Banana / “LLM GUI” are each covered only in this one paragraph-level mention — not yet substantial enough for their own concept pages per the 2+-source or “central to one source” threshold in _schema.md.
  • This is Karpathy’s own single-author retrospective; no independent industry source has been captured to corroborate or contest any of the six claims.

Sources

^[raw/feeds/https-karpathy-bearblog-dev-year-in-review-2025-604f493848c47ec0.md] ^[raw/feeds/https-karpathy-bearblog-dev-power-to-the-people-28a4b3a1adbe15c8.md] ^[raw/feeds/https-karpathy-bearblog-dev-verifiability-576bce59c7e3a574.md] ^[raw/feeds/https-karpathy-bearblog-dev-the-space-of-minds-0ee8301f13fcc405.md]