Verifiability (Software 2.0 / RLVR)
Overview (confidence: medium)
Andrej Karpathy analogizes AI to a new computing paradigm rather than to historical precedents like electricity or the industrial revolution, because both computing and AI are fundamentally about automating digital information processing. In the 1980s computing era, the most predictive feature of whether a job would be automated was specifiability: can the task be reduced to a rote, hand-writable algorithm? That era’s programs are what Karpathy calls “Software 1.0.”
With AI, we now write “Software 2.0” programs we could never hand-write: we specify an objective (classification accuracy, a reward function) and search the program space via gradient descent to find a neural network that satisfies it. In this new paradigm, the most predictive feature of automatability is verifiability — can an AI “practice” the task? A verifiable task/environment must be resettable (a new attempt can start), efficient (many attempts are cheap), and rewardable (an automated process can score any given attempt).
The Jagged Frontier (confidence: medium)
Verifiable tasks (math, code, puzzle-like domains with a checkable correct answer) can be optimized directly via reinforcement learning and progress rapidly — sometimes past top human-expert level. Non-verifiable tasks (creative work, strategy, tasks combining real-world knowledge, state, context, and common sense) cannot be optimized this directly; they must fall out of weaker imitation or the model’s general capacity to generalize. This asymmetry is what produces the “jagged” capability frontier observed across LLMs: genius in verifiable domains, unreliable elsewhere.
RLVR: Verifiability Operationalized in 2025 (confidence: medium)
Karpathy’s 2025 retrospective names Reinforcement Learning from Verifiable Rewards (RLVR) as the year’s defining new training stage, added on top of the prior stable recipe of Pretraining → Supervised Finetuning → RLHF. By training against automatically verifiable rewards across environments like math and code puzzles, LLMs spontaneously develop strategies that look like “reasoning”: breaking problems into intermediate steps, backtracking, and self-correction (see DeepSeek-R1 for a documented example). Unlike SFT/RLHF, which are relatively short finetuning stages, RLVR trains against objective, non-gameable reward functions that support much longer optimization runs — and introduced a new scaling knob: capability as a function of test-time “thinking time.” OpenAI’s o1 (late 2024) was the first RLVR demonstration; o3 (early 2025) was the point where the capability jump became intuitively obvious.
A downstream consequence: because benchmarks are themselves verifiable environments by construction, they became immediately susceptible to RLVR-style optimization and synthetic-data “benchmaxxing” — training directly on test-set-adjacent distributions became, in Karpathy’s words, “a new art form,” eroding trust in benchmark scores as a proxy for general capability.
Implications: Verifiability is a more actionable predictor of near-term AI automation than raw model scale — a task’s tractability to automation can be assessed by asking whether it can be made resettable, efficient, and rewardable, independent of how “hard” it looks to a human. This also explains why benchmark-driven progress narratives became unreliable in 2025: the same property that makes a domain automatable also makes it gameable.
Related
- technology-diffusion-llms
- ghosts-vs-animals
- fine-tuning
- 2025-llm-year-in-review
- andrej-karpathy
- sequoia-ascent-2026
Open Questions
- No independently-authored source on RLVR has been captured into this vault yet (e.g. the DeepSeek-R1 paper referenced by Karpathy) — both sources here are Karpathy’s own writing.
- How stable is the specifiability/verifiability framing against domains that are partially verifiable (e.g. long-horizon agentic tasks with sparse, delayed rewards)?
Sources
^[raw/feeds/https-karpathy-bearblog-dev-verifiability-576bce59c7e3a574.md]