Summary of Andrej Karpathy’s Fireside Chat at Sequoia Ascent 2026
Andrej Karpathy’s talk at Sequoia Ascent 2026 covered eleven key themes about the evolving landscape of AI-assisted software development and information processing. Here is a synthesized overview:
1. December 2025 — Agentic Inflection Point
Around December 2025, Karpathy noticed a step change in agent reliability. Generated code chunks became larger, more coherent, and more reliable. The unit of programming shifted from typing lines of code to delegating larger “macro actions” (implement features, refactor subsystems, research libraries, set up services, write tests). The programmer is increasingly an “orchestrator of agents.”
2. Software 3.0: The Context Window as the New Program
Software evolution sequence:
- Software 1.0: Humans write explicit code
- Software 2.0: Humans create datasets, objectives, neural networks; program learned into weights
- Software 3.0: Humans program LLMs through prompts, context, tools, examples, memory, and instructions. The context window becomes the main lever; the LLM is an interpreter over that context.
Examples: Software 3.0 installers (like OpenClaw) are blocks of text agents read and adapt to local environment; much more adaptive than brittle shell scripts.
3. MenuGen and the Moment Software Disappears
MenuGen comparison:
- Traditional web app: Take menu picture → OCR → Generate images via APIs → Render in UI. Required frontend code, APIs, image generation, deployment, auth, payments, secrets, infrastructure.
- Software 3.0 version: Take photo → Give to multimodal model → Render dish images directly onto menu image. Much of the app disappears; the neural network directly transforms input media into output media.
Key implication: Some apps should stop existing as apps. AI is not just a faster way to build old apps; some apps should be replaced by direct model transformations.
4. The New Opportunity Is Not Just Faster Programming
LLMs automate forms of information processing that were not previously programmable. The LLM Wiki pattern is the clearest example: instead of RAG (retrieval-augmented generation) to answer questions from raw documents each time, an agent incrementally compiles raw sources into a persistent Markdown wiki (summaries, entity pages, concept pages, contradictions, cross-links, logs, evolving synthesis). No classical program could robustly maintain that kind of knowledge base, but an LLM can.
Lesson: Ask not only “What existing workflow can AI speed up?” but also “What information transformation was impossible before, but is now natural?“
5. Verifiability Explains Where AI Moves Fastest
Core framework: Traditional software automates what you can specify; LLMs and reinforcement learning automate what you can verify. If a task has an automatic reward or success signal, models can practice it. This explains why math, coding, tests, benchmarks, games, and engineering tasks improve quickly — they are resettable, repeatable, and rewardable.
Also explains why coding agents feel better than ordinary chatbots: coding gives feedback (tests pass/fail, programs run/crash, diffs can be inspected).
6. Jagged Intelligence Has Two Axes: Verifiability and Training Attention
Model capability depends not only on verifiability but also on whether the task was emphasized during training, post-training, synthetic data generation, and RL. Rough formula: capability spike ≈ verifiability × training attention × data coverage × economic value.
Example: Chess improvement from GPT-3.5 to GPT-4 was partly because much more chess data was included in the training mix. Someone at OpenAI decided to add that data, creating a capability spike.
Practical question for founders: Are you on the model’s rails? If your task is in a verifiable, heavily-trained region, the model may fly. If not, it may fail in basic ways. You may need better context, tools, fine-tuning, your own evals, or your own RL environment.
7. Vibe Coding vs. Agentic Engineering
- Vibe coding: Raises the floor. Lets almost anyone create software by describing what they want. Fine for prototypes and personal tools.
- Agentic engineering: Raises the ceiling. Professional discipline of coordinating fallible agents while preserving correctness, security, taste, and maintainability.
Hiring should test agentic engineering directly: build a substantial project with agents, deploy it, have adversarial agents try to break it. The old “10x engineer” may become much more extreme — people who master agentic workflows may outperform others by far more than 10x.
8. Hiring Should Change
If agentic engineering is the new professional skill, hiring should test it: give candidates a big project, have them implement it, then have agents simulate activity and try to break it. Tests: Can the candidate decompose work for agents? Write useful specs? Preserve quality while moving fast? Review generated work? Secure and harden the system? Use agents as leverage rather than produce slop?
9. Founders Should Look for Valuable Verifiable Environments
For founders, one important opportunity is finding domains that are valuable, verifiable, and undertrained by frontier labs. If you can create a domain-specific environment where models can try actions and receive reliable rewards, you may improve performance with fine-tuning or RL even if the base model is not already excellent there.
Most obvious domains (coding, math) are already heavily targeted. But many economically important domains may have latent verifiable structure not yet exploited — that’s a startup wedge.
10. Agent-Native Infrastructure: Build for the Agent, Not Just the Human
Most software is still built for humans clicking through screens. Agent-native surfaces needed: Markdown docs, CLIs, APIs, MCP servers, structured logs, machine-readable schemas, copy-pasteable agent instructions, safe permissioning, auditable actions, headless setup flows.
Sensors turn state of world into digital information; actuators let agents change something. The future stack: agents using sensors and actuators on behalf of people and organizations.
MenuGen deployment remains a benchmark: building the app was easy compared to wiring Vercel, auth, payments, DNS, secrets, and production config. In a mature agent-native world, you should be able to say “build MenuGen” and have the agent deploy the whole thing without manual clicking.
11. Ghosts, Not Animals
LLMs are not animals — they do not have biological drives, embodied survival pressure, curiosity, play, or intrinsic motivation shaped by evolution. They are statistical simulations of human artifacts, shaped by pretraining, post-training, RL, product feedback, and economic incentives.
Right posture: neither dismissal nor blind trust. It is empirical familiarity: learn where they work, where they fail, what they were trained for, and how to build guardrails. The right posture is suspicious and empirical over time.
Related
- context-engineering — shaping the full context an agent sees
- agent-loop — the repeated control loop that agents run in
- verifiability — the main reason some tasks improve fastest
- llm-as-os — the broader agent-native framing behind the app example
External Links
Sources
raw/external/docs-getsemantica-ai-c0a0b760.md
Education: You Can Outsource Thinking, But Not Understanding
Key line: “You can outsource your thinking, but you can’t outsource your understanding.” Even if agents do more work, the human still needs understanding to direct them. You need to know what is worth building, what question matters, what result is suspicious, and what tradeoff is acceptable.
LLM knowledge bases are tools for transforming information into understanding. They are not just answer machines. The human expert contributes the distilled artifact and the taste behind it; the agent can then explain it interactively to each learner.
Main Thesis
AI is becoming a new operating layer for digital work. The scarce thing is shifting:
- Less scarce: code generation, API recall, boilerplate, first drafts, repetitive setup, simple transformations
- More scarce: understanding, taste, eval design, security, system boundaries, agent orchestration, domain-specific feedback loops, knowing when the model is off the rails
For founders, the most important questions:
- What becomes possible when the primary user is an agent acting for a human?
- What workflows can be rebuilt around sensors, actuators, and verifiable loops?
- What software should disappear into direct model transformations?
- What domains are valuable and verifiable but not yet heavily trained by frontier labs?
- What human judgment must remain in the loop to preserve quality?
Implications for Founders
-
Software reorganization: Work is being reorganized around agents. Software, research, education, infrastructure, and knowledge work are all becoming variations of the same pattern: define context → define tools → define feedback loop → define guardrails → let agents work → preserve human understanding.
-
Valuable verifiable domains: Create domain-specific environments where models can try actions and receive reliable rewards. Even if not heavily trained by frontier labs, you can use fine-tuning or your own RL environments.
-
Human judgment remains: Taste, engineering judgment, design, and oversight remain uniquely human. Agents fill in the blanks, but humans design the spec and plan.
-
Agent-native infrastructure: Products need agent-native surfaces first. Don’t tell humans “go to this URL” or “click here” — describe things to agents first, build automation around data structures legible to LLMs.
-
Education enhancement: Tools that enhance understanding are incredibly interesting. LLM knowledge bases help both humans and agents gain insight through synthetic data generation over fixed data.
References
- Full transcript: Sequoia Ascent 2026 fireside chat with Stephanie Zhan
- Key concepts: Software 3.0, Agentic Engineering, Jagged Intelligence, Verifiability, Vibe Coding
- Related: LLM Wiki pattern, context window as program, agent-native infrastructure