In 2024 we learned to talk to models. In 2025 we learned to feed them. In 2026 we're learning to build the machinery around them — and to ask whether language is even the right substrate for the next leap.
The story of the last four years is a progressive migration of "where the intelligence lives" — away from the model itself and into the layers wrapped around it. Each ring below is a discipline that emerged when the previous one hit a ceiling. Click a ring to explore it.
The pretrained weights. In 2026's framing, the model is increasingly treated as a "frozen utility" — a reasoning engine you don't retrain, you engineer around. Everything interesting happens in the rings.
"The recent history of LLM agents can be understood as a progressive movement outward from the model itself."
Click any ring on the diagram
Loop, harness, context, MCP: good instincts, all four are the headline acts of 2026. Here's the full radar, including six you didn't mention. Tap any card to expand.
A worked example our audience knows well: "Generate the daily portfolio-drift exceptions report and draft RM talking points." Step through how each 2026 discipline shows up in a single agent run.
A July 2026 snapshot of the major families. Gold border = current flagship. Prices are per million tokens (input / output) where public, and change often — treat as directional.
The single biggest strategic debate in AI right now. LLMs predict the next word; world models predict the next state of reality — learning physics, space, and cause-and-effect from video and interaction rather than text. Toggle below to feel the difference.
Given a sequence of tokens, the model outputs a probability distribution over the next one. Astonishingly powerful — but the model's grasp of physics, space, and time is second-hand, inferred from descriptions of the world rather than the world itself. That's why LLMs can ace an exam question about gravity yet "hallucinate physics" in a dynamic scene.
Trained on video and interaction, a world model builds an internal simulator: give it a scene and an action, and it predicts how the environment evolves. That unlocks the things language struggles with — planning, counterfactuals ("what if I grab it?"), and training agents safely "in imagination" before they act in the real world.
The model never read the word "gravity". It watched enough video to expect the arc.
Three arguments, in ascending order of ambition: (1) Robotics & physical AI — agents that act in the world need to predict consequences of actions, and simulators generate infinite safe training data. (2) Economics — as frontier LLMs commoditise and margins compress, investors are hunting the next defensible moat; over $3B of venture capital flowed to world-model startups in H1 2026 alone. (3) The AGI bet — Yann LeCun's camp argues text is a lossy shadow of reality and scaling LLMs will never yield general intelligence; his JEPA-style architectures predict in abstract representation space rather than pixels or tokens. The counterpoint: LLM labs argue reasoning + tools + loops gets there first. Both camps are now funded like they're right.