Key ideas
- The one idea — The trajectory is "omni" models — one model natively handling text, image, audio, and video in a single shared stream — and this whole act reduces to one idea: every modality becomes tokens in one shared mind.
- How it is shown — Separate per-sense models merging into one native any-to-any model sharing a single residual stream.
- The trap to avoid — Thinking multimodal means "a text model with vision plugged in"; native omni, trained on all modalities together, is qualitatively different — and where things are heading.
- What it sets up — One model, every sense — but it still forgets you the instant you leave.
Most 'multimodal' AI is a text brain with senses bolted on after. The next generation is different: trained on every modality together, from the start.
The one idea
The trajectory is "omni" models — one model natively handling text, image, audio, and video in a single shared stream — and this whole act reduces to one idea: every modality becomes tokens in one shared mind.
Everything this act covered — vision, diffusion, video, audio — has been senses added to a text core. The endgame: not four models bolted together, but one model natively handling text, image, audio, and video in one mind. That's omni, where the field points. Omni means one model, any sense to any sense. Text, image, audio, video all enter and leave as tokens in a single shared stream — the one-mind idea from episode one-fifty-two, scaled to every sense. Any input becomes any output. The distinction that matters: native versus bolted-on. Early multimodal stapled a vision adapter onto a text brain — seamed, and seams are where failures slip in.
How it works — the demo
Separate per-sense models merging into one native any-to-any model sharing a single residual stream.
Native omni trains on all modalities from the start, so senses share representations deeply. A native tongue, not an add-on. Why the trajectory? The world is multimodal, so a fused model matches reality better than any single-sense system. In one shared space, understanding transfers across senses — reading reinforces seeing and hearing. Plus efficiency, plus new cross-modal skills. It all pulls one way. What it unlocks: assistants you talk to while they watch live video and answer in voice.
The trap to avoid
Thinking multimodal means "a text model with vision plugged in"; native omni, trained on all modalities together, is qualitatively different — and where things are heading.
Why it matters — and what’s next
One model, every sense — but it still forgets you the instant you leave.
And it compresses the whole act into one line — images, CLIP, VLMs, diffusion, video, audio, every piece the same idea: every modality becomes tokens in one shared mind. The trap: underestimating it as "a chatbot with a camera plugged in." Native omni is a qualitative shift, not an accessory — senses sharing one mind, so the model reasons across them in ways an adapter can't fake. Not a text model with plugins; one model thinking in every modality. So that closes the senses — sight, hearing, voice, one mind. But it's missing something enormous: memory. End the conversation and it vanishes, greeting you as a stranger. Fixing that — memory, context, agents — is the next act, and it starts where it must: the model has amnesia.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:15 of narration
Everything this act covered — vision, diffusion, video, audio — has been senses added to a text core. The endgame: not four models bolted together, but one model natively handling text, image, audio, and video in one mind. That's omni, where the field points.
Omni means one model, any sense to any sense. Text, image, audio, video all enter and leave as tokens in a single shared stream — the one-mind idea from episode one-fifty-two, scaled to every sense. Any input becomes any output.
The distinction that matters: native versus bolted-on. Early multimodal stapled a vision adapter onto a text brain — seamed, and seams are where failures slip in. Native omni trains on all modalities from the start, so senses share representations deeply. A native tongue, not an add-on.
Why the trajectory? The world is multimodal, so a fused model matches reality better than any single-sense system. In one shared space, understanding transfers across senses — reading reinforces seeing and hearing. Plus efficiency, plus new cross-modal skills. It all pulls one way.
What it unlocks: assistants you talk to while they watch live video and answer in voice. And it compresses the whole act into one line — images, CLIP, VLMs, diffusion, video, audio, every piece the same idea: every modality becomes tokens in one shared mind.
The trap: underestimating it as "a chatbot with a camera plugged in." Native omni is a qualitative shift, not an accessory — senses sharing one mind, so the model reasons across them in ways an adapter can't fake. Not a text model with plugins; one model thinking in every modality.
So that closes the senses — sight, hearing, voice, one mind. But it's missing something enormous: memory. End the conversation and it vanishes, greeting you as a stranger. Fixing that — memory, context, agents — is the next act, and it starts where it must: the model has amnesia.