Key ideas
- The one idea — Context-extension tricks rescale RoPE's angles to stretch the window — but advertised context ≠ usable context, and the gap is measurable.
- How it is shown — The stretched ribbon accepting 128k tokens; then the honesty tests: needle-recall strong at edges, sagging at depth and middle; the banner vs the measured usable zone.
- The trap to avoid — Standardizing a workflow on the banner number — test retrieval at YOUR lengths and depths; place critical material at the edges.
- What it sets up — You can now judge models by machinery — next: judge what 'open' means.
Paste a huge document into an AI, and the middle is where facts go to die — measured, replicated, and never mentioned in the ads.
The one idea
Context-extension tricks rescale RoPE's angles to stretch the window — but advertised context ≠ usable context, and the gap is measurable.
The banner says a million tokens. The fine print no one prints says: reliable to a fraction, depending where your facts sit and what you ask of them. Advertised and usable context are different numbers. Today: how the stretch works, where it frays, and how to measure the window you're buying. The stretch works because of last episode's gift: position is angles, and angles rescale. Interpolation compresses the dial-scale so more positions fit the geometry the model knows — a hundred-twenty-eight thousand packed into dials woven for eight — then a light fine-tune settles the weave. Cheap, effective, and genuinely how short-trained models went long. The craft matured fast. Naive uniform compression squeezes every dial — including the fast hands keeping neighbors distinct, precision you can't afford to blur. YaRN scales each frequency by what it can afford: fast hands untouched, slow hands absorbing the stretch.
How it works — the demo
The stretched ribbon accepting 128k tokens; then the honesty tests: needle-recall strong at edges, sagging at depth and middle; the banner vs the measured usable zone.
Local crispness kept, document-scale extended. Tailoring, not pulling — and it ships in models you use. Now the fraying, measured. Retrieval is strongest near the edges and sags through the middle — "lost in the middle," documented and replicated. Task matters too: finding one fact survives far deeper than reasoning across several, which thins earlier. Usable context isn't one number — it's a curve, shaped by depth and task, sitting below the banner for nearly every model tested. So work the window you have. Critical things at the edges — instructions up front, questions at the end — the middle carrying bulk, not criticality. Staged passes over one heroic gulp for work that must be right. And calibrate once: hide needles in your own documents at your lengths; draw your private curve.
The trap to avoid
Standardizing a workflow on the banner number — test retrieval at YOUR lengths and depths; place critical material at the edges.
Why it matters — and what’s next
You can now judge models by machinery — next: judge what 'open' means.
Twenty minutes, and the banner never surprises you again. The trap: building on the banner. A workflow that stuffs the advertised window and trusts the output has planted its failures in the sagging middle, where dropped clauses return in perfect confidence. The fix costs an afternoon: measure, then build on your curve. The spec is a claim. Your measurement is the fact. And with that, you can judge the machinery — experts, routers, caches, clocks, and stretched windows. The act's next literacy is legal and economic: the marketplace. Every release now wears a label — open weights, open source, API-only — and two of those labels rarely mean what they say. What you can download, audit, and legally ship: next group.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:45 of narration
The banner says a million tokens. The fine print no one prints says: reliable to a fraction, depending where your facts sit and what you ask of them. Advertised and usable context are different numbers. Today: how the stretch works, where it frays, and how to measure the window you're buying.
The stretch works because of last episode's gift: position is angles, and angles rescale. Interpolation compresses the dial-scale so more positions fit the geometry the model knows — a hundred-twenty-eight thousand packed into dials woven for eight — then a light fine-tune settles the weave. Cheap, effective, and genuinely how short-trained models went long.
The craft matured fast. Naive uniform compression squeezes every dial — including the fast hands keeping neighbors distinct, precision you can't afford to blur. YaRN scales each frequency by what it can afford: fast hands untouched, slow hands absorbing the stretch. Local crispness kept, document-scale extended. Tailoring, not pulling — and it ships in models you use.
Now the fraying, measured. Retrieval is strongest near the edges and sags through the middle — "lost in the middle," documented and replicated. Task matters too: finding one fact survives far deeper than reasoning across several, which thins earlier. Usable context isn't one number — it's a curve, shaped by depth and task, sitting below the banner for nearly every model tested.
So work the window you have. Critical things at the edges — instructions up front, questions at the end — the middle carrying bulk, not criticality. Staged passes over one heroic gulp for work that must be right. And calibrate once: hide needles in your own documents at your lengths; draw your private curve. Twenty minutes, and the banner never surprises you again.
The trap: building on the banner. A workflow that stuffs the advertised window and trusts the output has planted its failures in the sagging middle, where dropped clauses return in perfect confidence. The fix costs an afternoon: measure, then build on your curve. The spec is a claim. Your measurement is the fact.
And with that, you can judge the machinery — experts, routers, caches, clocks, and stretched windows. The act's next literacy is legal and economic: the marketplace. Every release now wears a label — open weights, open source, API-only — and two of those labels rarely mean what they say. What you can download, audit, and legally ship: next group.