Act 05 · Architectures & the model zoo 2:45 Position and long context: RoPE, YaRN

Advertised vs. usable context (YaRN).

Context-extension tricks rescale RoPE's angles to stretch the window — but advertised context ≠ usable context, and the gap is measurable.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — Context-extension tricks rescale RoPE's angles to stretch the window — but advertised context ≠ usable context, and the gap is measurable.
  • How it is shown — The stretched ribbon accepting 128k tokens; then the honesty tests: needle-recall strong at edges, sagging at depth and middle; the banner vs the measured usable zone.
  • The trap to avoid — Standardizing a workflow on the banner number — test retrieval at YOUR lengths and depths; place critical material at the edges.
  • What it sets up — You can now judge models by machinery — next: judge what 'open' means.

Paste a huge document into an AI, and the middle is where facts go to die — measured, replicated, and never mentioned in the ads.

The one idea

Context-extension tricks rescale RoPE's angles to stretch the window — but advertised context ≠ usable context, and the gap is measurable.

The banner says a million tokens. The fine print no one prints says: reliable to a fraction, depending where your facts sit and what you ask of them. Advertised and usable context are different numbers. Today: how the stretch works, where it frays, and how to measure the window you're buying. The stretch works because of last episode's gift: position is angles, and angles rescale. Interpolation compresses the dial-scale so more positions fit the geometry the model knows — a hundred-twenty-eight thousand packed into dials woven for eight — then a light fine-tune settles the weave. Cheap, effective, and genuinely how short-trained models went long. The craft matured fast. Naive uniform compression squeezes every dial — including the fast hands keeping neighbors distinct, precision you can't afford to blur. YaRN scales each frequency by what it can afford: fast hands untouched, slow hands absorbing the stretch.

How it works — the demo

The stretched ribbon accepting 128k tokens; then the honesty tests: needle-recall strong at edges, sagging at depth and middle; the banner vs the measured usable zone.

Local crispness kept, document-scale extended. Tailoring, not pulling — and it ships in models you use. Now the fraying, measured. Retrieval is strongest near the edges and sags through the middle — "lost in the middle," documented and replicated. Task matters too: finding one fact survives far deeper than reasoning across several, which thins earlier. Usable context isn't one number — it's a curve, shaped by depth and task, sitting below the banner for nearly every model tested. So work the window you have. Critical things at the edges — instructions up front, questions at the end — the middle carrying bulk, not criticality. Staged passes over one heroic gulp for work that must be right. And calibrate once: hide needles in your own documents at your lengths; draw your private curve.

The trap to avoid

Standardizing a workflow on the banner number — test retrieval at YOUR lengths and depths; place critical material at the edges.

Why it matters — and what’s next

You can now judge models by machinery — next: judge what 'open' means.

Twenty minutes, and the banner never surprises you again. The trap: building on the banner. A workflow that stuffs the advertised window and trusts the output has planted its failures in the sagging middle, where dropped clauses return in perfect confidence. The fix costs an afternoon: measure, then build on your curve. The spec is a claim. Your measurement is the fact. And with that, you can judge the machinery — experts, routers, caches, clocks, and stretched windows. The act's next literacy is legal and economic: the marketplace. Every release now wears a label — open weights, open source, API-only — and two of those labels rarely mean what they say. What you can download, audit, and legally ship: next group.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:45 of narration

The banner says a million tokens. The fine print no one prints says: reliable to a fraction, depending where your facts sit and what you ask of them. Advertised and usable context are different numbers. Today: how the stretch works, where it frays, and how to measure the window you're buying.

The stretch works because of last episode's gift: position is angles, and angles rescale. Interpolation compresses the dial-scale so more positions fit the geometry the model knows — a hundred-twenty-eight thousand packed into dials woven for eight — then a light fine-tune settles the weave. Cheap, effective, and genuinely how short-trained models went long.

The craft matured fast. Naive uniform compression squeezes every dial — including the fast hands keeping neighbors distinct, precision you can't afford to blur. YaRN scales each frequency by what it can afford: fast hands untouched, slow hands absorbing the stretch. Local crispness kept, document-scale extended. Tailoring, not pulling — and it ships in models you use.

Now the fraying, measured. Retrieval is strongest near the edges and sags through the middle — "lost in the middle," documented and replicated. Task matters too: finding one fact survives far deeper than reasoning across several, which thins earlier. Usable context isn't one number — it's a curve, shaped by depth and task, sitting below the banner for nearly every model tested.

So work the window you have. Critical things at the edges — instructions up front, questions at the end — the middle carrying bulk, not criticality. Staged passes over one heroic gulp for work that must be right. And calibrate once: hide needles in your own documents at your lengths; draw your private curve. Twenty minutes, and the banner never surprises you again.

The trap: building on the banner. A workflow that stuffs the advertised window and trusts the output has planted its failures in the sagging middle, where dropped clauses return in perfect confidence. The fix costs an afternoon: measure, then build on your curve. The spec is a claim. Your measurement is the fact.

And with that, you can judge the machinery — experts, routers, caches, clocks, and stretched windows. The act's next literacy is legal and economic: the marketplace. Every release now wears a label — open weights, open source, API-only — and two of those labels rarely mean what they say. What you can download, audit, and legally ship: next group.

ArchitecturesLong ContextEvals & Testing