Act 05 · Architectures & the model zoo 2:30 Model cards and the family tree

How to read a model card without getting fooled.

A model card mixes verifiable facts with vendor-run claims — separate them, demand the conditions, and trust independent replication.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — A model card mixes verifiable facts with vendor-run claims — separate them, demand the conditions, and trust independent replication.
  • How it is shown — A card dissected into two inks: facts (checkable from the files) vs claims (benchmark tables with tunable conditions); the red-flag checklist glowing.
  • The trap to avoid — Reading the benchmark table as measurement rather than marketing — the conditions column is where the game lives.
  • What it sets up — Every card descends from somewhere.

That benchmark table in every AI release? The conditions column is where the game lives — their best settings versus everyone else's defaults.

The one idea

A model card mixes verifiable facts with vendor-run claims — separate them, demand the conditions, and trust independent replication.

Every release arrives wearing a model card — part spec sheet, part résumé. And like a résumé, it mixes two inks: facts you can verify, and claims arranged to flatter. Telling them apart is the last literacy this act owes you. The auditor's lamp, on. The silver ink: facts checkable from the artifact. Parameter counts — both bars, episode one-twenty-three — architecture, context length, license: all verifiable against the files themselves, the manifest from episode one hundred. Honest cards make this easy. Any card where the prose and the config disagree has told you everything already. The shimmer: the benchmark table. Vendor-run, and every number wears hidden conditions — how many examples shown, how many attempts sampled, which harness, and crucially, whether rivals were measured under the same settings. The classic move is their-best versus others'-default. The table isn't false. It's curated — episode one-sixteen's asterisk-shake, now applied to software.

How it works — the demo

A card dissected into two inks: facts (checkable from the files) vs claims (benchmark tables with tunable conditions); the red-flag checklist glowing.

The red flags. Comparison rows under mismatched conditions. An unnamed harness — can't rerun it, it's a story. No contamination note, when episode one-forty-nine shows why that silence matters. Boilerplate limitations. And the obvious rival, conspicuously absent. Each is a tell. Two is a verdict. And the green flags. A named public harness — rerun me, it says. Conditions printed beside every number. A volunteered contamination analysis. Limitations specific enough to be slightly embarrassing — the surest sign a human told the truth.

The trap to avoid

Reading the benchmark table as measurement rather than marketing — the conditions column is where the game lives.

Why it matters — and what’s next

Every card descends from somewhere.

The best cards invite their own audit, and that tells you about everything else the lab does. The trap: hiring on the résumé. The card opens diligence; it never closes it. Cross-check independent leaderboards under one harness. Watch for community replications early. And hold final judgment for the only reference that matters: your own evaluation, on your own work — its own episode soon. Résumés start conversations; references end them. One thing every card claims without saying so: a pedigree. Because none of these models appeared from nowhere — nearly everything you've ever used descends from one small experiment published in twenty-nineteen. Reading cards is easier when you know the family. Next: the tree — GPT-2 to today.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:30 of narration

Every release arrives wearing a model card — part spec sheet, part résumé. And like a résumé, it mixes two inks: facts you can verify, and claims arranged to flatter. Telling them apart is the last literacy this act owes you. The auditor's lamp, on.

The silver ink: facts checkable from the artifact. Parameter counts — both bars, episode one-twenty-three — architecture, context length, license: all verifiable against the files themselves, the manifest from episode one hundred. Honest cards make this easy. Any card where the prose and the config disagree has told you everything already.

The shimmer: the benchmark table. Vendor-run, and every number wears hidden conditions — how many examples shown, how many attempts sampled, which harness, and crucially, whether rivals were measured under the same settings. The classic move is their-best versus others'-default. The table isn't false. It's curated — episode one-sixteen's asterisk-shake, now applied to software.

The red flags. Comparison rows under mismatched conditions. An unnamed harness — can't rerun it, it's a story. No contamination note, when episode one-forty-nine shows why that silence matters. Boilerplate limitations. And the obvious rival, conspicuously absent. Each is a tell. Two is a verdict.

And the green flags. A named public harness — rerun me, it says. Conditions printed beside every number. A volunteered contamination analysis. Limitations specific enough to be slightly embarrassing — the surest sign a human told the truth. The best cards invite their own audit, and that tells you about everything else the lab does.

The trap: hiring on the résumé. The card opens diligence; it never closes it. Cross-check independent leaderboards under one harness. Watch for community replications early. And hold final judgment for the only reference that matters: your own evaluation, on your own work — its own episode soon. Résumés start conversations; references end them.

One thing every card claims without saying so: a pedigree. Because none of these models appeared from nowhere — nearly everything you've ever used descends from one small experiment published in twenty-nineteen. Reading cards is easier when you know the family. Next: the tree — GPT-2 to today.

Open ModelsEvals & TestingGovernance