Act 10 · Our systems 2:45 Second brains, the Cortex, the ledger

Your reviewer runs on a different brain.

The worker never grades its own homework. R1 uses cross-model adversarial review: if Claude implemented, Codex reviews, and vice versa — reading the diff, the acceptance criteria, and the tool-call log. It works where "a second AI checks the first" fails, because it's cross-model and anchored to mechanical criteria, not vibes.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — The worker never grades its own homework. R1 uses cross-model adversarial review: if Claude implemented, Codex reviews, and vice versa — reading the diff, the acceptance criteria, and the tool-call log. It works where "a second AI checks the first" fails, because it's cross-model and anchored to mechanical criteria, not vibes.
  • How it is shown — A finished diff handed across the table from the model that wrote it to a different model that reviews it — against the ACs, not on feel.
  • The trap to avoid — Assuming this is the doomed "AI judges AI" setup. It isn't: a different model has different blind spots, and the review is anchored to acceptance criteria — a grounded check, not a shared-bias rubber stamp.
  • What it sets up — One reviewer on a different brain is powerful — but what if many specialist minds could reason in parallel over one shared context? the Cortex.

AI checking AI is supposed to fail — shared blind spots. But swap in a rival model from a different company and anchor it to hard criteria, and review starts working.

The one idea

The worker never grades its own homework. R1 uses cross-model adversarial review: if Claude implemented, Codex reviews, and vice versa — reading the diff, the acceptance criteria, and the tool-call log. It works where "a second AI checks the first" fails, because it's cross-model and anchored to mechanical criteria, not vibes.

The worker never grades its own homework. In R1, a rival model does. This is cross-model review: if Claude implemented the change, Codex reviews it, and vice versa — reading the diff, the acceptance criteria, and the full tool-call log of how the work was done. A different brain, checking a different brain's work. But wait — didn't the last act prove this fails? "A second AI checks the first" was doomed: a verifier sharing the generator's blind spots, rubber-stamping the same errors. BS checking BS. So why would R1 stake its trust on that? Because R1 fixes the two things that broke the naive version — precisely the two the last act told us to fix. Fix one: a genuinely different brain.

How it works — the demo

A finished diff handed across the table from the model that wrote it to a different model that reviews it — against the ACs, not on feel.

The naive verifier failed because it was a sibling — same training, same architecture, same blind spots. R1 makes the reviewer a different model, blind to different things, so its mistakes and the worker's don't line up. It sees what the worker couldn't. Fix two: anchor the review to criteria, not vibes. The naive judge floated on a subjective sense of quality — easy to share, easy to fool. R1's reviewer grades against the SOW's mechanical acceptance criteria, not "looks good to me." That's the last act's rule — ground the check outside the models — made concrete. Why both? A different brain grading on vibes just disagrees by feel. Shared criteria judged by the same brain still share its blind spots. You need both: independent eyes and an objective target.

The trap to avoid

Assuming this is the doomed "AI judges AI" setup. It isn't: a different model has different blind spots, and the review is anchored to acceptance criteria — a grounded check, not a shared-bias rubber stamp.

Why it matters — and what’s next

One reviewer on a different brain is powerful — but what if many specialist minds could reason in parallel over one shared context? the Cortex.

Together they turn "a second AI checks the first" into a real check — the last act's hardest problem, answered in the runtime — reduced, not eliminated. The trap: dismissing this as the doomed "AI judges AI" setup. It isn't — it fixes the two things that broke it: a different brain for independent blind spots, and an anchor to criteria for a grounded target. Whether AI checking AI works isn't a slogan; it's a question of how it's wired. So R1's reviewer runs on a different brain and grades against mechanical criteria — the built answer to "why a second AI fails." Cross-model gives it different blind spots; the criteria keep it grounded. One outside mind is powerful. But what if many specialist minds reasoned in parallel over one shared context? Next: six minds in one context — the Cortex.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:45 of narration

The worker never grades its own homework. In R1, a rival model does. This is cross-model review: if Claude implemented the change, Codex reviews it, and vice versa — reading the diff, the acceptance criteria, and the full tool-call log of how the work was done. A different brain, checking a different brain's work.

But wait — didn't the last act prove this fails? "A second AI checks the first" was doomed: a verifier sharing the generator's blind spots, rubber-stamping the same errors. BS checking BS. So why would R1 stake its trust on that? Because R1 fixes the two things that broke the naive version — precisely the two the last act told us to fix.

Fix one: a genuinely different brain. The naive verifier failed because it was a sibling — same training, same architecture, same blind spots. R1 makes the reviewer a different model, blind to different things, so its mistakes and the worker's don't line up. It sees what the worker couldn't.

Fix two: anchor the review to criteria, not vibes. The naive judge floated on a subjective sense of quality — easy to share, easy to fool. R1's reviewer grades against the SOW's mechanical acceptance criteria, not "looks good to me." That's the last act's rule — ground the check outside the models — made concrete.

Why both? A different brain grading on vibes just disagrees by feel. Shared criteria judged by the same brain still share its blind spots. You need both: independent eyes and an objective target. Together they turn "a second AI checks the first" into a real check — the last act's hardest problem, answered in the runtime — reduced, not eliminated.

The trap: dismissing this as the doomed "AI judges AI" setup. It isn't — it fixes the two things that broke it: a different brain for independent blind spots, and an anchor to criteria for a grounded target. Whether AI checking AI works isn't a slogan; it's a question of how it's wired.

So R1's reviewer runs on a different brain and grades against mechanical criteria — the built answer to "why a second AI fails." Cross-model gives it different blind spots; the criteria keep it grounded. One outside mind is powerful. But what if many specialist minds reasoned in parallel over one shared context? Next: six minds in one context — the Cortex.

Our SystemsR1Cross-Model Review