Key ideas
- The one idea — The worker never grades its own homework. R1 uses cross-model adversarial review: if Claude implemented, Codex reviews, and vice versa — reading the diff, the acceptance criteria, and the tool-call log. It works where "a second AI checks the first" fails, because it's cross-model and anchored to mechanical criteria, not vibes.
- How it is shown — A finished diff handed across the table from the model that wrote it to a different model that reviews it — against the ACs, not on feel.
- The trap to avoid — Assuming this is the doomed "AI judges AI" setup. It isn't: a different model has different blind spots, and the review is anchored to acceptance criteria — a grounded check, not a shared-bias rubber stamp.
- What it sets up — One reviewer on a different brain is powerful — but what if many specialist minds could reason in parallel over one shared context? the Cortex.
AI checking AI is supposed to fail — shared blind spots. But swap in a rival model from a different company and anchor it to hard criteria, and review starts working.
The one idea
The worker never grades its own homework. R1 uses cross-model adversarial review: if Claude implemented, Codex reviews, and vice versa — reading the diff, the acceptance criteria, and the tool-call log. It works where "a second AI checks the first" fails, because it's cross-model and anchored to mechanical criteria, not vibes.
The worker never grades its own homework. In R1, a rival model does. This is cross-model review: if Claude implemented the change, Codex reviews it, and vice versa — reading the diff, the acceptance criteria, and the full tool-call log of how the work was done. A different brain, checking a different brain's work. But wait — didn't the last act prove this fails? "A second AI checks the first" was doomed: a verifier sharing the generator's blind spots, rubber-stamping the same errors. BS checking BS. So why would R1 stake its trust on that? Because R1 fixes the two things that broke the naive version — precisely the two the last act told us to fix. Fix one: a genuinely different brain.
How it works — the demo
A finished diff handed across the table from the model that wrote it to a different model that reviews it — against the ACs, not on feel.
The naive verifier failed because it was a sibling — same training, same architecture, same blind spots. R1 makes the reviewer a different model, blind to different things, so its mistakes and the worker's don't line up. It sees what the worker couldn't. Fix two: anchor the review to criteria, not vibes. The naive judge floated on a subjective sense of quality — easy to share, easy to fool. R1's reviewer grades against the SOW's mechanical acceptance criteria, not "looks good to me." That's the last act's rule — ground the check outside the models — made concrete. Why both? A different brain grading on vibes just disagrees by feel. Shared criteria judged by the same brain still share its blind spots. You need both: independent eyes and an objective target.
The trap to avoid
Assuming this is the doomed "AI judges AI" setup. It isn't: a different model has different blind spots, and the review is anchored to acceptance criteria — a grounded check, not a shared-bias rubber stamp.
Why it matters — and what’s next
One reviewer on a different brain is powerful — but what if many specialist minds could reason in parallel over one shared context? the Cortex.
Together they turn "a second AI checks the first" into a real check — the last act's hardest problem, answered in the runtime — reduced, not eliminated. The trap: dismissing this as the doomed "AI judges AI" setup. It isn't — it fixes the two things that broke it: a different brain for independent blind spots, and an anchor to criteria for a grounded target. Whether AI checking AI works isn't a slogan; it's a question of how it's wired. So R1's reviewer runs on a different brain and grades against mechanical criteria — the built answer to "why a second AI fails." Cross-model gives it different blind spots; the criteria keep it grounded. One outside mind is powerful. But what if many specialist minds reasoned in parallel over one shared context? Next: six minds in one context — the Cortex.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:45 of narration
The worker never grades its own homework. In R1, a rival model does. This is cross-model review: if Claude implemented the change, Codex reviews it, and vice versa — reading the diff, the acceptance criteria, and the full tool-call log of how the work was done. A different brain, checking a different brain's work.
But wait — didn't the last act prove this fails? "A second AI checks the first" was doomed: a verifier sharing the generator's blind spots, rubber-stamping the same errors. BS checking BS. So why would R1 stake its trust on that? Because R1 fixes the two things that broke the naive version — precisely the two the last act told us to fix.
Fix one: a genuinely different brain. The naive verifier failed because it was a sibling — same training, same architecture, same blind spots. R1 makes the reviewer a different model, blind to different things, so its mistakes and the worker's don't line up. It sees what the worker couldn't.
Fix two: anchor the review to criteria, not vibes. The naive judge floated on a subjective sense of quality — easy to share, easy to fool. R1's reviewer grades against the SOW's mechanical acceptance criteria, not "looks good to me." That's the last act's rule — ground the check outside the models — made concrete.
Why both? A different brain grading on vibes just disagrees by feel. Shared criteria judged by the same brain still share its blind spots. You need both: independent eyes and an objective target. Together they turn "a second AI checks the first" into a real check — the last act's hardest problem, answered in the runtime — reduced, not eliminated.
The trap: dismissing this as the doomed "AI judges AI" setup. It isn't — it fixes the two things that broke it: a different brain for independent blind spots, and an anchor to criteria for a grounded target. Whether AI checking AI works isn't a slogan; it's a question of how it's wired.
So R1's reviewer runs on a different brain and grades against mechanical criteria — the built answer to "why a second AI fails." Cross-model gives it different blind spots; the criteria keep it grounded. One outside mind is powerful. But what if many specialist minds reasoned in parallel over one shared context? Next: six minds in one context — the Cortex.