Act 04 · The physical machine 2:45 Cards, ops, and the building

Which GPU for which job? The card taxonomy.

GPUs ship in classes — flagship SXM, PCIe datacenter, inference fleet, workstation, consumer, unified-memory, and non-NVIDIA lanes — and each class has its own power, cooling, and interconnect needs that decide which jobs it can hold.

Video rendering soonThe cinematic render for this supplement is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — GPUs ship in classes — flagship SXM, PCIe datacenter, inference fleet, workstation, consumer, unified-memory, and non-NVIDIA lanes — and each class has its own power, cooling, and interconnect needs that decide which jobs it can hold.
  • How it is shown — The showroom: seven card classes on podiums, each with its wattage-aura, cooling need, and the jobs it serves; the decision table assembling at the end.
  • The trap to avoid — Buying by memory size alone — a card's class (interconnect, form factor, cooling) decides its jobs as much as its gigabytes.

This act has said "GPU" as if it were one animal. It's seven.

The one idea

GPUs ship in classes — flagship SXM, PCIe datacenter, inference fleet, workstation, consumer, unified-memory, and non-NVIDIA lanes — and each class has its own power, cooling, and interconnect needs that decide which jobs it can hold.

From thousand-watt liquid-cooled flagships to silent laptop chips, the card market is a taxonomy — each class with power, cooling, and plumbing needs that decide its jobs. The showroom tour, then the decision table. Top tier: the datacenter cards, in two costumes. SXM modules — the H100-and-Blackwell class — bolt onto carrier boards: seven hundred to a thousand-plus watts, liquid or forced-air cooled, full NVLink. Act four's chassis organisms; they cannot live outside a datacenter. Their PCIe siblings trade wattage and NVLink for a standard slot — same silicon family, tamer needs, humbler ceilings. The middle tiers do the volume work. Inference fleet cards — the L4 and L40S class — run seventy to three-fifty watts, air-cooled: cheap to rack by the dozen, serving models more than training them. Workstation cards put forty-eight gigabytes and certified drivers under an engineer's desk — development, fine-tunes, and quiet dignity. Fleets serve; desks develop. Home tier, two philosophies.

How it works — the demo

The showroom: seven card classes on podiums, each with its wattage-aura, cooling need, and the jobs it serves; the decision table assembling at the end.

Consumer flagships pack twenty-four to thirty-two gigabytes at up to five-seventy-five watts: act four's local-lab workhorses, mighty and loud. Unified-memory machines — Apple's silicon — share one huge pool — hundreds of gigabytes at the top end — CPU and GPU together: bigger models than any consumer card, at gentler speeds and whisper power. Capacity versus velocity, both home-sized. And the other lanes. AMD's MI-series brings enormous memory and aggressive pricing — with tooling still thinner than CUDA's dense web, the market's real moat. Google's TPUs live only as cloud pods. Edge modules run models on robots and cameras. The pattern: silicon competes well; tooling decides adoption. Price the software tax before crossing lanes. The table: frontier training wants SXM pods. Volume serving: PCIe and fleet cards.

The trap to avoid

Buying by memory size alone — a card's class (interconnect, form factor, cooling) decides its jobs as much as its gigabytes.

Why it matters — and what’s next

Fine-tuning and development: workstation or consumer silicon. Local AI: consumer or unified memory. Edge: modules. And the trap — buying by gigabytes alone. A card's class carries the real constraints: interconnect for sawn models, cooling for its room, wattage for its panel. Memory is one column. The class is the row. Cards chosen, racks humming — and now the part every glossy build-video skips: keeping models alive in production. Engines and driver stacks, rollouts that can't break Friday traffic, monitoring that catches quality rot, and the pager that goes off at three a.m. Episode one-eighteen priced the ops payroll. Next: what that payroll actually does.

This is a supplement in AI: Zero → Frontier — a side-trip that deepens the act it sits beside, one file and one loop at a time.

Full transcript 2:45 of narration

This act has said "GPU" as if it were one animal. It's seven. From thousand-watt liquid-cooled flagships to silent laptop chips, the card market is a taxonomy — each class with power, cooling, and plumbing needs that decide its jobs.

The showroom tour, then the decision table. Top tier: the datacenter cards, in two costumes. SXM modules — the H100-and-Blackwell class — bolt onto carrier boards: seven hundred to a thousand-plus watts, liquid or forced-air cooled, full NVLink.

Act four's chassis organisms; they cannot live outside a datacenter. Their PCIe siblings trade wattage and NVLink for a standard slot — same silicon family, tamer needs, humbler ceilings. The middle tiers do the volume work.

Inference fleet cards — the L4 and L40S class — run seventy to three-fifty watts, air-cooled: cheap to rack by the dozen, serving models more than training them. Workstation cards put forty-eight gigabytes and certified drivers under an engineer's desk — development, fine-tunes, and quiet dignity. Fleets serve; desks develop.

Home tier, two philosophies. Consumer flagships pack twenty-four to thirty-two gigabytes at up to five-seventy-five watts: act four's local-lab workhorses, mighty and loud. Unified-memory machines — Apple's silicon — share one huge pool — hundreds of gigabytes at the top end — CPU and GPU together: bigger models than any consumer card, at gentler speeds and whisper power.

Capacity versus velocity, both home-sized. And the other lanes. AMD's MI-series brings enormous memory and aggressive pricing — with tooling still thinner than CUDA's dense web, the market's real moat.

Google's TPUs live only as cloud pods. Edge modules run models on robots and cameras. The pattern: silicon competes well; tooling decides adoption.

Price the software tax before crossing lanes. The table: frontier training wants SXM pods. Volume serving: PCIe and fleet cards.

Fine-tuning and development: workstation or consumer silicon. Local AI: consumer or unified memory. Edge: modules.

And the trap — buying by gigabytes alone. A card's class carries the real constraints: interconnect for sawn models, cooling for its room, wattage for its panel. Memory is one column.

The class is the row. Cards chosen, racks humming — and now the part every glossy build-video skips: keeping models alive in production. Engines and driver stacks, rollouts that can't break Friday traffic, monitoring that catches quality rot, and the pager that goes off at three a.m.

Episode one-eighteen priced the ops payroll. Next: what that payroll actually does.

Hardware & InferenceHardwareProcurement