Google Ironwood TPU: Inside the 42.5-Exaflop Powerhouse Built for the “Age of Inference”
Saturday, November 08, 2025Google Ironwood TPU: Inside the 42.5-Exaflop Powerhouse Built for the “Age of Inference”
Google has pulled back the curtain on Ironwood, its seventh-generation Tensor Processing Unit—and the first TPU purpose-built for inference rather than training. Unveiled at Google Cloud Next ’25 and detailed again at Hot Chips 2025, Ironwood is already being deployed in 9,216-chip “SuperPods” that deliver a mind-bending 42.5 exaflops of FP8 compute, more than 24× the horsepower of today’s top supercomputer.
Why Ironwood Exists – The Age of Inference
Google argues that AI is shifting from “ask-and-answer” workloads to continuous, agentic reasoning—models that proactively retrieve, plan and act. These “thinking” pipelines need three things:
- Massive, low-latency tensor throughput
- Terabytes of directly addressable memory
- Chip-to-chip bandwidth that scales linearly with model size
Ironwood is Google’s answer to all three.
Ironwood at a Glance
| Spec | Ironwood (per chip) |
|---|---|
| Peak FP8 Compute | 4,614 TFLOPS |
| High-Bandwidth Memory | 192 GB HBM3e |
| Memory Bandwidth | 7.37 TB/s |
| Inter-Chip Interconnect | 1.2 TB/s bidirectional |
| Power Efficiency vs Trillium | 2× perf/W |
| Max Scale per Pod | 9,216 chips / 42.5 EFLOPS |
From Chip to City-Block Scale
- Die level: Two compute chiplets on one interposer—Google’s first multi-die TPU
- Tray level: Four Ironwood TPUs on a liquid-cooled blade
- Rack level: 16 trays = 64 chips in a 3-D torus (4×4×4)
- SuperPod: 144 racks = 9,216 chips with 1.77 PB of directly addressable HBM
- Cluster: Optical Circuit Switches (OCS) glue multiple SuperPods into 100,000+ chip “AI Hypercomputers”
Breakthroughs That Matter
1. SparseCore Gen-4 – dedicated silicon for embeddings and mixture-of-experts gating, doubling sparse-model throughput
2. Integrated root-of-trust & confidential-compute – secure boot, live-integrity checks and silent-data-correction protect trillion-parameter weights
3. AI-designed ALUs – floor-plan and circuit paths co-optimised by Google’s AlphaChip team, shaving 8 % power and 5 % latency
4. Third-gen liquid cooling – dual-loop cold-plate design keeps inlet water below 35 °C even at 10 MW rack density
Real-World Muscle – Customer Wins
Anthropic has already committed to one million Ironwood TPUs to serve Claude 3.5 and future reasoning models. Google’s own Gemini 2.0 “Flash Thinking” is trained and served exclusively on Ironwood pods, cutting inference cost per 1 k tokens by 42 % versus the prior Trillium fleet.
Availability & Pricing
Google Cloud will offer Ironwood in two slices:
- Ironwood-256 – 256-chip slice for mid-size training or high-QPS inference
- Ironwood-9216 – full SuperPod for frontier-model pre-training or mega-batch serving
General availability opens “in the coming weeks” with pay-per-second billing and committed-use discounts up to 70 %.
Bottom Line
Ironwood isn’t just a faster TPU—it’s Google’s declaration that the future of AI is memory-bound, inference-centric and agent-driven. With 192 GB of HBM3e per chip, 42.5 exaflops per pod and a 2× jump in power efficiency, Ironwood gives Google Cloud the biggest hammer in the hyperscale toolbox. For customers, that translates to lower cost per token, larger context windows and the head-room to build models that think instead of merely answer.
Ironwood SuperPods are rolling out to Google Cloud regions us-central1, europe-west4 and asia-southeast1 throughout Q4 2025.