Google's TPU Split Shows Inference Chips Taking Over
Google's TPU plans split inference from training, lifting MediaTek while cutting into Broadcom's position.
TL;DR:
- Specialized inference chips make more economic sense than big all-purpose AI designs.
- MediaTek moves up as a real AI chip supplier with growing cloud business.
- Broadcom loses some leverage as inference work shifts to cheaper partners.
- AI agents and reinforcement learning make memory layout and fast decoding the key limits.
Google's move to add a higher-end inference version to its TPU v9 plans shows that custom chips now work better when split by task. The Triggerfish program turns MediaTek from a side player into a main partner in Google's supply chain. It shows how big cloud companies are splitting workloads to get costs and performance that Nvidia can't match at scale.
Inference split weakens Broadcom and boosts MediaTek
| Narrative / Interpretation Camp | Evidence / Signal / Source of Conviction | How This Affected Industry Thinking or Positioning | Your Strategic Judgment | |--------------------------------|------------------------------------------|-----------------------------------------------------|-------------------------| | MediaTek moves into top-tier AI ASIC spot | Kuo's notes on the exclusive Triggerfish deal, 2-3x SRAM, simulation die, HBM4E, and 30% price bump over Humufish; backed by Counterpoint and CryptoBriefing on 25% market share by 2028 | Old view of MediaTek as just a phone chip maker fades; investors now see Cloud ASIC revenue hitting $18B+ by 2027 | MediaTek margins should grow faster than people think because inference chips can charge more than training ones | | Google is pulling back from Broadcom | TPU v8 split (Broadcom Sunfish for training vs MediaTek Zebrafish for inference) plus four-partner setup with Marvell and Intel | The picture changes from Nvidia versus custom chips to mixed ASIC groups; this cuts Broadcom's bargaining power | Broadcom's custom AI forecasts look too high once inference moves to lower-cost partners | | AI agent and RL work change hardware needs | Simulation die and SRAM growth tied to agent coordination and low-latency decode; production starts 2028 | Talk moves from raw FLOPS to memory efficiency as the real limit on agent systems | Early players in agent tools get lasting cost edges; late ones pay more for inference |
The real change is the clear split between training and inference silicon. It shows that agent workloads create their own hardware math instead of just tweaking old designs.
- Training chips stay flexible and high-margin; inference chips get pushed hard on cost.
- MediaTek's SerDes skills and TSMC ties now feed straight into AI revenue that phone cycles never gave.
- Google's 1-2 million extra Triggerfish units sit on top of committed Humufish volumes, so inference demand is not taking from training.
Ignore the idea that this is just supply-chain theater. The 30% price premium and simulation die show the upgrade fixes real gaps in reinforcement learning and agent work; Google wouldn't pay more otherwise.
Significance: High Categories: Industry Trend, Market Impact, Partnership
Verdict: Builders and investors in inference-focused custom silicon are ahead; those still treating all AI chips as training-heavy or Broadcom-only are behind and will stay mispriced through 2028.