Google bets big on marvell for ai inference boost
Google is doubling down
on its AI ambitions, but not with another in-house chip breakthrough. Instead, the tech giant is reportedly in talks with Marvell Technology to develop two new chips specifically tailored for AI inference—the computationally intensive phase where models deliver real-world value. This signals a strategic shift, suggesting Google recognizes the need to diversify its silicon supply chain and accelerate inference capabilities.The memory and tpu partnership
The planned chips are a critical pairing. One will be a memory chip designed to work in tandem with Google's Tensor Processing Units (TPUs). The other? A brand-new TPU architecture explicitly built for inference tasks. This isn’t about replacing Google’s existing TPU efforts, but rather complementing them, alongside collaborations with Broadcom, MediaTek, and TSMC. It's a calculated move to bolster performance across the board.
Google’s Ironwood, the seventh generation of its AI accelerators, launched just last November as part of Google Cloud, already promises a tenfold performance improvement over the v5p TPU and a more than fourfold increase in both training and inference workloads compared to the v6e. But the limitations of relying solely on internal chip development are becoming apparent. The sheer scale of AI deployment demands a multifaceted approach to silicon sourcing.
Consider this: Ironwood can integrate up to 9,216 chips into a 'superpod'—a massive AI supercomputer—connected via Google's proprietary ICI interconnect. That network moves a staggering 9.6 Tb/s of data, creating a shared 1.77 Petabytes of high-bandwidth memory, a resource far more efficient than traditional RAM. While impressive, building and maintaining such colossal systems is a resource-intensive challenge.

Microsoft’s challenge and the inference race
Google isn’t alone in recognizing the importance of inference. Microsoft’s recent unveiling of Maia 200, an AI accelerator boasting three times the FP4 performance of Amazon’s Trainium Gen 3 and outperforming Google’s seventh-generation TPU in FP8, demonstrates the intensifying competition. The race isn't just about training models; it's about delivering insights and actions at scale, and that demands optimized inference hardware.
The deal with Marvell suggests Google is prioritizing agility and scale in its AI infrastructure. By partnering with a leading semiconductor manufacturer, Google can potentially bypass some of the bottlenecks associated with in-house chip fabrication, gaining a competitive edge in the burgeoning AI landscape. The move signifies a pragmatic recognition that even the deepest pockets can benefit from strategic outsourcing—especially when the stakes are this high.
