Posts

Showing posts with the label NVLink

BWR Episode 8: Nvidia@CES, Upscale AI, Marvell Xconn, RISC-V

Image
The latest episode of The Byrne-Wheeler Report covers a slew of deals from funding to outright acquisitions. In my coverage space, Upscale AI has raised another $200 million at a unicorn valuation. Marvell is acquiring CXL-switch startup Xconn Technologies. Joe and I discuss two announcements in the RISC-V arena. In case you're wondering about Episode 7, it was a special episode with guest Mike Demler discussing Nvidia's Christmas Eve $20 billion bombshell Groq deal.

BWR Episode 5: Celestial AI, More NVLink, and Merchant TPUs

Image
The Byrne-Wheeler Report Episode 5 is now available. In this episode: Marvell Technology has announced the acquisition of Celestial AI, a move that positions Marvell as a leader in next-generation co-packaged optics (CPO). Unlike traditional CPO focused on standards compliance, this deal targets the bleeding edge of scale-up interconnects for AI accelerators. In a surprising shift at AWS re:Invent, Amazon disclosed that its upcoming Trainium 4 AI accelerator will support NVLink Fusion. While Amazon typically relies on its proprietary NeuronLink, this move allows for a cookie-cutter rack design. AWS will be able to mix and match Nvidia GPUs and Trainium chips within the same physical infrastructure, speeding up deployment velocity. Reports indicate a massive shift in Google’s strategy: the company may be moving from strictly using TPUs for its own cloud services to acting as a merchant silicon supplier. Rumors suggest Meta is planning to deploy its own on-premise TPU cluster. Please exc...

Decoding Nvidia's Rubin Networking Math

Image
At GTC DC last month, Jensen Huang showed off components of the Vera Rubin NVL144 platform. First, here's the latest roadmap, which now includes BlueField-4 and BlueField-5. For more on that, see BWR Episode 4 .  Source: Nvidia Below is the Vera Rubin compute tray, which includes four Rubin GPUs. By GPU, we mean package not die. Note that the Blackwell NVL72 and Rubin NVL144 both have 72 GPU packages, but the NVL144 moniker denotes Nvidia's new math counting die. The company didn't rename the Blackwell configuration, even though that GPU also has two die. Each compute tray has two Vera CPUs, which are 88-core Arm processors. Two GPUs connect with one CPU using NVLink-C2C, a coherent variant of NVLink. Although the roadmap above shows CX9 as 1600G, each ConnectX-9 is actually 800Gbps, requiring eight chips to deliver the aggregate 800GB/s quoted for the tray. That means each GPU has a pair of 800G Ethernet/InfiniBand NICs for scale-out networking. Finally, a single BlueField...

LightCounting Publishes October 2025 Ethernet, InfiniBand, and Optical Switches for Cloud Data Centers Report

Image
October was a busy month, between the OCP Global Summit and being the lead author on the LightCounting switch report. The new report adds another level of granularity to the scale-up switch forecast, which incorporates NVLink, UALink, Scale-Up Ethernet (SUE), and non-Nvidia proprietary interconnects. There are many other changes to the report, including revisions to the methodology for the co-packaged optics (CPO) forecast. The newsletter summary is freely available: Co-packaged Optics Grow the Scale-out Switch Pie Source: Wheeler's Network

Introducing The Byrne-Wheeler Report

Image
My colleague Joe Byrne and I launched a video podcast covering the latest news in the chip world. Episode 1 covers Hot Interconnects and Hot Chips news from August as well as Arm's Lumex announcement and Nvidia's Rubin CPX. We hope you enjoy this new format, which will complement our respective blogs.  

Broadcom Adds New Architecture With Tomahawk Ultra

Image
Source: Broadcom Tomahawk Ultra is a misnomer. Although the name leverages Tomahawk's brand equity, Tomahawk Ultra represents a new architecture. In fact, when it began development, Broadcom's competitive target was InfiniBand. During development, however, AI scale-up interconnects emerged as a critical component of performance scaling, particularly for large language models (LLMs). Through luck or foresight, Tomahawk Ultra suddenly had a new and fast-growing target market. Now, the leading competitor was NVIDIA's NVLink. Also happening in parallel, Broadcom built a multi-billion-dollar business in custom AI accelerators for hyperscalers, most notably Google. At the end of April, Broadcom announced its Scale-Up Ethernet (SUE) framework, which it published and contributed to the Open Compute Project (OCP). Appendix A of the framework includes a latency budget, which allocates less than 250ns to the switch. At the time, we saw this as an impossibly low target for existing Eth...

Broadcom Pitches Ethernet for AI Scale Up

Image
Tomahawk 6 is First to 102.4T Through relentless execution, Broadcom has been first to market generation after generation in data-center switching. The company just announced sampling of Tomahawk 6 (TH6), its 102.4T Ethernet switch ASIC. This generation actually consists of two switch chips, TH6-200G with 512x200G SerDes, and TH6-100G with 1,024x100G SerDes, both of which are sampling now. A version with fully co-packaged optics, TH6-Davisson, will follow on a to-be-announced schedule.  Whereas Tomahawk 5 (TH5) is a monolithic 5nm chip, TH6 comprises a core die and separate chiplets for the two SerDes options, all of which use 3nm technology. Source: Broadcom For AI scale-out networks, TH6 enables a 128K-XPU network using only two switch tiers. Fewer tiers mean lower latency, simpler load balancing and congestion control, and fewer optics. The new chip is first to handle 1.6T Ethernet ports, but it also handles up to 512x200GbE ports for maximum radix. Beyond sheer port density, TH...

AI Unsurprisingly Dominates Hot Chips 2024

Image
This year's edition of the annual Hot Chips conference represented the peak in the generative-AI hype cycle. Consistent with the theme, OpenAI's Trevor Cai made the bull case for AI compute in his keynote. At a conference known for technical disclosures, however, the presentations from merchant chip vendors were disappointing; despite a great lineup of talks, few new details emerged. Nvidia's Blackwell presentation mostly rehashed previously disclosed information. In a picture-is-worth-a-thousand-words moment, however, one slide included the photo of the GB200 NVL36 rack shown below. GB200 NVL36 rack (Source: Nvidia) Many customers prefer the NVL36 over the power-hungry NVL72 configuration, which requires a massive 120kW per rack. The key difference for our readers is that the NVLink switch trays shown in the middle of the rack have front-panel cages, whereas the "non-scalable" NVLink switch tray used in the NVL72 has only back-panel connectors for the NVLink spin...

NVIDIA Reveals Roadmap at Computex

Image
The annual Computex trade show in Taipei has traditionally been PC-centric, with ODMs showing their latest motherboards and systems. The 2024 event, however, included keynotes from Nvidia and others that revealed details of forthcoming datacenter GPUs, demonstrating the importance of the ODM ecosystem to the explosion of AI. The fact that Jensen Huang was born on the island made his keynote all the more impactful for the local audience. In the week following the CEO's keynote, Nvidia's market capitalization surpassed $3 trillion. From a networking perspective, the keynote focused on Ethernet rather than InfiniBand, as the former is a better fit in the ecosystem messaging. Source: NVIDIA The datacenter section of Jensen's talk largely reminded the audience of what Nvidia announced at GTC in March. The Blackwell GPU, now in production, introduces NVLink5, which operates at 200Gbps per lane. It includes 18 NVLink ports with two lanes each, or 36x200Gbps serdes. The new NVLink...

All Eyes on NVIDIA

Image
Aside from CEO Jensen Huang, the DGX GB200 NVL72 was the star of the GTC 2024 keynote. The rackscale system integrates 72 next-generation Blackwell GPUs connected by NVLink to form “1 Giant GPU.” Jensen’s description of the NVLink passive-copper “backplane” caused a brief panic among investors that believed it somehow replaced InfiniBand, which it does not. The NVL72 represents next-generation AI systems, but Nvidia also revealed new details of its deployed Hopper-generation clusters. Next-generation 800G (XDR) InfiniBand won’t reach customers until 2025, so early Blackwell systems will use 400G (NDR) InfiniBand instead. Source: NVIDIA Jensen said the Hopper-generation EOS supercomputer had just come online. This cluster uses 608 NDR switches with 64 ports each for a total of 38,912 switch ports. This system places the leaf switches in racks at the end of the row, so all InfiniBand links employ optical transceivers. We estimate the servers add 5,120 ports for a system total of 44,032 N...

AMD Looks to Infinity for AI Interconnects

Image
With the formal launch of the MI300 GPU, AMD revealed new plans for scaling the multi-GPU interconnects vital to AI-training performance. The company's approach relies on a partner ecosystem, which stands in stark contrast with NVIDIA's end-to-end solutions. The plans revolve around AMD's proprietary Infinity Fabric and its underlying XGMI interconnect. Infinity Fabric Adopts Switching As with its prior generation, AMD uses XGMI to connect multiple MI300 GPUs in what it calls a hive. The hive shares a homogeneous memory space formed by the HBM attached to each GPU. In current designs, the GPUs connect directly using XGMI in a mesh or ring topology. Each MI300X GPU has up to seven Infinity Fabric links, each with 16 lanes. The 4th-gen Infinity Fabric supports up to 32Gbps per lane, yielding 128GB/s of bidirectional bandwidth per link. At the MI300 launch, Broadcom announced that its next-generation PCI Express (PCIe) switch chip will add support for XGMI. At last October...

NVIDIA Reveals DGX GH200 System Architecture

Image
We industry analysts sometimes get out over our skis when trying to project details of new products. Following NVIDIA's DGX GH200 announcement at Computex, we noticed industry press making the same mistake. Rather than correct our previous NVLink Network post, we'll explain here what we've since learned. Our revelation came when NVIDIA published a white paper titled NVIDIA Grace Hopper Superchip Architecture . The figure below from that paper shows the interconnects at the HGX-module level. The GPU (Hopper) side uses NVLink as a coherent interconnect, whereas the CPU (Grace) side uses InfiniBand, in this case connected through a Bluefield-3 DPU. From a networking perspective, the NVLink and InfiniBand domains are independent, that is, there is no bridging between the two protocols. HGX Grace Hopper Superchip System With NVLink Switch (Source: NVIDIA) The new DGX GH200 builds a SuperPOD based on this underlying module-level architecture. You've probably seen the headlin...

NVIDIA Networks NVLink

Image
I attended several sessions at last week's Hot Chips, and NVIDIA's NVSwitch talk was a standout. Ashraf Eassa did a great job of covering the talk's contents in an NVIDIA blog, so I will focus on analysis here.  SuperPOD Bids Adieu to InfiniBand From a system-architecture perspective, the biggest change is extending NVLink beyond a single chassis. NVLink Network is a new protocol built on the NVLink4 link layer. It reuses 400G Ethernet cabling to enable passive-copper (DAC), active-copper (AEC), and optical links. The build its DGX H100 SuperPOD, NVIDIA designed a 1U switch system around a pair of NVSwitch chips. The pod ("scalable unit") includes a central rack with 18 NVLink Switch systems, connecting 32 DGX H100 nodes in a two-level fat-tree topology. This pod interconnect yields 460.8Tbps of bisectional bandwidth. NVLink Network replaces InfiniBand (IB) as the first level of interconnect in DGX SuperPODs. The A100 generation pod uses an IB HDR leaf/spine inte...