• | 9:00 am

Where Nvidia is going, it doesn’t need cables

In data centers, like everywhere else, cables are a hassle. So the AI chip giant’s new platform reduces 50 or 60 of them down to 2.

Where Nvidia is going, it doesn’t need cables
[Source photo: Courtesy of Nvidia]

Hello again, and welcome back to Fast Company’s Plugged In.

Last week, I attended an Nvidia media event that included a field trip to an undisclosed location in Silicon Valley—a building without any identifying signage. Though hundreds of racked GPUs were chugging away at real AI tasks inside, Nvidia calls the facility an “engineering superlab” rather than a data center. Among the AI being crunched during our visit: a test version of OpenAI’s GPT model optimized for Nvidia’s new Vera Rubin platform, named after the legendary astronomer.

For show-and-tell purposes, several powerful computers were on display in the building’s lobby. Containing GPUs, CPUs, and supporting components and known as trays, they slide into racks like the ones in the superlab. Data centers equipped with vastly larger quantities of these racks power a huge percentage of the AI in our lives.

Basically, we all use these Nvidia tray computers all day long. It’s just that under normal circumstances, we never see them or even have much reason to give them any thought. Getting a close-up peek at the ones at Nvidia’s unmarked superlab was a rare opportunity to contemplate how they do what they do.

For the sake of comparison, the ones in the lobby included a Vera Rubin tray sittting next to one based on Nvidia’s previous-generation platform for its GB200 and GB300 chips. At a glance, it was obvious that the redesign went far beyond the new tray’s more powerful chips. The older tray looked like a rat’s nest, with an array of cables snaking every which way to connect components. The Vera Rubin tray was radically more streamlined. It was cable-free.

Well, not literally cable-free. “We call it cableless—it actually has two cables,” explained Senior VP of Hardware Engineering Andrew Bell, a 25-year Nvidia veteran. One cable remains to hook up the tray to power, another to connect it to networking resources. But they were unobtrusive enough that I didn’t notice them until he pointed them out.

By water-cooling its Vera Rubin computing tray, Nvidia was able to replace dozens of cables with a midplane connector. [Photo: Courtesy of Nvidia]

It also might be unfair to call the earlier design a rat’s nest. it was way more fastidiously organized than the bedside tangle of USB cables I use to charge my phone, laptop, and iPad. But when a computer has 50 or 60 cables, as this one did, it gives off an overwhelming impression of steampunk-y chaos. Especially when it’s sitting next to one that’s so close to being cable-free.

Two overarching questions are relevant to this redesign: Why did Nvidia decide to go (almost) cableless, and how did the company pull it off? The first answer might be pretty obvious if you’ve ever built a desktop PC, or at least opened one up.

Inside a computer, cables just complicate things. Some 50 or 60 cables add up to 100 to 120 additional steps during assembly. Each connection is a place something could go wrong during the building process and, after that, a potential point of failure. Cables create thickets that get in the way of other components, both during manufacturing and later servicing.

“It takes two to three hours to assemble [the GB200/GB300 tray],” Bell said. “And when we first started building it, you had about a 20% chance of getting it right the first time.”

Ian Buck, VP and GM of hyperscale and high-performance computing, added, “Just imagine yourself plugging in every one of those cables and trying to do it at a scale of Nvidia.”

The success rate did gradually increase, but it wasn’t just Nvidia’s own scale that was an issue. Rather than mass-producing trays itself, the company shipped parts to contractors and manufacturing partners, such as Dell. Each time production expanded, it set off a Groundhog Day-style cycle as new teams learned how to assemble the cable-heavy trays.

Nvidia’s previous-generation tray required cables, more cables, and a few cables on top of that. [Photo: Courtesy of Nvidia]

If cables’ downsides are so many and varied, why have they been so fundamental to computer designs for so long? One way or another, components do need to talk to each other. For a long time, it wasn’t obvious how to achieve that sans cabling.

For Nvidia, the breakthrough cascaded out of another major design achievement reflected in the Vera Rubin platform: It’s fanless. Instead of using air cooling to prevent overheating, the tray is 100% liquid cooled, using a new approach that works with water at up to 113 degrees Fahrenheit, dramatically reducing energy requirements compared to earlier liquid-cooling techniques.

As a bonus, removing fans also promises to dramatically reduce the noise pollution inside data centers. (At the superlab, which is full of fan-cooled racks, we wore earplugs, and Bell spoke to us with a microphone despite being only a few feet away.)

When Nvidia figured out how to ditch the fans, it opened up room inside the Vera Rubin trays. More room allowed the company to create a midplane connector, based on its NVLink technology, that established cable-free communications between components. Even if Nvidia had somehow found room for the midplane connector in an air-cooled tray, it wouldn’t have worked: It would have obstructed airflow.

In short, one elegant technical breakthrough enabled another. You don’t need to get that backstory from Nvidia executives to see the elegance in the resulting Vera Rubin tray. It’s sleek, black, and bordering on brutalism in its lack of ornamentation. Compared to it, its predecessors look like highly compressed versions of 1951’s Univac.

All of this is ultimately in service of a faster, more reliable ramp-up of the Vera Rubin platform, in factories that can now be more fully automated. “This [tray] can be assembled in five minutes instead of two to three hours,” Bell said. “The first time you put it together, you have a 95% chance of getting it right. You can take it to a new factory. The first time they put it together, it’s 95%. We expect after a month, they’ll be at 98% or 99%.”

With other AI chip makers ever more intent on stealing market share from Nvidia—hello, AMD!—those efficiencies matter. But even if Nvidia wasn’t after elegance for elegance’s sake, the company has achieved it—and I, for one, will never assume that cables are one of computing’s eternal necessary evils again.

  Be in the Know. Subscribe to our Newsletters.

ABOUT THE AUTHOR

Harry McCracken is the technology editor for Fast Company, based in San Francisco. In past lives, he was editor at large for Time magazine, founder and editor of Technologizer, and editor of PC World. More More

More Top Stories:

FROM OUR PARTNERS