Nvidia Is Jolting Silicon Valley to Think Like China

When a typical business finds its supply chain costs have risen, it will absorb the increase to keep its customers happy. Nvidia isn’t a typical business. It enjoys gross profit margins of 75% and commands between 70% and 90% of the global market for artificial-intelligence chips. So when the companies that make memory components put their prices up, Nvidia passed the hike on. It’s now raising the prices of servers containing its accelerator processors by more than 15%, according to Bloomberg News.

That won’t deter big tech customers — some of the most well-capitalized companies in history — from buying Nvidia chips. But it will probably force a reevaluation of how companies like Meta Platforms Inc., Microsoft Corp. and OpenAI use them. The US firms will need to start thinking more like their counterparts in China, who’ve learned from years of export bans to work around chip constraints by introducing more efficient software.

Take Hangzhou-based DeepSeek. After US export controls blocked it from buying Nvidia’s most powerful chips, it figured out how to compress the memory needs of its AI models by writing a clever mathematical shorthand called multi-head latent attention. It cut the company’s chip memory needs by about 96%, allowing it to build a world-class product at a fraction of the typical hardware costs.

Silicon Valley could have pursued similar innovations to use semiconductors more efficiently. In fact, it should have. The use of Nvidia chips to process tasks for end customers is remarkably inefficient. One founder of an AI chip startup recently told me that AI firms were using less than 15% of the theoretical capacity of Nvidia’s graphics processing units (GPUs). Others estimate that figure to be as low as 5%.

The problem isn’t the GPUs themselves — remarkably powerful machines that can do millions of mathematical calculations at once — but the flood of information they must handle when thousands are stitched together in a data center. The software1 powering them is good at telling a single chip how to process graphics or perform a calculation. But it wasn’t designed to distribute the data from a large AI model across thousands of chips simultaneously. That leads to massive data traffic jams, and chips sitting idle as they wait for the software to move numbers between them.

Until now, tech giants have dealt with the problem by just buying more chips to make up for the inefficiencies. They could afford to, and being in a fast-moving AI race made it difficult to take the time to rewrite millions of lines of entrenched code for running processors, especially when Nvidia was constantly releasing new chip architectures.

See more: The Quantum Computer Revolution Is Tantalizingly Close