AWS to deploy 2 million more Nvidia GPUs as AI demand runs ahead of expectations
ADVERTISEMENT

Amazon Web Services (AWS) and Nvidia are expanding their partnership, with AWS planning to deploy an additional 2 million Nvidia GPUs across its global infrastructure in 2027 and 2028, as demand for AI computing continues to grow faster than expected.
The announcement comes just five months after AWS said it would add more than 1 million Nvidia GPUs to its infrastructure starting in 2026. Nvidia said demand for that capacity has since exceeded expectations, prompting the companies to expand their plans further. The new deployment will include Nvidia’s Blackwell Ultra, Rubin and Rubin Ultra GPUs.
The expanded partnership is not limited to GPUs. AWS and Nvidia are also working together on CPUs, networking, AI models, data processing and robotics as businesses move more AI workloads from experimentation into production.
“Customers want the freedom to choose the best tools for their AI workloads, and they want confidence that everything works seamlessly together,” said Matt Garman, CEO of AWS. “That’s why we’ve invested deeply with NVIDIA to make AWS the best place to run NVIDIA AI technologies, optimizing performance across our infrastructure from networking and security to deployment. This expanded collaboration gives frontier labs, enterprises and governments even more ways to build and deploy AI on AWS.”
Nvidia CEO Jensen Huang said the companies were seeing demand accelerate even as they continued to scale their businesses. “NVIDIA and AWS have built one of the great growth engines of the AI era, and demand is running ahead of every forecast,” Huang said. “For 16 years, we have scaled NVIDIA computing in the cloud together. Now, we are expanding our partnership across the full stack — GPUs, CPUs, networking, open models and software — to make agentic and physical AI real at an unprecedented pace and scale that only AWS and NVIDIA can deliver.”
Nvidia’s Vera CPUs coming to AWS
As part of the expanded agreement, AWS and Nvidia will bring Nvidia’s Vera CPU-based infrastructure to AWS. Vera is designed to handle the CPU-heavy parts of newer AI workloads, including agentic AI, where AI systems can execute code, use tools, process data and coordinate multiple tasks. The addition gives AWS customers another option alongside the company’s own custom chips, including Trainium, as well as Nvidia’s GPUs and CPUs.
The companies are also working to connect AWS’s Trainium chips with Nvidia technology. Amazon’s Annapurna Labs is expanding its work with Nvidia’s NVLink Fusion, a high-speed chip interconnect technology, to support Nvidia’s new custom high-bandwidth memory, or NVHBM. The combination is intended to allow Trainium chips and Nvidia GPUs to work together within the same rack-scale system. This could give AWS more flexibility in building AI infrastructure while allowing customers to use different types of compute for different parts of an AI workload. AWS and Nvidia are also expanding their work on networking. The companies are collaborating on Nvidia Spectrum networking technology to improve network performance for large AI training workloads that use large clusters of GPUs.
100,000 GPUs planned for US government AI
The expanded partnership also includes plans to build AI infrastructure for the US government.
AWS and Nvidia said they plan to deliver Nvidia’s AI stack, including 100,000 GPUs, on secure AWS infrastructure for federal and national-security workloads. The systems will support workloads classified at Impact Level 6 (IL6) and above, which covers highly sensitive government data. The two companies said the infrastructure will use AWS’s Nitro System and Elastic Fabric Adapter (EFA) to provide security, reliability and high-speed networking across GPU and Trainium-based systems.
The partnership will also make Nvidia’s Nemotron family of open AI models available through Amazon Bedrock and Amazon SageMaker. Bedrock allows customers to use managed AI models through AWS, while SageMaker is aimed at customers that want to deploy and fine-tune models on their own infrastructure.
The companies are also targeting the growing amount of data needed to train and run AI applications. Nvidia’s cuDF library will be used with Amazon EMR to accelerate data processing. The companies said this can deliver up to 3.7 times faster processing and 30% better price performance than CPU-based configurations. For Amazon OpenSearch, GPU-based vector indexing can deliver up to nine times faster indexing at a quarter of the cost, according to the companies. Vector databases are increasingly important for AI applications such as retrieval-augmented generation and semantic search.
Amazon brings Nvidia’s physical AI stack to robotics
The partnership is also moving into physical AI, as Amazon Robotics works with Nvidia on next-generation warehouse robots.
Amazon Robotics will use Nvidia’s Jetson platform, Omniverse libraries and Isaac robotics platform for areas including simulation, synthetic data generation, robot training, route optimisation and safety testing. The companies said the work will allow robotics systems to be trained and tested using GPU-powered AWS infrastructure before being deployed in real-world environments.
The broader expansion comes as AI workloads are moving beyond training large language models. Companies are increasingly using AI for autonomous agents, scientific research, enterprise automation and robotics, creating demand for more computing capacity across GPUs, CPUs, networking and data infrastructure.
For AWS, the deal also strengthens its position as a major customer and distribution channel for Nvidia at a time when Amazon is simultaneously developing its own AI chips. For Nvidia, the additional AWS capacity provides another indication that demand for AI infrastructure remains strong across the major cloud providers.
Neither company disclosed financial terms for the expanded partnership.