Amazon Web Services plans to deploy 2 million additional NVIDIA GPUs across its global infrastructure in 2027 and 2028. The two companies are also extending their relationship into CPUs, rack-scale interconnects, memory, models, data processing and robotics.
The August 26 expansion announcement covers NVIDIA Blackwell Ultra, Rubin and Rubin Ultra GPUs. It comes on top of AWS’s earlier plan to add more than 1 million NVIDIA GPUs starting in 2026.
The important qualification is timing. These are planned deployments, not two million GPUs already installed or generally available to AWS customers. The companies did not provide a regional rollout, instance catalogue, pricing schedule or financial terms. TechCrunch also reported that the commercial terms were not disclosed.
The announcement is a capacity plan, not a completed rollout
AWS and NVIDIA attribute the expansion to growing demand from AI laboratories, enterprises, startups and governments. That explanation comes from the companies involved and should not be treated as an independently audited measurement of future AWS utilization.
NVIDIA’s business results provide useful context without proving the AWS forecast. In its second-quarter fiscal 2027 results, NVIDIA reported $96.2 billion in quarterly revenue, including $89 billion from its Data Center segment. Those figures show the current scale of NVIDIA’s data-centre business. They do not establish when AWS will receive the announced GPUs, how much capacity customers will consume or whether the deployment will produce the returns either company expects.
For cloud customers, the practical details will arrive later. Instance types, regions, reservation models, service quotas and prices will determine whether the capacity changes an actual workload plan. A large fleet announcement alone does not answer those questions.
The collaboration reaches into AWS custom silicon
The expansion is broader than a GPU purchase. AWS and NVIDIA say they are working to bring NVIDIA Vera CPU-based infrastructure to AWS and to connect NVIDIA technology more deeply with AWS’s own Trainium accelerator roadmap.
The companies plan to extend NVLink Fusion support for next-generation Trainium chips with NVIDIA’s custom high-bandwidth memory, called NVHBM. Their stated goal is a common rack-scale architecture in which Trainium and NVIDIA GPUs can be integrated more closely. That is an engineering plan, not evidence that the two accelerator families will expose identical software, performance or operating characteristics.
AWS also says NVIDIA GPU-based and Trainium-based EC2 instances use the AWS Nitro System and Elastic Fabric Adapter. The Nitro System moves infrastructure functions such as networking and I/O onto dedicated AWS hardware and software components. The Elastic Fabric Adapter provides low-latency, high-throughput communication for distributed AI, machine-learning and high-performance-computing workloads.
Those AWS-controlled layers matter because they let the cloud provider standardize parts of the operating environment even when the underlying processors differ.
The announcement also combines existing services with planned expansion. NVIDIA Nemotron models are already listed in the Amazon Bedrock model catalogue, while the companies say they will continue supporting Nemotron through Bedrock and SageMaker. They are also collaborating on GPU acceleration for EMR data processing, OpenSearch vector-index construction and Amazon Robotics workloads.
AWS and NVIDIA published performance and price-performance figures for some of those data and vector-processing integrations. The announcement does not provide enough independent methodology to treat those figures as generally proven results, so operators should validate them against their own data, configurations and cost boundaries.
More NVIDIA does not make Trainium irrelevant
Buying more supplier silicon while developing custom chips can look contradictory only if a hyperscaler is expected to choose one processor for every workload.
AWS has several reasons to maintain both paths. Customers may already depend on NVIDIA-oriented software and operational practices. Making that capacity available can keep those workloads on AWS rather than pushing teams to another cloud. Trainium gives AWS a separate design path that can increase control over its roadmap, supply choices, service packaging and infrastructure differentiation.
The deeper integration can also support the custom-silicon strategy. If Trainium can participate in a rack-scale environment that uses NVIDIA interconnect and memory technologies, AWS can offer a broader hardware mix while keeping Nitro, EFA, storage, orchestration and managed services as the common cloud layer.
That common layer creates its own form of platform stickiness. A customer may gain a choice of accelerator while remaining committed to AWS networking, data services, deployment tooling and operational controls. Hardware choice and cloud portability are not the same thing.
For technical leaders, this means the relevant comparison is not simply NVIDIA versus Trainium. It is the complete workload path: model and framework support, compiler maturity, distributed communication, memory requirements, regional capacity, observability, failure recovery and total operating cost.
Delivery and workload economics remain unresolved
The plan spans several processor generations and two future calendar years. Its execution depends on chip production, high-bandwidth memory, networking equipment, data-centre power, cooling, construction and AWS’s ability to convert components into usable cloud capacity.
Software integration presents another boundary. A workload tuned for one accelerator stack may require changes to run efficiently on another. Even when both options sit behind AWS services, teams still need to examine supported frameworks, numerical behaviour, profiling tools, debugging workflows and operational skills.
Several facts remain unknown: the financial size and contractual structure of the arrangement, the regional allocation of the two million GPUs, the mix among Blackwell Ultra, Rubin and Rubin Ultra, the precise delivery schedule, and how much of the capacity will be available through general EC2 services rather than dedicated arrangements.
The announcement therefore supports a clear conclusion, but not a winner. AWS is building a multi-silicon AI infrastructure strategy. NVIDIA supplies capacity and an established technology stack, while Trainium gives AWS a custom path it can integrate into its own platform. The result will be judged by delivered services, usable capacity and workload economics, not by the announced GPU count alone.



