HomeArtificial Intelligence

Artificial Intelligence

AMD Says ROCm Lifted MI355X Inference by Up to 38% Without New Hardware

AMD's MLPerf Inference 6.1 results show sizable gains on the same Instinct MI355X hardware, highlighting how quickly software optimization is reshaping the AI accelerator race.

AMD corporate logo
AMD corporate logo
Research-based guidePrimary references and a decision framework are included below.How we research →

AMD's latest MLPerf results make a case that the AI accelerator race is increasingly a software race as much as a silicon race. In MLPerf Inference 6.1, published September 16, AMD says the same Instinct MI355X hardware used in the previous benchmark round delivered up to 38% more GPT-OSS-120B server throughput after continued work on ROCm and the surrounding inference stack.

That same-hardware comparison is the most useful part of the announcement. New accelerator generations routinely produce bigger benchmark numbers, but those comparisons mix architectural improvements, memory changes, power budgets and software updates. Here, the underlying MI355X GPU did not change. AMD attributes the improvement to software optimization.

The result does not mean every production workload will suddenly run 38% faster after an update. MLPerf is a standardized benchmark suite with specific models, configurations and performance rules. But it does demonstrate that the amount of useful inference work extracted from an already deployed accelerator can change substantially over time.

The same MI355X hardware did more work

AMD says an eight-GPU MI355X system improved GPT-OSS-120B throughput by 28% in MLPerf's Offline scenario and 38% in Server compared with its Inference 6.0 submission. For the Wan 2.2 text-to-video workload, SingleStream performance improved 70%.

At cluster scale, AMD reports that 72 MI355X GPUs in round 6.1 produced more GPT-OSS-120B throughput than 94 of the same GPUs did in round 6.0. That is an important infrastructure metric because operators care not only about peak chip performance but about how much demand an installed fleet can serve.

If software lets a cluster process more requests while staying inside its latency target, an operator may be able to delay expansion, serve additional customers or reduce the number of accelerators assigned to a workload. Those outcomes can affect power, cooling and capital costs even though no physical hardware has changed.

ROCm maturity is becoming easier to measure

For years, one of the central questions around AMD's Instinct business has been whether ROCm could mature quickly enough to compete with NVIDIA's deeply established CUDA ecosystem. That is broader than raw accelerator speed. Framework support, kernels, model compatibility, distributed execution, debugging tools and deployment reliability all influence whether theoretical compute turns into useful application performance.

MLPerf cannot answer every part of that question, but repeated tests on identical hardware create a measurable signal. Instead of relying only on claims that a software ecosystem is improving, teams can compare how a fixed platform performs across successive benchmark rounds.

AMD says ROCm 7 powered its 6.1 submissions. The company has also been moving toward a faster software release cadence, making software velocity a more explicit part of its accelerator strategy.

The practical implication is that organizations evaluating AI infrastructure should not freeze their performance assumptions at purchase time. A benchmark captured when a GPU launches may become stale as kernels, runtimes and serving frameworks improve.

AMD also posted competitive cross-vendor results

AMD highlighted selected comparisons against NVIDIA B200 and B300 systems. On GPT-OSS-120B with eight GPUs, AMD says its MI355X submission exceeded selected B200 results by 33% in Offline and 26% in Server, and selected B300 results by 14% and 13% respectively.

Those figures are worth watching, but they need the usual benchmark discipline. A result for one model and one configuration is not evidence that one accelerator is universally faster. Production systems differ in model mix, context lengths, latency requirements, quantization, networking, batching, power limits and software versions.

The stronger signal is that AMD is participating across a wider set of MLPerf workloads and configurations. MLCommons says Inference 6.1 received submissions from a record 30 organizations, including AMD, NVIDIA, Google, Intel, Microsoft Azure, CoreWeave, Oracle, Dell and multiple system vendors. That broader participation makes standardized comparisons more useful.

Scale and reproducibility matter beyond a single node

AMD's submission also emphasized multi-node inference. The company says a 72-GPU MI355X configuration retained 95% scaling efficiency on GPT-OSS-120B. Through Crusoe, a 512-GPU submission reached 5.75 million Offline tokens per second on the model, which AMD describes as the highest aggregate token throughput submitted to MLPerf so far.

Partner participation is another relevant detail. AMD says systems submitted by seven partners produced comparable MI355X results that landed within 4% of AMD's numbers on average. Reproducibility across vendors matters because a reference-system benchmark is less useful if customers cannot approach it with systems they can actually deploy.

What infrastructure buyers should take from the round

The headline should not be that one benchmark settles the AMD-versus-NVIDIA debate. It does not. The more durable lesson is that AI infrastructure has a moving performance envelope.

For buyers, that means accelerator evaluation should include the software roadmap, not only memory capacity, theoretical FLOPS and purchase price. Teams should ask how quickly a vendor supports new models, whether performance improvements reach existing hardware, how easily results reproduce across partner systems and whether the serving stack works with their operational tooling.

For organizations already running Instinct MI355X, the round provides a reason to benchmark newer ROCm and inference-stack releases before assuming additional demand requires additional GPUs. For prospective buyers, it provides a clearer data point that AMD's software stack is capable of extracting more performance from a fixed generation over time.

MLPerf remains a benchmark rather than a substitute for testing a real application. But same-hardware gains are unusually informative because they isolate something infrastructure teams can often change without replacing the data center: the software.

Editorial research note

How we reached this guidance

We reviewed AMD's September 16 MLPerf Inference 6.1 announcement, MLCommons' independent round announcement and analysis of the same-hardware ROCm gains. Vendor-to-vendor comparisons are identified as benchmark results rather than generalized real-world superiority, and we focus on the unusually useful same-hardware comparison between rounds 6.0 and 6.1.

Decision framework

ScenarioRecommendationWhy
An infrastructure team assumes accelerator performance is fixed when hardware shipsTrack software-stack performance across benchmark roundsAMD reports substantial throughput and latency gains on unchanged MI355X hardware, showing that serving software can materially change deployed-system economics.
A buyer treats one MLPerf result as proof of superiority for every workloadMatch benchmark scenario and model to the intended production workloadMLPerf contains different models and serving scenarios, while production performance also depends on latency targets, batching, networking and application behavior.
A team is comparing AMD and NVIDIA from vendor chartsCheck the underlying MLPerf submissions and system configurationsSelected comparisons can be valid within a benchmark but do not automatically describe every configuration or deployment constraint.
An operator already owns MI355X systemsEvaluate current ROCm and inference-stack releases before buying more hardwareThe round-over-round results indicate that software updates can unlock meaningful additional capacity from existing accelerators.

Primary references

Reviewed on September 16, 2026. Unless an article explicitly states that TECHMUNDI performed hands-on testing, our guides are research-based and do not present specification or documentation review as first-hand product testing.