d-Matrix Is Bringing Its AI Inference Chips Into NVIDIA's NVLink Fusion Ecosystem
d-Matrix plans to connect its Raptor inference processors directly into NVIDIA rack-scale infrastructure, showing how the AI hardware race is expanding beyond training GPUs.
The AI hardware race is increasingly splitting into two related but different problems: training the largest models and serving those models efficiently once millions of users begin asking them questions. A new collaboration between d-Matrix and NVIDIA is a useful example of how the second problem—AI inference—is reshaping data-center architecture.
d-Matrix announced on September 10 that it plans to integrate its next-generation Raptor XPUs into NVIDIA's rack-scale infrastructure using NVLink Fusion. Reuters also reported the agreement, describing it as a path for d-Matrix's inference-focused processors to connect more directly with NVIDIA-centered AI server systems.
The announcement does not mean d-Matrix chips are replacing NVIDIA GPUs. Instead, it points toward a more modular future in which specialized accelerators can plug into an ecosystem of NVIDIA CPUs, networking, switches and rack designs.
Inference is becoming its own hardware market
Training receives most of the attention because building a frontier model can require enormous clusters of accelerators running for weeks or months. But once a model is trained, the economics change.
Inference is the work performed every time a chatbot generates a response, a coding assistant writes a function, a voice model processes a conversation or an enterprise agent executes a task. The workload happens continuously and can become extremely expensive at large scale.
For many inference services, the critical metrics are not only raw compute. Providers care about how quickly the first token appears, how many tokens a system can generate per second, how many users can share the hardware and how much power and memory are required for each completed request.
That creates room for processors designed specifically around inference rather than general-purpose GPU workloads.
d-Matrix has positioned its architecture around that opportunity. Its Raptor processors are intended for low-latency generative-AI inference, particularly the premium token-serving workloads used by AI labs, hyperscalers and specialized cloud providers.
Why NVLink Fusion matters
A specialized accelerator is only useful at data-center scale if it can connect efficiently to the rest of the system.
Modern AI racks contain much more than accelerator chips. They include CPUs, network interfaces, high-speed switches, memory systems, storage and software that has to coordinate data movement across the entire machine.
NVIDIA introduced NVLink Fusion as a way for partners to integrate custom processors into this broader architecture. Instead of forcing a customer to build an entirely separate island around a new accelerator, a partner can connect its technology into parts of the NVIDIA rack ecosystem.
According to d-Matrix, its planned system will combine Raptor XPUs with NVIDIA Vera CPUs, NVLink switches, BlueField-4 data-processing units, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking. The company is also working with Astera Labs on high-speed connectivity inside the system.
That is strategically important because NVIDIA's advantage in AI infrastructure is larger than the GPU itself. Networking, rack design, software and deployment standards can make it difficult for a competing accelerator to enter an existing environment even if the chip performs well on a benchmark.
NVLink Fusion gives NVIDIA a way to remain part of the architecture even when another company's silicon executes part of the workload.
NVIDIA can benefit even when another chip does the inference
At first glance, helping alternative accelerators connect into NVIDIA systems might seem counterintuitive. NVIDIA sells GPUs, and specialized inference chips compete for some of the same computing budget.
But the broader strategy can expand NVIDIA's role in the data center.
If a cloud provider chooses a d-Matrix accelerator but still uses NVIDIA CPUs, networking, switches and rack architecture, NVIDIA continues capturing value from the deployment. The ecosystem also becomes harder to displace because third-party hardware is designed to operate inside it rather than around it.
This resembles a platform strategy more than a single-chip strategy.
It also acknowledges that AI infrastructure is becoming heterogeneous. The most efficient processor for training a huge multimodal model may not be the most economical choice for serving a particular model at high volume. A future AI factory could combine different accelerators for different stages of the workload.
The announcement is a roadmap, not a benchmark result
There is an important limitation to the news: the integrated Raptor systems are not yet a mature product with independently verified production performance.
Reuters reported that d-Matrix expects its next-generation chip design to be completed toward the end of 2026, with integrated systems expected later. The companies are describing a multi-year product roadmap.
That means buyers should not interpret the announcement as proof that Raptor already provides better economics than current GPUs or other inference accelerators.
Real systems will need to demonstrate latency, throughput, energy efficiency, reliability, software compatibility and total cost under production workloads.
The infrastructure around a chip also matters. A theoretically efficient accelerator can lose its advantage if data movement, memory limits or software integration prevent the hardware from reaching high utilization.
Cost per token is becoming a major competitive metric
As AI services mature, the economics of inference are becoming increasingly visible.
A provider serving billions of tokens can save meaningful amounts of money from even modest improvements in hardware utilization or energy efficiency. At the same time, users increasingly expect AI systems to respond faster while models become larger and more computationally demanding.
This creates a strong incentive for specialized inference hardware.
But buyers should measure the complete service, not only a chip's theoretical performance. Power consumption, networking, memory, software support, utilization and operational complexity all affect the real cost of delivering a token to a customer.
The same principle applies when evaluating AI software. Our guide to choosing AI tools for work recommends measuring the full workflow rather than comparing one headline metric in isolation.
What this says about NVIDIA's position
The d-Matrix deal suggests NVIDIA is preparing for a market in which it does not need every AI operation to run on an NVIDIA GPU for the company to remain central to AI infrastructure.
If NVLink Fusion attracts more custom accelerators, NVIDIA can increasingly become the connective fabric of heterogeneous AI systems: CPUs, networking, rack architecture and interconnects surrounding a mixture of processors.
For challengers such as d-Matrix, that can reduce the friction of entering data centers already built around NVIDIA technology. For NVIDIA, it can turn potential hardware competition into ecosystem participation.
Bottom line
The d-Matrix collaboration is not evidence that NVIDIA GPUs are being displaced. It is evidence that the AI infrastructure market is becoming more specialized.
Training and inference have different economic pressures, and inference at massive scale creates room for purpose-built accelerators. NVLink Fusion gives those processors a path into an established rack ecosystem while allowing NVIDIA to remain deeply embedded in the system.
The real test will come when Raptor-based racks reach production. Until then, the most important part of the announcement is architectural: the next phase of AI hardware may be less about one chip winning everything and more about who controls the platform that lets many kinds of chips work together.
Editorial research note
How we reached this guidance
We reviewed d-Matrix's September 10 announcement, NVIDIA's own ecosystem listing and Reuters coverage of the collaboration. The article separates announced integration plans from products already shipping and focuses on what the deal reveals about inference economics and rack-scale AI architecture.
Decision framework
| Scenario | Recommendation | Why |
|---|---|---|
| An AI infrastructure buyer assumes every workload should run on the same accelerator | Match hardware to training and inference characteristics | Inference emphasizes latency, throughput and cost per generated token differently from large-scale training. |
| A custom accelerator must fit into an existing NVIDIA-centered data center | Evaluate interconnect and rack compatibility, not just chip benchmarks | NVLink Fusion is designed to let partner processors participate in a broader rack architecture with NVIDIA networking and CPUs. |
| A vendor announces a future rack-scale product | Separate roadmap commitments from deployable systems | d-Matrix says Raptor integration is part of a multi-year roadmap, so production availability and real-world performance still need validation. |
| Teams compare accelerator cost using chip price alone | Measure cost per useful inference outcome | Power, networking, memory, utilization, latency and software integration can matter as much as the accelerator's purchase price. |
Primary references
- d-Matrix: Raptor and NVIDIA NVLink Fusion collaboration
- Reuters: d-Matrix to use NVIDIA chip-linking technology in AI servers
- NVIDIA newsroom: d-Matrix adopts NVLink Fusion
Reviewed on September 13, 2026. Unless an article explicitly states that TECHMUNDI performed hands-on testing, our guides are research-based and do not present specification or documentation review as first-hand product testing.