AI inference chip developer d-Matrix announced on September 10, 2026, that its next-generation Raptor XPUs will connect to NVIDIA’s infrastructure through NVLink Fusion and the MGX rack architecture. Reuters, NVIDIA and d-Matrix descriptions confirm that the collaboration is intended to place specialized inference processors inside a broadly deployed data-center platform. The companies did not disclose financial terms. The announced roadmap focuses on technical integration, networking and the deployment of complete rack-scale systems rather than an immediate commercial shipment.
The infrastructure requirement for generative AI is expanding from model training toward inference. Training creates and refines a model, while inference is the repeated process of generating responses during everyday use. Coding assistants, chatbots and voice agents need low latency and high throughput, sometimes under tight power and memory limits. d-Matrix develops memory-centered processors for this stage of the workload. Connecting Raptor XPUs to NVIDIA’s rack platform is meant to let customers add the specialized accelerators without building an entirely separate data-center environment.
NVIDIA says the planned systems will combine NVLink scale-up connectivity, Spectrum-X scale-out networking and the MGX rack architecture. d-Matrix is also working with Astera Labs on custom connectivity intended to move data efficiently across the system. Reuters reported that the final design stage for Raptor is scheduled to be completed by the end of 2026, while compatible racks are expected in 2027. StorageReview identified the fourth quarter of 2027 as the target for initial systems. These dates remain company targets and should not be treated as completed availability.
NVLink Fusion extends NVIDIA’s interconnect approach beyond NVIDIA GPUs by allowing partners to attach custom processors to the wider platform. In a heterogeneous system, different accelerators can handle the parts of a workload for which they are best suited while sharing processors, networking, cooling and management components. For d-Matrix, using an established rack design may reduce the engineering and supply-chain work required to create a separate standard. For NVIDIA, the strategy can keep its interconnect and networking technology relevant as more companies develop custom AI silicon.
The announcement does not mean that a finished Raptor rack is already shipping. Chip design completion, system validation, software compatibility and customer qualification remain ahead. Performance, energy-efficiency and cost claims from vendors will also need independent testing on real workloads before firm comparisons are possible. The verified development at this stage is the decision to integrate Raptor with NVLink Fusion, MGX and Spectrum-X, together with a multi-year product roadmap toward rack-scale deployment.
For customers, operating cost, power demand, cooling, software support and model portability will matter alongside raw specifications. Those factors can only be measured reliably after production systems become available and undergo independent testing.
The collaboration illustrates the broader movement toward heterogeneous computing in AI data centers. Instead of requiring one processor type to handle training, data movement and inference, operators can assign different stages to specialized hardware. Reuters, NVIDIA’s technical blog, d-Matrix’s announcement, HPCwire, The Register and StorageReview independently support the core details of the integration and timing. The most important measures will come later: whether systems arrive on schedule, whether the software toolchain is reliable and whether customer deployments demonstrate repeatable latency, throughput and efficiency gains under transparent testing conditions.
