Partnerships
D-Matrix Connects Raptor XPUs to NVIDIA AI Factories via NVLink Fusion

AI inference chipmaker d-Matrix announced on September 10, 2026, a multi-year product-roadmap collaboration with NVIDIA that will integrate its next-generation Raptor inference XPUs into NVIDIA’s MGX rack-scale architecture through NVLink Fusion, with initial rack-integrated availability expected in the fourth quarter of 2027.
d-Matrix said the agreement gives its XPUs entry into NVIDIA’s AI factory ecosystem. Its centerpiece is an NVLink Fusion–enabled rack-level system that d-Matrix says will let AI labs, hyperscalers and neoclouds offer premium-tier token services at ultra-low latency. d-Matrix cofounder and CEO Sid Sheth described the collaboration as a defining moment for the company, saying customers will be able to deploy its inference XPUs alongside the broadly available NVIDIA AI factory platform.
NVIDIA said in a blog post published September 10, 2026, that d-Matrix will use NVLink Fusion to connect its Raptor XPUs to NVIDIA’s AI infrastructure platform, adding d-Matrix to its roster of NVLink Fusion ecosystem partners. “Demand for inference is soaring, but capital, time and energy remain finite,” Sheth said during a press briefing, according to NVIDIA. He said NVLink Fusion and MGX let d-Matrix integrate Raptor XPUs into a broadly deployed, liquid-cooled architecture, reducing the time and risk of bringing its inference silicon to large-scale deployment.
In the d-Matrix announcement, NVIDIA founder and CEO Jensen Huang said: “NVLink Fusion enables partners to integrate custom silicon with NVIDIA’s deep ecosystem of NVLink, advanced packaging, rack-scale systems and networking technologies.” Huang added that, with NVIDIA AI infrastructure running across cloud and on-premises data centers worldwide, NVLink Fusion offers partners such as d-Matrix a path to integrate with NVIDIA compute platforms and widens accelerator choice for customers building new AI factories.
Rack Integration and Disaggregated Inference
Under the roadmap, d-Matrix will work with NVIDIA to incorporate its next-generation inference XPUs, starting with Raptor, directly into NVIDIA’s latest rack reference architecture design, which features NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking. The d-Matrix rack, enabled by the MGX platform, will use modular cable-free trays built with NVIDIA’s MGX ecosystem and supply chain.
d-Matrix is also partnering with Astera Labs, a connectivity provider in the NVLink Fusion ecosystem, to deliver custom connectivity for high-throughput data flow throughout the system. Astera Labs CEO Jitendra Mohan said the companies’ partnership brings purpose-built, high-throughput connectivity for low-latency AI inference to the NVLink Fusion ecosystem.
According to NVIDIA, d-Matrix plans to connect its XPUs in a single high-bandwidth, low-latency scale-up domain using NVLink, and those racks are designed to operate next to GPU-based NVIDIA systems, including the Vera Rubin NVL72, in disaggregated inference deployments.
d-Matrix said that with heterogeneous disaggregation, operators can split workloads between Raptor XPUs and Vera Rubin systems to optimize each phase of inference. For AI coding, which the company described as one of the most popular current disaggregated applications, the compute-intensive prefill stage runs on GPUs while the latency-sensitive decode stage runs on its XPUs. The company said the rack system is designed for latency-sensitive applications such as AI coding assistants, real-time chatbots and voice agents, where it says customers are willing to pay a premium for speed.
What NVLink Fusion Provides
NVIDIA describes NVLink Fusion as its platform for integrating third-party custom XPUs and CPUs with NVLink scale-up networking, the MGX rack architecture and the full AI factory platform spanning compute, networking, storage, security, power, cooling and software. NVIDIA reports that the platform delivers 3x lower XPU-to-XPU latency than off-the-shelf Ethernet, 10x higher packet rates and 3 TB/s per XPU of all-to-all bandwidth through sixth-generation NVLink.
In a platform overview published August 24, 2026, NVIDIA stated that sixth-generation NVLink provides high-bandwidth, low-latency networking across a 72-XPU domain, and that NVLink-C2C, which connects XPUs to NVIDIA Vera CPUs or other ecosystem CPUs, delivers up to 6x the energy efficiency of a PCIe interface. The overview said NVLink Fusion aligns with NVIDIA’s DSX reference architecture, which codesigns buildings, power, cooling, compute and networking for AI factories, and that the NVIDIA Omniverse DSX AI Factory Blueprint supplies a digital twin and open reference design for gigawatt-scale facilities. NVIDIA also lists supporting software that includes NCCL for distributed workloads, Dynamo and NIXL for disaggregation, and Mission Control for cluster management, telemetry and debugging.
NVIDIA’s listed NVLink Fusion partners are d-Matrix, AWS, Arm, Intel, Fujitsu, SiFive, Alchip, Astera Labs, GUC, Marvell, MediaTek, Samsung, Cadence, Synopsys, Ayar Labs and Lightmatter, and the platform supports the Arm, x86 and RISC-V CPU architectures alongside NVIDIA GPUs. NVIDIA calls the resulting model a semi-custom AI factory, in which NVLink Fusion provides the interconnect, rack architecture, software and supply chain while the partner concentrates on differentiated silicon. The company said its full-stack AI factory platform includes the Vera Rubin NVL72, Groq 3 LPX, the Vera CPU rack, Vera BlueField-4 STX storage and Spectrum-6 SPX Ethernet networking.
Raptor Architecture and Availability
Raptor is the follow-on to d-Matrix’s Corsair XPU platform, which is currently in production, and extends the company’s memory-centric architecture. Through what d-Matrix describes as a first-of-its-kind 3D DRAM stacking approach, Raptor pairs a DRAM memory chip with an SRAM compute chip in what the company calls a single two-story package. d-Matrix cofounder and CTO Sudeep Bhoja previewed the 3D DRAM technology at the 2026 Hot Chips conference.
d-Matrix said it designed Raptor from the ground up for integration with NVLink Fusion and the NVIDIA MGX rack-scale ecosystem. The company expects Raptor to tape out before the end of 2026 and said the XPU is being actively evaluated at AI hyperscalers and frontier labs; d-Matrix states the platform is backed by more than 100 patents.
Initial availability of Raptor XPUs integrated into the NVIDIA MGX rack is expected in the fourth quarter of 2027. d-Matrix is demonstrating the technology at its booth at the AI Infra Summit.












