Partnerships
NVIDIA and Palantir Announce Sovereign AI Stack for Supply Chains

NVIDIA and Palantir Technologies announced on September 10, 2026, a collaboration to bring sovereign AI to critical supply chains, with the jointly built AI stack first being deployed inside NVIDIA’s own supply chain operations.
The stack brings NVIDIA’s Nemotron open models into Palantir Foundry and Palantir’s Artificial Intelligence Platform (AIP), grounded in the Palantir Ontology. The companies said the system gives supply chain teams a shared command center, beginning with materials allocation decisions, while customers retain control and ownership of their proprietary data. The stack will be showcased at Palantir’s AIPCon 11 conference, and the companies said organizations in agriculture, manufacturing, pharmaceutical, retail, technology and government can apply it to their own supply chains.
NVIDIA said its supply chain spans millions of parts, thousands of suppliers and a global network of manufacturing partners, and that bringing a rack-scale AI system to production requires the coordinated availability of compute, memory, networking, power, cooling and mechanical components. Each NVIDIA Vera Rubin rack contains 1.3 million parts, according to the announcement.
In the joint announcement, Palantir cofounder and CEO Alex Karp said: “NVIDIA has arguably the most valuable, intricate and complex supply chain in the world. Our sovereign stack, powered by Nemotron models and Ontology, is delivering capabilities that exceed the frontier while providing alpha protection qualities unavailable otherwise.”
NVIDIA founder and CEO Jensen Huang said the companies aim to turn the operational graph behind AI infrastructure into sovereign intelligence, pairing Nemotron models with Palantir’s Ontology to coordinate the path from wafer production to token output.
NVIDIA’s Supply Chain and the Foundry Command Center
In a technical blog post published alongside the announcement, NVIDIA said it measures supply chain performance from wafer-out to first token. The interval splits in two: time-to-rack runs from silicon leaving the fab to an assembled system arriving on a data center floor, and time-to-token covers the power, cooling, networking and software work that makes the infrastructure productive.
Grace Blackwell NVL72 platforms draw on millions of parts and thousands of suppliers, with final systems assembled by dozens of OEMs and ODMs. A single GB200 NVL72 compute tray — one of eighteen in a rack — requires two Grace CPUs, four Blackwell GPUs and 32 HBM3e memory stacks, and NVIDIA said the supply chain it created for Vera Rubin is twice as large as the one behind Grace Blackwell. To track how long material sits idle, NVIDIA runs a metric called Time of Ownership: the elapsed time between a manufacturing site receiving material and that material leaving as part of a sub-assembly or product.
Each week, planners manually rework what NVIDIA calls the critical material allocation problem, deciding what material and how much of it goes to each manufacturing site across the current quarter and the next. To support that work, NVIDIA’s supply chain operations team built a Digital Supply Chain Intelligence command center with Palantir on Foundry, where the Ontology connects materials, manufacturing sites, commits, capacity, allocations, production outputs and unstructured qualitative signals into one governed data layer.
NVIDIA cuOpt, the company’s open-source library for GPU-accelerated decision optimization, solves the weekly allocation as a mixed-integer linear program whose objective is to minimize Time of Ownership, drawing inputs from the Ontology and writing results back as allocation decisions. The post reports that cuOpt also identifies which constraints are binding, so planners can see whether capacity or memory supply limited a given week, and can test scenarios such as a ten percent cut in memory supply or a new manufacturing site coming online.
Post-Training Nemotron on Allocation Decisions
NVIDIA and Palantir back-tested historical allocation decisions against what actually happened and found that human planners outperformed the solver, the post reports, because planners acted on information it could not see: emails exchanged with partners, severe weather forecasts for key regions, geopolitical events and supplier debrief transcripts, alongside years of accumulated expertise. The teams rebuilt the workflow around that expertise, capturing each allocation decision, its rationale, the expected result and the actual outcome inside the Ontology.
That record became training data. The teams post-trained Nemotron 3.5 Lightning, an open-weight mixture-of-experts model with 30 billion parameters and roughly 3 billion active per forward pass, which the post says was selected because it is purpose-built for the execution layer of an agentic workflow and because open models can be post-trained inside an organization’s own compute boundary. The pipeline uses NeMo Anonymizer to strip personally identifiable information, NeMo Data Designer to generate and balance synthetic examples, and NeMo AutoModel to train a small set of LoRA adapter parameters while the base weights stay frozen, with Palantir Autopilot managing the lifecycle from Ontology data to deployed model.
On the development benchmark described in the post, the post-trained model reached 86.7% allocation-decision accuracy, compared with 55.5% for the larger Nemotron 3 Ultra and 17.5% for the base Nemotron 3.5 Lightning. NVIDIA reports balanced accuracy of 58.6% against Ultra’s 42.0% and a Macro-F1 score of 57.5% against 39.5%. The authors say those metrics weight each decision type equally, which matters because planners cut allocations far more often than they raise them.
The post cautions that the gains are concentrated in the domain the model was post-trained on and that forecasting future production risk remained difficult despite fine-tuning. The LoRA training run finished on two NVIDIA B200 GPUs in minutes, light enough to repeat as feedback accumulates, according to the post.
The model recommends an allocation range with attached rationale and risks, but a human planner reviews each recommendation and makes the final call. Every acceptance, edit, override and production outcome is written back into the Ontology until enough representative data exists for another governed training run. The post states that the model never retrains itself in production and that the accumulated feedback is intended to support reinforcement learning in the future.
Sovereign Deployment and Infrastructure Partners
Given the sensitivity of NVIDIA’s supply chain, the deployment runs on NVIDIA reference architectures and the jointly developed Palantir Sovereign AI Operating System Reference Architecture, known as SAIOS, which is supported by Dell Technologies and Cisco. The companies said the stack can be deployed on premises with system manufacturers including Cisco and Dell, or in co-location and cloud environments with Rackspace and Nebius.
Within Palantir AIP, cuOpt supports optimization and scenario planning, while NeMo Data Libraries, NeMo AutoModel and NeMo RL integrate with Palantir Autopilot in what the companies describe as a governed learning loop that measures decisions against real-world results. The companies said they plan to extend what they learn from NVIDIA’s deployment to enterprises across manufacturing, energy, healthcare, automotive and aerospace.












