Partnerships

Thinking Machines Lab Signs $65M Annual Deal for Crusoe Cloud Inference

mm
Add Unite.AI to your preferred sources on Google

Crusoe and Thinking Machines Lab announced a $65 million annual agreement on September 23, 2026, to power inference for the lab’s open models on Crusoe Cloud, with production workloads served on a dedicated deployment of NVIDIA HGX B200 systems through Crusoe Managed Inference.

The announcement, datelined San Francisco, states that Thinking Machines Lab will serve a range of production workloads, including its Inkling models, GLM 5.2 and 5.3, and its own fine-tuned variants, on a dedicated Tailored Deployment of NVIDIA HGX B200 systems connected with NVIDIA Quantum-2 InfiniBand networking. Crusoe characterized the deployment as engineered for high-throughput, cost-efficient serving at scale.

Crusoe will run and support the cluster directly as a Tailored Deployment, which the release describes as a dedicated, benchmarked, SLA-backed endpoint intended to give Thinking Machines Lab price-performance and headroom to scale without managing its own inference stack. In the release, Crusoe describes itself as the industry’s first vertically integrated AI infrastructure provider.

Managed Inference Options and the MemoryAlloy Engine

On Crusoe’s Managed Inference product page, the company presents Tailored Deployments as its highest level of optimization: the customer brings the model and defines the requirements while Crusoe handles the rest, with a dedicated, benchmarked endpoint positioned for proprietary models, bespoke inference optimization, and SLA-backed performance.

The page also details Serverless Inference, usage-based consumption of open-source models through a fully managed API in Crusoe Intelligence Foundry, and Self-Serve Deployments, now generally available, which run models in the Foundry with inference optimized for throughput or responsiveness, including models customized with the company’s Serverless Fine-Tuning feature. The Foundry itself is presented as a developer platform for discovering models, generating API keys, monitoring performance metrics, and deploying managed endpoints.

Crusoe says its inference engine is powered by MemoryAlloy, a cluster-native memory fabric that persists across nodes and enables intelligent routing to eliminate redundant prefill computation. The company reports up to 9.9x faster time-to-first-token and 5x higher throughput using speculative decoding and dynamic batching, figures it benchmarked against vLLM serving the Llama-3.3-70B model in a four-node deployment.

The Inkling and GLM Models

The agreement’s named workloads include the lab’s own Inkling family. Thinking Machines Lab released Inkling on July 15, 2026, describing it as an open-weights Mixture-of-Experts transformer with 975B total parameters and 41B active parameters, supporting a context window of up to 1M tokens. According to the lab, Inkling was pretrained on 45 trillion tokens spanning text, images, audio and video, and the effort was the lab’s first major training run, conducted on NVIDIA GB300 NVL72 systems.

Alongside Inkling, the lab previewed Inkling-Small, a lighter-weight 276B-parameter Mixture-of-Experts model with 12B active parameters that carries a different performance and latency trade-off. Inkling’s full weights are published on Hugging Face, both as the original checkpoint and as an NVFP4 checkpoint for efficient inference on NVIDIA Blackwell systems, and the model is available for fine-tuning on Tinker, the lab’s model-customization platform, with 64K and 256K context length options.

The agreement also places GLM 5.2 and 5.3 among the lab’s production workloads on the service. Crusoe’s Managed Inference model hub lists Z.ai’s GLM 5.3 at 753B parameters with a 1,000,000-token context and GLM 5.3 Flash at 320B parameters, both available for Serverless Inference.

Executive Statements and Planned Expansion

“Crusoe Managed Inference took us from evaluation to production quickly and met the standards we hold our own systems to, giving us the scale and economics to reinvest in the research at the core of what we do,” said Myle Ott, ML Infra Lead at Thinking Machines Lab.

Erwan Menard, SVP of Product Management for Crusoe Cloud, said Thinking Machines Lab has a sophisticated engineering team that expects partners to meet its high bar as it grows. He said Crusoe built Managed Inference to deliver cost-efficient, reliable infrastructure, inference, and model lifecycle tooling as a service so that companies can run their chosen models in production at scale.

The companies said they are also looking to expand into batch inference for large-scale synthetic data generation, which the release frames as turning inference into a flywheel for research. Crusoe said Thinking Machines Lab joins a customer roster that has taken Crusoe Managed Inference past $100 million in contracted ARR less than a year from launch.

Theo Nash is an AI-generated specialist at Unite.AI, covering AI infrastructure, compute, and the hardware systems that power modern artificial intelligence. His work focuses on the technical foundations behind large-scale AI workloads, including data centers, accelerators, networking, and the software stacks that tie them together.

With an analytical and engineering-driven perspective, Theo examines how advances in GPUs, custom silicon, memory architectures, and distributed systems enable new generations of AI models. He pays particular attention to performance trade-offs, energy efficiency, scalability, and the practical constraints that shape real-world deployment of AI infrastructure.

Articles authored by Theo Nash are AI-generated and reviewed by Unite.AI’s editorial team to ensure technical accuracy, clarity, and responsible coverage of the rapidly evolving AI compute landscape.