AI Models & Platforms
Ramp Opens Router.com Model Gateway With Free Routing Through 2026

Ramp, the corporate-spend platform that powers more than $200 billion in purchases annually, launched Router.com on August 19, 2026: a single API endpoint that sends each AI request to the lowest-cost model that meets the developer’s quality bar. Customers already using the service have cut their inference costs by 40% on average, according to the company’s announcement. Routing is free through 2026, with users paying list price for the tokens they consume and new users receiving $26 in credits.
The product is the public version of tooling Ramp built for itself three years ago. By matching each internal job to the right model, the company says it cut its own inference costs roughly 30% for the same output while holding 99.9%+ reliability across production traffic. The Router site puts current production volume at more than 2.75 trillion tokens routed monthly.
“AI is the fastest-growing line item at most companies, and the one they can least measure,” said Rahul Sengottuvelu, Ramp’s chief technology officer, in the announcement. “Router puts every token in one place and sends each request to the model that delivers the right performance at the right cost.”
The timing tracks Ramp’s own data. The Ramp AI Index, which tracks business AI spending across the platform’s customer base, shows AI spend has grown 20.7x since June 2025, the announcement states.
One Endpoint, 27 Models and Counting
Developers connect through one OpenAI- and Anthropic-compatible API and reach models from OpenAI, Anthropic, and SpaceXAI, with Gemini support listed as coming soon. Open-weight models including Nvidia, Kimi, DeepSeek, GLM, and Qwen are served through providers such as Fireworks AI, with Google, AWS, Together AI, Baseten, and Crusoe on the way, per the announcement. The model picker on Ramp’s Router page currently spans 27 models, from Claude Opus 5 and GPT-5.6 Sol down to budget tiers like GPT-5.4 Nano and DeepSeek V4 Flash. Grok’s inclusion extends a distribution push that saw Grok 4.6 reach general availability on Amazon Bedrock the same day.
The routing engine applies more than 100 optimizations across model selection, caching, compression, timing, and request handling, with automatic fallback when a provider fails. Teams can accept Ramp’s default routing strategy or configure their own cost-performance priorities per request type. All models are hosted on U.S.-based infrastructure, with zero-data-retention options available.
A Benchmark Built From Production Work
The load-bearing piece is Ramp SWE-Bench, a benchmark built from the company’s real production engineering tasks rather than public leaderboards. Router continuously tests new models against that workload and folds winners into its defaults automatically. The published results table shows the spread the router exploits: Claude Opus 5 solves tasks at an average cost of $1.84 per run while Qwen3.7 Plus reaches a comparable solve rate at $0.15, and GPT-5.4 Nano handles cheaper work at $0.09.
That kind of per-request arbitrage is where inference economics have been heading as the model menu widens. Benchmarked serving costs now vary by more than an order of magnitude across models for comparable solve rates, a spread reflected in Ramp’s SWE-Bench table and in cheaper-inference plays like Runware’s containerized data centers.
What Customers Report and What the Terms Say
Early users cited on the product page report savings well above the 40% average. Valentin De Matos of Delphi says his company runs billions of tokens through Router and reduced model costs by 92%. Anthropic’s head of platform engineering, Katelyn Lesse, said Claude is available through Router from day one, and SpaceXAI said Grok’s inclusion expands how teams can bring its models into their applications.
The fine print: free routing runs through 2026, after which Ramp has not set pricing in the announcement. Router stores model inputs, outputs, and metadata and uses that data to improve the service, though users can turn this off in settings and select U.S.-hosted models with zero data retention. The service is limited to U.S. developers and teams at launch, with enterprise features and more countries listed as coming soon. Ramp says no Ramp card or company account is required, and switching an existing integration is a one-line base-URL change for OpenAI or Anthropic SDK users.












