Ramp, the corporate expense management company, went public with a product called Router on August 19, 2026 — a single API endpoint that dispatches developer requests to whichever AI model delivers the right result at the lowest price. The company claims early customers have cut their AI inference bills by 40% on average, a number that lands with some weight given that AI has become, in Ramp's own index, one of the fastest-growing expense categories in the modern tech stack.
Background You Need
The model routing problem is not new, but it has become urgent. A year ago, a company running AI workloads might have been calling one or two APIs; today, the realistic landscape includes OpenAI, Anthropic, open-weight models from DeepSeek and Alibaba's Qwen team, inference providers like Fireworks AI, and a handful of others all competing on different price-to-performance tradeoffs depending on the task. Routing requests manually — or defaulting to a single frontier model for everything — is expensive and often unnecessary. A customer service triage query does not need the same compute budget as a multi-step code generation task.
OpenRouter emerged as the dominant intermediary for this problem, offering a unified API layer that developers could drop in and immediately access dozens of models. It worked well enough that Stripe reportedly agreed to acquire it for approximately $7.5 billion — a deal that, if it closes, hands control of the most-used neutral routing layer to a major financial infrastructure company. That context matters: developers who built OpenRouter into their stacks now have a reason to ask whether a Stripe-owned version will stay as permissive, as independently priced, and as open to the full model catalog going forward.
Ramp, meanwhile, has been running a version of Router internally for three years. This is not a startup with a proof of concept — it is a company that built a routing system to manage its own rising AI costs, watched it work, and then decided to open it to the market at the moment OpenRouter's neutrality is most in question.
What's New
Router, now live at router.com, provides a single API key that connects to models from OpenAI, Anthropic, xAI, DeepSeek, Nvidia, Kimi, GLM, Qwen, and others — with Gemini listed as coming soon. On the infrastructure side, providers like Fireworks AI are already connected, with Google, AWS, and additional compute providers reportedly in the pipeline. The Ramp product page describes the offering as "one endpoint, one bill, every model," and the company has set pricing to zero through the end of 2026, bundling $26 in model credits for new accounts.
The differentiation Ramp is pitching beyond the free tier is its routing logic. Rather than simply load-balancing or sending every request to the cheapest available model, Router lets developers specify performance requirements using up to three benchmark dimensions — think latency ceiling, cost target, and a task-specific capability score like SWE-bench for coding workloads. The system then selects the model that satisfies all three constraints at the lowest cost. Ramp's own cost-impact data, shown on the product page, charts a steady decline in spending as the routing share shifts from fixed to flexible: over eighteen sampled data points, a team that started with fully fixed routing and gradually increased flexible routing saw their total cost index drop from 100 to 70 — a 30% reduction from routing changes alone, before any model-level optimization.
The dashboard surfaces the metrics that matter for this kind of cost management: token spend by model, latency distributions, fallback attempts when a preferred model is unavailable, and a consolidated bill across providers. Ramp's CTO, Rahul Sengottuvelu, framed the problem in Ramp's press release as an expense-visibility failure: AI has become one of the largest and fastest-growing line items at many companies, and also one of the least measurable. The pitch is that Router fixes the measurement problem first, then optimizes spend on top of it — which is a natural angle for an expense management company to take.
According to Ramp's AI Index, enterprise AI spending has grown 20.7 times over since June 2025. That figure comes from Ramp's own transaction data across its corporate card and expense platform, giving it a visibility into actual AI line items that most analysts lack. Our read is that this number, even if it overstates the typical company's experience, reflects a real and broadly confirmed trend: AI spend is no longer a pilot budget category, it is a meaningful recurring operating cost that finance teams are starting to scrutinize.
The Pushback
The 40% cost-savings claim deserves scrutiny. Ramp reports this as an average across "customers already using Router," but the baseline matters enormously. A team that has been sending every request to GPT-5.6 Sol or Anthropic's Opus for months will see dramatic savings the moment cheaper models handle a share of their workload — not because the routing is especially clever, but because the starting point was inefficient by design. Teams that have already done even basic model tiering will see a much smaller delta.
The data retention policy is also a meaningful limitation that Ramp's marketing language does not lead with. Router is not a zero-data-retention service by default. Inputs and outputs are stored for up to one year, with opt-out as the mechanism rather than opt-in. For developers working in regulated industries, handling user-generated content, or routing any prompts that touch personally identifiable data, that default is a problem. Ramp's product page does note that US-hosted options with ZDR are available, but the default configuration is not zero-retention — and "opt-out" is the kind of detail that gets missed during rapid onboarding.
There is also the question of what happens after 2026. Router is free through December 31, but no pricing structure for 2027 has been announced. For a team that rebuilds infrastructure around a single routing endpoint, that dependency is not trivial to unwind. Ramp is a well-funded company and has no obvious reason to price developers out, but the absence of any stated post-free-tier terms means teams are signing up without knowing what the service costs when the free window closes.
The broader competitive landscape is also not static. OpenRouter's Stripe acquisition, if completed, may not produce the constraints developers fear; Stripe has historically been developer-friendly and has no obvious incentive to restrict model access. Ramp entering the market as the "neutral alternative" works as a narrative today, but that framing depends on OpenRouter actually behaving differently post-acquisition — which has not happened yet.
Sources
techcrunch.com 2026 | TechCrunch Ramp Launches Router.com to Cut Companies' Rising AI Bills Latest News | TechCrunch Ramp Router: The LLM Gateway That Cuts Inference CostsBased on
https://techcrunch.com/2026/08/20/ramp-launches-its-own-ai-model-router-called-router/— techcrunch.comThis article is an original, AI-assisted summary and analysis. Credit for the underlying reporting or footage belongs to the source above.

Written by the vybecoding.ai editorial team
Published on August 20, 2026