Start Free

Real Enterprise Queries. Honest, Verified Results.

Verified Cost Savings benchmark showing DeepSeek 88.3%, Qwen 86.3%, OpenAI 84.9%, Azure OpenAI 84.9%, Google 84.4%, Anthropic 78.8%, Mistral 77.7%, Cohere 73.9%, and Grok 30.9%.

We benchmarked CARDIAC-PURR against 100 real enterprise queries per provider, live, with real spend — not a synthetic easy-classification task. Nine providers were included in this specific benchmark run for testing purposes; the platform itself works with any LLM provider.

Methodology

How These Numbers Were Produced.

  • 100 queries per provider, across 4 verticals: Legal, Healthcare, Financial, IT (25 per vertical)
  • Query mix: 75% simple/definitional, 17% medium/explanatory, 8% complex/open-ended — designed to approximate real support/copilot traffic
  • Run date: July 12, 2026
  • Measured: routing accuracy — did the query go to the tier (small/medium/large) a human pre-labeled as correct for that query's complexity?
  • Savings computed from real per-query dollar costs against what the same queries would have cost at that provider's most expensive tier

Results

9 Providers — Routing Accuracy & Savings.

ProviderRouting AccuracyErrorsVerified Savings vs. Always-Large
DeepSeek100%0/10088.3%
Qwen100%0/10086.3%
OpenAI100%0/10084.9%
Azure OpenAI100%0/10084.9%
Google100%0/10084.4%
Anthropic100%0/10078.8%
Mistral100%0/10077.7%
Cohere98.0%2/10073.9%
Grok100%0/10030.9%

DeepSeek (88.3% — highest)

DeepSeek's small and medium tiers call the same model, which is dramatically cheaper than its large tier — 92% of traffic in this benchmark never touches the expensive model at all.

Grok (30.9% — lowest)

All three of Grok's tiers currently call the same underlying model, differentiated only by a compute-effort setting rather than a genuinely cheaper model. Every other provider saves money by routing easy queries to a smaller model; Grok doesn't offer that today. This is an honest ceiling given how Grok is currently offered, not a routing weakness.

Cohere (73.9% savings, 98.0% accuracy)

Cohere prices its medium and large tiers identically, capping its best-case savings regardless of routing quality. Its 98.0% accuracy reflects two provider-side read timeouts, not routing misclassification — the router selected the correct tier both times; the provider's API failed to respond in time.

Why the savings figures vary so much:

Savings percentage isn't a router-quality score — it's a function of how far apart each provider prices its own tiers. Routing accuracy is the fair way to judge the router itself; on that measure, 8 of 9 providers hit 100%.

Latency

Routing Overhead & End-to-End Latency.

Routing Decision (pre-inference)

This is what CARDIAC-PURR AI Control Plane itself adds, measured in isolation: sub-millisecond (p50 0.42ms, p99 0.97ms). Pure in-process computation, no I/O.

Per-Provider Proxy Overhead

ProviderOverhead
DeepSeek12.7ms
Azure OpenAI12.8ms
Grok13.3ms
Qwen13.8ms
Google14.4ms
OpenAI14.8ms
Anthropic14.9ms
Mistral15.2ms
Cohere558.9ms

Cohere's overhead is confirmed elevated across separate test runs — a genuine characteristic of that integration, not a one-off measurement artifact.

End-to-End Latency (routing + provider response)

This reflects each provider's own response time, not CARDIAC-PURR's routing overhead.

ProviderP50P95P99
Google0.77s5.38s6.22s
Mistral0.77s4.00s5.90s
OpenAI0.81s2.86s3.59s
Qwen0.86s4.20s5.03s
Grok1.07s6.97s9.34s
Azure OpenAI1.16s2.40s3.30s
Anthropic1.81s25.54s48.36s
DeepSeek1.96s9.04s21.75s
Cohere3.84s11.25s16.27s

Agentic Framework

Agentic Framework Benchmark — August 12, 2026.

A separate, smaller-scope benchmark of agent-framework routing — 350 factual-QA requests across 7 frameworks, all routed to the small tier. This is distinct from the 9-provider AI Control Plane benchmark above.

100%

Accuracy on this 350-request factual-QA workload, small tier

80%

Savings on this same factual-QA workload

Frameworks tested: LangChain, LlamaIndex, AutoGen, CrewAI, Haystack, DSPy, Semantic Kernel — 50 requests each, 350 total, Anthropic claude-haiku-4-5-20251001, small tier.

These figures apply only to this factual-QA workload at small-tier routing — not a general accuracy or savings claim, and medium/large tiers were not exercised by this run. Caching affected the workload and is disclosed here rather than omitted. Results are reported as an aggregate across all 7 frameworks combined; no per-framework breakdown is published from this run.

Full architecture, framework integration paths, and routing detail are documented separately.

Explore Agentic Framework →

Scope

What This Measures, and What It Doesn't.

This is one benchmark run against a query mix designed to approximate real enterprise traffic — a reasonable pilot, not a claim of production-scale robustness across every workload. If you're evaluating this for your own traffic, run your own workload through it rather than take this number as a universal guarantee.

Full per-query results are available on request.

FAQ

Common Questions.

Why does Cohere's latency look so different from the other providers?

Cohere's routing overhead (558.9ms) is meaningfully higher than the other 8 providers (12.7–15.2ms), confirmed across separate test runs — a genuine characteristic of that integration, not a measurement error.

Why was Azure tested but folded into the OpenAI row for cost data?

Azure OpenAI runs the same underlying models as OpenAI.

Does this mean CARDIAC-PURR AI Control Plane only works with these 9 providers?

No. These 9 were used for this specific benchmark run. The platform itself works with any LLM provider — see Platform for how routing and governance work regardless of which provider you use.

Can I see the full per-query data?

Yes, available on request — full methodology and per-item results are provided under standard technical documentation request.

Why does Grok show such a lower savings figure (30.9%) than the rest?

The 30.9% figure is reported deliberately, alongside the higher-performing providers, because this benchmark is presented as honest results including where the platform underperformed — not a curated best-case number. It reflects that specific provider's characteristics on this specific 100-query, 4-vertical test run, not a universal ceiling.

What counts as "Accuracy" in the results table — routing accuracy or response correctness?

Routing accuracy, specifically: whether the query went to the tier (small/medium/large) a human pre-labelled as correct for that query's complexity — see Methodology above. It is not a measure of whether the underlying model's answer was factually correct.

Are these real enterprise queries or synthetic test data?

Real: 100 real enterprise queries per provider, run live with real spend against each provider's actual API — not a synthetic, easy-classification task built to flatter the result.

Are the savings percentages based on public list price, or negotiated/volume pricing?

Savings are computed from real per-query dollar costs against what the same queries would have cost at that provider's own most expensive tier — i.e., standard published pricing, not a customer-specific negotiated or volume-discounted rate. If you have your own negotiated provider pricing, your actual savings may differ from the published figures.

What's the confidence interval on these numbers, and how is it calculated?

95% confidence intervals using the Wilson score method (standard statistical practice for proportions), documented in enterprise_benchmark.py. The benchmark is re-run multiple times per month; the results above are from the July 12, 2026 run.

Which providers were excluded from this run, and why?

Not excluded so much as out of scope for this specific run: this benchmark covered 9 providers (DeepSeek, Qwen, OpenAI, Azure OpenAI, Google, Anthropic, Mistral, Cohere, Grok). Other supported providers, including Llama and Custom endpoints, weren't part of this particular test — see Platform for the full supported-provider list.

Will these numbers hold for our own traffic?

Not necessarily as a guarantee — these results reflect 100 real enterprise queries per provider across Legal, Healthcare, Financial, and IT verticals. If you're evaluating this for your own workload, running your own traffic through it is the recommended way to validate savings for your specific mix of tasks.

How often is this benchmark re-run?

Multiple times per month. This run reflects the latest results, from July 12, 2026.

Start Free →