Model routing gateway

Route every call to the cheapest model that holds quality.

Fairmeter drops in front of chatbots, RAG pipelines, agent loops, and finance workloads. Cut spend 40 to 80 percent, measured on our own traffic, live in thirty seconds.

Plans start at $9.99/month.

00%

Spend cut

~0 ms

Cache-hit latency

Live resolverrouting
app

gpt-4o-mini

cheap · fast

$0.01

gpt-4o

balanced

$0.06

claude-opus-4-8

frontier reasoning

$0.41
complexity 2.3 → routed to gpt-4o-miniquality held

Drop in front of the OpenAI or Anthropic SDK you already use

Any OpenAI- or Anthropic-shaped endpoint drops in behind one base URL. Kimi and GLM are next on the roadmap.

OpenAIAnthropicGoogle
The problem

Every call hits your top model. The bill keeps climbing and nobody can say why.

Pin one frontier model on classification, retrieval, tool calls, and cleanup alike and you pay reasoning rates for work a cheap model would nail. Then the invoice lands as one line, shipping outruns anything finance can categorise, the cost report is a guess, and the audit trail is a Slack thread.

Which team spent $14k last week?

team_id=engineering · 9,212 calls · 89% on opus

Why was this call so cheap?

complexity 2.3 · routed to gpt-4o-mini · quality held

Did we pay for this twice?

cache hit · same question, different words · $0.00

One worked example, before and after

The same five-step LangChain loop, routed two ways.

The Before trace pins every step on the same top-tier reasoning model, which totals $1.84. The After trace lets the resolver pick the cheapest variant that holds the quality bar on each step, which totals $0.77 — 58 percent less for the same five steps. Every per-step price is on the page, so the arithmetic is yours to check.

Before · vanilla LangChain

one provider, one budget — every step on claude-opus-4-8 at thinking=high

Planning$0.42

claude-opus-4-8 · thinking=high

Retrieval$0.38

claude-opus-4-8 · thinking=high

Tool call$0.28

claude-opus-4-8 · thinking=high

Synthesis$0.41

claude-opus-4-8 · thinking=high

Verification$0.35

claude-opus-4-8 · thinking=high

Loop total$1.84

After · Fairmeter

per-step resolver, session envelope — cheapest variant that holds quality

58% less
Planning$0.25

cap:reason-heavy · gpt-5.6-terra high

Retrieval$0.06

cap:long-context-128k · gpt-4o

Tool call$0.04

cap:tool-call-strict · gpt-4o no thinking

Synthesis$0.41

cap:reason-heavy · claude-opus-4-8 high

Verification$0.01

cap:cheap-fast · gpt-4o-mini

Loop total$0.77
The fix

One gateway. Six superpowers.

Each one shows up on the dashboard the moment your first request lands.

State

Some tools only allow one model family. Point Claude Code, Cursor, or your own app at Fairmeter and it stays inside the family they require, routing to the right tier for each step and never leaking to another vendor. It holds the conversation too, so switching models mid-thread stays safe.

Read the docs

Cache

When a request matches one Fairmeter has already handled, the answer comes straight back and you are not charged for it again. Fairmeter catches exact repeats and the ones that just mean the same thing, and every team's cache stays walled off from the rest.

Read the docs

Route

Routing well is the whole product. A purpose-built classifier reads each request for true difficulty, weighs it against live cost and latency across providers, and picks the model that clears your quality bar for the least spend, with instant fallback if one slips.

Read the docs

Book

Per-team budgets, per-key limits, and a four-way input, output, cached, and thinking cost split, written to the ledger and enforced before the provider call.

Read the docs

Session

A session is one accountable unit. Every multi-step agent loop tracks per-turn cost, latency, tokens, and step class, and a governor caps how much thinking the session can spend.

Read the docs

Step

An agent runs many steps, and most of them don't need your most expensive model. Fairmeter sizes up each one and routes it to the cheapest model that can still do it well, so you get frontier quality where it counts and pay a fraction everywhere else.

Read the docs
Try it yourself

What the dashboard looks like

Spend, budgets, and per-team usage on one page. Illustrative dashboard, sample data.

Operations

Live request feed, agent sessions, spend, and per-step routing for your org.

live

$184.31

Spend today

72.6%

Cache hit rate (24h)

49

Agent sessions (24h)

1,564 ms

P99 latency (24h)

Live feed

req_gr7oclaude-haiku-4-51576 tok79ms
req_s8kuopenai/o42492 tok9256ms
req_hkcjgpt-4o-mini3544 tok5999ms
req_k3grgpt-4o-mini568 tok312ms
req_gr7oclaude-haiku-4-51576 tok79ms
req_s8kuopenai/o42492 tok9256ms
req_hkcjgpt-4o-mini3544 tok5999ms
req_k3grgpt-4o-mini568 tok312ms

Cost by model (7d)

GPT-4o46%
Opus34%
Haiku20%
Pricing

Priced so it pays for itself.

Get started in thirty seconds for $9.99/month. Scale into Team and Enterprise when the bill makes the case.

Basic

$9.99per month

Billed monthly, cancel anytime.

One URL for every provider. Bring your own provider keys (BYOK).

Usage limit — 240 requests/min

  • OpenAI, Anthropic, and Google, with one URL for all three
  • Bring your own keys (BYOK)
  • Semantic cache
  • Email support
Start with Basic
Most popular

Team

$70per month

Billed monthly, no annual plan.

Everything in Basic plus alerting, org-scoped access, and custom band maps.

Usage limit — 600 requests/min

  • Everything in Basic
  • Per-team budgets
  • Audit log
  • Per-team alerting
  • Org-scoped access controls
  • Custom band maps (auto, cheap, frontier)
  • Priority email support
Upgrade to Team

Enterprise

CustomAnnual

Custom annual contract.

On-prem deployment, SOC2 evidence pack on request, dedicated support.

Usage limit — 6,000+ requests/min

  • Everything in Team
  • On-prem or VPC deployment
  • SOC2 evidence pack on request
  • ERP integrations (on the roadmap)
  • Dedicated support engineer
Contact sales

Three lines today. A clean bill tomorrow.

Route real traffic in under five minutes.

Plans start at $9.99/month.