Real-time frontier model health

Know which frontier model is actually good right now.

Frontier models drift. The same model, same settings, can be brilliant one minute and slow or shallow the next — demand, maintenance and changes you never see. FrontierScore continuously tests them and scores live quality, speed and reasoning, so you and your routers always pick the best.

View live data →
live.frontierscore.ai is now public for everyone — full dashboard free, data delayed by 1 hour. Sign in (free) for realtime.
Preview the dashboard
Become R&D member for a free Premium subscription
Tracking the latest Claude · GPT · Gemini · Grok · DeepSeek versions — and growing
Frontier model health
Delayed 1h
The problem

Frontier performance isn't constant.

You're paying frontier prices for a moving target. Without live visibility, you can't tell a great run from a degraded one until the output disappoints.

Quality swings

The same prompt and settings can return deep, holistic reasoning one hour and shallow, rushed answers the next — with no warning and no changelog.

Speed swings

Latency and throughput move with demand, maintenance and silent infrastructure changes. "Fast" is not a fixed property of a model.

You're flying blind

Benchmarks are run once and published. They tell you how a model did last month — not which model is the right call for the request you're about to send.

How it works

Continuous frontier tests.

Deep analytical tasks, run around the clock against every provider — turned into live, comparable scores.

Probe

Run curated deep-analysis "frontier tests" across providers, continuously — not once a month.

Measure

Capture answer quality, reasoning depth and holistic approach, plus latency and throughput.

Score

Normalize into live quality, speed and an overall preferred score per model and setting.

Serve

Publish to the live dashboard and a high-performance API — ready for routers and agents.

Capabilities

One score for the whole frontier.

Live health dashboard

See current quality, speed and overall score for every tracked model, refreshed continuously.

High-performance API

Query the current score for any model in milliseconds — built to sit in the hot path of a request.

Intelligent routing

Feed live scores to smart routers and MCP so they route to whichever model is best right now.

Quality · speed · reasoning

Three dimensions, not one number — so you weight what matters for each workload.

Multi-provider coverage

Claude, GPT, Gemini, Grok, DeepSeek and open-weight models, scored on the same scale.

Built into Gateward

Drop FrontierScore into gateward.ai to power its routing decisions out of the box.

Coverage

Every provider, every new version.

We track the latest models across providers — and add new versions the moment they ship, so the score always reflects today's frontier.

Anthropic
Claude Opus 4.8Sonnet 4.6Fable 5
OpenAI
GPT-5.5GPT-5.5 ProCodex
Google
Gemini 3.1 ProGemini 3.5 Flash
xAI
Grok 4.3reasoning levels
DeepSeek
V4 ProV4 Flash
Open-weight
Llama 4Qwen 3.7Kimi K2.6
And growing
New models & versions added as they ship
Live model health

The frontier, scored in real time.

These are real scores, delayed 1 hour and shown hourly. Live, realtime data is unlocked in the beta.

Healthy Elevated latency Degraded Real measurements · delayed 1 hour · join the beta to unlock realtime
Public scores are delayed 1 hour — join the beta to unlock realtime data.View live data →Create a free account →
The API

A score your router can act on.

One fast call returns the current quality, speed and overall score for any model — turning a plain router or MCP server into an intelligent one.

  • Millisecond responses, built for the request hot path
  • Per-model and per-setting scores (e.g. reasoning effort)
  • Status flags so you can fail over from a degraded model
  • Drop-in for smart routers, MCP and Gateward
# current score for a model
GET /v1/score?model=claude-opus-4-8

{
  "model": "claude-opus-4-8",
  "quality": 96.2,
  "speed_tps": 142,
  "reasoning": 94.8,
  "score": 95.1,
  "status": "healthy",
  "updated": "2026-06-24T14:05Z"
}
Pricing

Start free. Scale when you're ready.

The full dashboard is free — signed in for real-time, or open with a 1-hour delay. Paid tiers add the API, SDK, MCP and higher limits.

 
Free
$0 / forever
Explore the whole frontier — public scores, one hour delayed.
  • Full dashboard, 1-hour delayed
  • Every model & region
  • Up to 3 starred models
  • No API access
  • Community support
Open the dashboard
Most popular
Premium
Coming soon
Real-time scores in the hot path — full API, SDK, MCP and generous quotas.
★ Starts with a free 90-day trial — full access, no card
  • Real-time scores (no delay)
  • Full API + SDK + Gateward
  • MCP access
  • Unlimited starred models
  • 30-day history & alerts
  • Email support
Start free 90-day trial
 
Enterprise
Let's talk
Run it fully private on-prem, test your own models, with SSO and SLA-backed support.
  • Everything in Premium
  • Private on-prem worker — fully secure, air-gapped (keys stay in your network)
  • Test your own local / private models
  • SSO / SAML · custom quotas · unlimited history
  • SLA & dedicated support
  • EU-sovereign deployment
Contact sales

Become an R&D member for a free Premium subscription.

Connect · intelligent routing

Three ways to put the score in the loop.

From a one-line change to full control — consume the routing intelligence however your stack prefers, with built-in failover on every path.

1 · Through Gateward

Point your app at Gateward's OpenAI-compatible endpoint — change nothing else. It routes every call to the best model and keeps your keys vaulted. x-route-task: reasoning

2 · Via MCP

Add the FrontierScore MCP server; your agent calls best_model() before a step and routes accordingly — ideal when your orchestrator already speaks MCP.

3 · Direct routing API

Call /v1/route?task=reasoning from your own router and act on the recommendation yourself — maximum control when you already have a routing layer.

AI-native · new

Or let your AI agent set it up.

Connect Claude, Cursor or any MCP-capable assistant to Digital One and it can onboard you to FrontierScore, issue your API keys, and even join you to the contributor fleet — on your behalf, with your approval, and only what you allow.

Onboard & get keys

Your agent enables FrontierScore and issues the API key you'll route with — without you ever leaving the chat.

Join the fleet

“Join me to the FrontierScore fleet” returns a one-command worker installer, ready to run — your path to a free Premium subscription.

You stay in control

Approve once in your browser via a standard OAuth 2.1 device flow — scoped to exactly what you allow, every action audited, and revocable anytime.

Connect via MCP · mcp.frontierscore.ai · works with Claude, Cursor & more Connect your agent →
Pairs with Gateward

The intelligence behind the gateway.

Gateward routes and governs your models; FrontierScore tells it which model is best moment to moment. Run them together for a local-first gateway that always reaches for the strongest available frontier model.

Visit gateward.ai →
Contributor · R&D support program

Run a worker. Earn a free Premium subscription.

Help measure the AI frontier from your region. Run one lightweight worker that probes the major models with your own API keys — you share only the measurements, never your keys, prompts or data. While it qualifies, your whole organization is on the Premium tier, free.

You contribute

Regional measurements from a small worker that probes the major models using your own keys — latency, success rates and scoring inputs, signed and sent back. Stats, not data.

You receive

The Premium tier across the suite — the high-performance API and full dashboard (the complete live model-health view, history and routing signals) — free while you qualify.

Your keys stay yours

Your API keys live only inside your worker, in your environment. It uses them locally to call the models; only the measurements it produces are signed and sent to us. We never receive, store or use your keys, and we never see your prompts, traffic or business data.

To keep the tier: each month your worker tests at least 4 models from the major providers and stays online ≥ 80%. Drop below either and it simply converts back to free — no charge, re-qualify anytime.

View live data & R&D program.

We're onboarding early testers and research partners to shape the frontier tests, the scoring and the API. If you route, build agents, or just want to stop guessing — let's talk.