Independent AI intelligence01 / LIVE INDEX

THE MODEL IS NOT THE MARKETING.

Deep benchmarks. Real tool use. Hidden tests. Adversarial audits. Cost, tokens, latency and reliability — measured under one reproducible harness.

Explore
rankings
MODEL / PROVIDER / PRODUCTv0.3 · CRUCIBLE 2026SCROLL ↓
CAPABILITY RELIABILITY COST / QUALITY TOKEN EFFICIENCY LATENCY AUTONOMY CAPABILITY RELIABILITY COST / QUALITY TOKEN EFFICIENCY LATENCY AUTONOMY

02 / LEADERBOARD

One score is never
the whole story.

Switch tracks instead of mixing apples, agents and gateways. Pure model runs use the same API harness. Product runs measure the full agent stack.

MODEL OF THE MONTH

CURRENT EVIDENCE

Awaiting pure-model cohort

The first public Agent Stack results are seeded; Model of the Month activates when comparable Pure Model runs exist.

QUALITY
VALUE
CONFIDENCEPENDING
#Model / stackTrackQualityReliabilityTimeQ / hourStatus

03 / CAPABILITY SCALE

70 is not a C.
70 is frontier.

Crucible is deliberately calibrated to leave headroom. The bands describe what a score means inside this benchmark, while category gates decide whether a model earns labels such as complex-app builder, security specialist or research specialist.

Bands are provisional until the 2026 anchor cohort is complete. They are capability evidence, never deployment or safety certification.

BENCHMARKBenchAtlas Crucible 2026

24 tasks · 12 domains · frozen season

INTENDED FRONTIER LINE60+

Broad high-end performance. Frontier+ begins at 70.

CEILING TARGET< 80

If every frontier model scores 90+, the benchmark failed to be hard enough.

CATEGORY MAP

Strong models should fail differently.

Overall score shows breadth. Domain scores show what the model is actually good at. Category badges require both score thresholds and adversarial reliability.

04 / VALUE LAB

Intelligence
per dollar.

A model that is 1% better and 5× slower is a different product decision. BenchAtlas keeps capability and efficiency separate — then lets you combine them transparently.

QUALITY ↑
COST →
QWEN 3.8
85.63
OPUS 5
86.47

Illustrative Agent Stack position. API-cost scatter activates when Pure Model runs contain normalized cost data.

04.9×

more observed score per hour

In the first audited Agent Stack duel, Qwen reached nearly the same quality in a fraction of the observed time. BenchAtlas records that distinction instead of flattening it into a single marketing number.

Qwen reviewed70.9 Q/h
Opus one-shot14.4 Q/h

04 / PRICE INTELLIGENCE

What does
$20 actually buy?

API prices, cache rates and subscriptions live in different billing universes. We never pretend opaque credits equal a fixed number of tokens. Use official token prices where they exist; use measured runs-per-plan where they do not.

ESTIMATED API COST$0.00

Select a priced model.

Source dated

Subscriptions ≠ token buckets.

For plans with credits or dynamic limits, the honest metric is empirical: how many benchmark-shaped workloads fit into the plan?

06 / PROVIDER UNIVERSE

One runner.
Many endpoints.

Choose a preset or bring any OpenAI-compatible URL. BenchAtlas discovers /models when possible, while native adapters cover protocols that need different message/tool semantics.

benchatlas — interactive runner
$ benchatlas run

Proveedor      OpenRouter
Base URL       https://openrouter.ai/api/v1
API key        ••••••••••••••••
Descubriendo   400+ models →
Modelo         anthropic/claude-opus-5
Evaluación     strict

 sandbox isolated · network off
 fresh conversation per task
 hidden tests remain invisible
 cost / tokens / latency / tools recorded

07 / METHODOLOGY

Green tests are
the beginning.

The benchmark that produced our first public results taught us something useful: both candidates passed large test suites and still contained real bugs. So the evaluator attacks the solution after it finishes.

01

Visible task

Exact frozen prompt, same tool vocabulary, isolated workspace. No judge files. No internet unless a pack explicitly allows it.

02

Hidden grader

Deterministic tests and reference checks run only after the candidate is done. The candidate never mounts the evaluator tree.

03

Blind rubric

Manual-design tasks are scored from the published rubric in a fresh evaluator context with candidate identity hidden.

04

Adversarial audit

A disposable copy is attacked for authorization, crash recovery, idempotency, migration, concurrency and false-confidence tests.

TRACK APURE MODEL

Same harness + API tool protocol. Best answer to “which model is strongest under equal conditions?”

TRACK BPROVIDER

Same model through different inference providers. Measures speed, price, routing and behavioral drift.

TRACK CAGENT STACK

Model + Kiro/OpenCode/Codex/etc. Best answer to “which product actually gets the work done?”

NO SINGLE SCORE
WITHOUT EVIDENCE.

PromptsHashesRaw runsCostsJudge modeFailure evidence