CURRENT EVIDENCE
Awaiting pure-model cohort
The first public Agent Stack results are seeded; Model of the Month activates when comparable Pure Model runs exist.
Independent AI intelligence01 / LIVE INDEX
Deep benchmarks. Real tool use. Hidden tests. Adversarial audits. Cost, tokens, latency and reliability — measured under one reproducible harness.
Explore02 / LEADERBOARD
Switch tracks instead of mixing apples, agents and gateways. Pure model runs use the same API harness. Product runs measure the full agent stack.
CURRENT EVIDENCE
The first public Agent Stack results are seeded; Model of the Month activates when comparable Pure Model runs exist.
| # | Model / stack | Track | Quality | Reliability | Time | Q / hour | Status |
|---|
03 / CAPABILITY SCALE
Crucible is deliberately calibrated to leave headroom. The bands describe what a score means inside this benchmark, while category gates decide whether a model earns labels such as complex-app builder, security specialist or research specialist.
Bands are provisional until the 2026 anchor cohort is complete. They are capability evidence, never deployment or safety certification.
24 tasks · 12 domains · frozen season
Broad high-end performance. Frontier+ begins at 70.
If every frontier model scores 90+, the benchmark failed to be hard enough.
CATEGORY MAP
Overall score shows breadth. Domain scores show what the model is actually good at. Category badges require both score thresholds and adversarial reliability.
04 / VALUE LAB
A model that is 1% better and 5× slower is a different product decision. BenchAtlas keeps capability and efficiency separate — then lets you combine them transparently.
Illustrative Agent Stack position. API-cost scatter activates when Pure Model runs contain normalized cost data.
04.9×
In the first audited Agent Stack duel, Qwen reached nearly the same quality in a fraction of the observed time. BenchAtlas records that distinction instead of flattening it into a single marketing number.
04 / PRICE INTELLIGENCE
API prices, cache rates and subscriptions live in different billing universes. We never pretend opaque credits equal a fixed number of tokens. Use official token prices where they exist; use measured runs-per-plan where they do not.
Select a priced model.
Source datedFor plans with credits or dynamic limits, the honest metric is empirical: how many benchmark-shaped workloads fit into the plan?
06 / PROVIDER UNIVERSE
Choose a preset or bring any OpenAI-compatible URL. BenchAtlas discovers /models when possible, while native adapters cover protocols that need different message/tool semantics.
$ benchatlas run
Proveedor OpenRouter
Base URL https://openrouter.ai/api/v1
API key ••••••••••••••••
Descubriendo 400+ models →
Modelo anthropic/claude-opus-5
Evaluación strict
✓ sandbox isolated · network off
✓ fresh conversation per task
✓ hidden tests remain invisible
✓ cost / tokens / latency / tools recorded
07 / METHODOLOGY
The benchmark that produced our first public results taught us something useful: both candidates passed large test suites and still contained real bugs. So the evaluator attacks the solution after it finishes.
Exact frozen prompt, same tool vocabulary, isolated workspace. No judge files. No internet unless a pack explicitly allows it.
Deterministic tests and reference checks run only after the candidate is done. The candidate never mounts the evaluator tree.
Manual-design tasks are scored from the published rubric in a fresh evaluator context with candidate identity hidden.
A disposable copy is attacked for authorization, crash recovery, idempotency, migration, concurrency and false-confidence tests.
Same harness + API tool protocol. Best answer to “which model is strongest under equal conditions?”
Same model through different inference providers. Measures speed, price, routing and behavioral drift.
Model + Kiro/OpenCode/Codex/etc. Best answer to “which product actually gets the work done?”
NO SINGLE SCORE
WITHOUT EVIDENCE.