Ranked 5 of 5 overall on 202 tasks. Strongest on monitoring at 84.8, weakest on credit analysis at 58.5.
Score
65.2
rank 5 of 5
Latency
41.8s
median per task
Cost
$0.085
per task at list prices
Capability profile
Mean score per category on question-answering, 0–100.
Score by category
5 categories, 152 tasks on question-answering.
Credit Analysis58.5 · 62
Monitoring84.8 · 20
Structured Products69.4 · 18
Research Intelligence74.8 · 10
Discovery & Prediction79.7 · 2
Where it sits
haiku-4.5 against the rest of the board. Everything else is drawn back so the one model reads.
haiku-4.5 beaten on both by gpt-5.6-luna and 1 otherFrontier · 3 models nothing beats on both
Best per task is the diamond: for each of the 200 tasks every model attempted, the highest score any of them got, averaged — 95.0 against 91.3 for opus-5 used on everything. Nobody can build it. The choice is made with the result in hand, which a router does not have when it has to choose, so it marks the edge of what having 5 different models is worth rather than a score anyone can reach. It is also flattered by noise — the highest of 5 measurements rises with the count even when the models behind them do not improve — so it will drift up as models are added and is not comparable between boards with different numbers of them. 2 tasks are left out of it, for not having been attempted by every model. Its cost is measured the way the board measures a model’s, on the question-answering track that holds more of the tasks.
The lowest-scoring tasks on question-answering, with the rubric gates that fired where they are known.
FAIL
Spread Analysis
Credit Analysis · spread, discount margin, floating rate penalty: Does not report the quoted margin of 125bp as the bond's current spread. The quoted margi…