FIBRE
Fixed Income Benchmark for Rigorous Evaluation

How well do agents handle fixed-income tasks?

FIBRE scores frontier models on 202 practitioner-written fixed-income tasks across two tracks — answering from a written brief, and researching filings live with tools — and reports what each answer costs in seconds and dollars.

Tasks
202
152 question-answering · 50 agentic-research
Models evaluated
5
3 labs · 31 task types
Judge panel
3
3 model families, majority vote
Last published
Sep 25, 2026

Leaderboard

#
01
opus-5
Anthropic · high
Score
91.3
Question-answering
90.0
Agentic-research
95.0
Latency143s
Cost$1.16
02
sonnet-5
Anthropic · high
Score
86.3
Question-answering
85.8
Agentic-research
88.0
Latency86.1s
Cost$0.519
03
Score
80.4
Question-answering
81.3
Agentic-research
77.8
Latency24.5s
Cost$0.013
04
Score
75.5
Question-answering
74.9
Agentic-research
77.3
Latency34.7s
Cost$0.024
05
haiku-4.5
Anthropic
Score
65.2
Question-answering
68.7
Agentic-research
54.4
Latency41.8s
Cost$0.085

Score weights every task equally regardless of track, so a model scored on more of them carries more of its own result. The # column reports standing on whatever the board is ranked on — that score, or the category chosen above it — and keeps reporting it whichever column you sort by. Question-answering and Agentic-research break that score down by track and are scored against the same rubric per task. Latency is the median wall clock per task and cost is the mean at list API prices from the v1-2026-08 table; a model with no published price shows no cost rather than a guess.

Ranked by lab, the board keeps one row per lab: that lab’s best model on Score, with every figure still that one model’s. Best rather than an average, which would only penalise a lab for also entering its small and cheap models.

Trade-offs

No model here.Better on both axes at once.$0.01$0.02$0.05$0.1$0.2$0.5$165758595Cost per task · better ←Score · better ↑Best per taskgpt-5.6-lunasonnet-5opus-5
Frontier · 3 models nothing beats on bothBetter on both — nothing built yetBeaten outright by a model on the lineBest per task — a ceiling, not a model: the score if every task had gone to whichever of the 5 turned out to win it

Best per task is the diamond: for each of the 200 tasks every model attempted, the highest score any of them got, averaged — 95.0 against 91.3 for opus-5 used on everything. Nobody can build it. The choice is made with the result in hand, which a router does not have when it has to choose, so it marks the edge of what having 5 different models is worth rather than a score anyone can reach. It is also flattered by noise — the highest of 5 measurements rises with the count even when the models behind them do not improve — so it will drift up as models are added and is not comparable between boards with different numbers of them. 2 tasks are left out of it, for not having been attempted by every model. Its cost is measured the way the board measures a model’s, on the question-answering track that holds more of the tasks.