FIBRE
Methodology

How FIBRE is built and scored

FIBRE measures whether a model can do fixed income work end to end: read the document, get the arithmetic right, cite the source, and then carry a multi-step task through tools without losing the thread.

Two tracks

The question-answering track hands the model a bounded brief in its context window — a filing extract, a term sheet, a rate snapshot — and asks for one written answer. There are no tools. Everything the model needs is in front of it, and everything it claims has to be supported by what it was given.

The agentic-research track gives the model a question and a live toolchain instead: web search, filing search, page retrieval, and a Python sandbox. It has to find its own evidence. A run that burns its step budget without producing a deliverable scores zero rather than being dropped, because failing to deliver is a result about the model and not a failed API call.

Question-answering · 152 tasks
Answered from a written brief with no tools. Graded on the same rubric as the task it shares with the research track.
Agentic-research · 50 tasks
Answered by researching live sources under a step budget. Trajectory, tool mix and termination reason are recorded alongside the score.

Scoring

Every task carries a rubric written by a practitioner: a list of specific claims the answer has to make, each worth points, each tagged with the stage of work it belongs to. A panel of judge models votes on each criterion independently, and the majority carries. An item’s score is what the answer earned against what it could have earned.

S = (T + B − P) / C

T is the core points earned, B the bonus points, P the penalties, and C the core points available. Bonus criteria sit outside the denominator, so a model can exceed 1.0 on a task by doing more than the rubric demanded. Must-pass criteria are gates rather than points: missing one zeroes the item however well the rest went. A task counts as correct at 80 or above.

Latency and cost are reported next to the score rather than folded into it. A model that trades three seconds for four points is a different purchase depending on whether it sits in a trading loop or an overnight batch, so FIBRE does not make that trade on the reader’s behalf. Prices come from the v1-2026-08 table; a model with no published price shows no cost.

The judges

Question-answering · 3 judges
judge-opus-5, judge-gpt-5.6-sol, judge-deepseek-v4-pro
Majority vote across 3 model families.
Agentic-research · 3 judges
judge-opus-5, judge-gpt-5.6-sol, judge-deepseek-v4-pro
Majority vote across 3 model families.

The panel shown is read off the verdicts that actually graded each track, not from the configuration file. The two tracks were graded under different panels and the site says so rather than quoting the panel that was supposed to run.

Run protocol

01Responses are the primary record. Grading can be re-run against a revised rubric without re-running the model, and it frequently is.
02Scores are read from the record written when the run happened, so an item regraded after its rubric changed does not silently rescore against criteria it never saw.
03Runs are consolidated in chronological order and a later result for the same task, model and track supersedes an earlier one.
04Identical tool schemas for every model on the research track. No model-specific prompt tuning beyond each vendor's system-prompt format.
05Publishing is a deliberate step. 6 runs contributed to this snapshot; a run that was entirely superseded is not listed among them.

A worked example

One task from the public split, with the graded output of the best and worst model on it side by side. The judge saw the same material the model did. This is read live from the run records rather than from the published snapshot, so it is the actual answer that produced the actual score.

Question-answeringCredit Analysis

Spread Analysis

Tagged
spreaddiscount marginfloating rate
Bond E is a SOFR-linked floating-rate note priced at 99.50. Analyze its spread characteristics: determine the appropriate market-implied spread measure, explain what the price discount reveals about the market's current credit assessment versus issuance conditions, assess how sensitive your spread estimate is to the assumed path of SOFR, and compare Bond E's effective spread to the Z-spreads on the fixed-rate bonds (D and F) for relative value purposes — accounting for the fundamentally different rate risk profiles of floating vs fixed-rate bonds.

5 rubric criteria, worth 14 points in total. A criterion is earned when the majority of the panel votes for it. An item counts as correct at 80 or above.

gpt-5.6-luna
PASS
54.6s
## Bond E — Vanguard Telecom SOFR + 125 bp FRN ### 1. Appropriate spread measure For a floating-rate note, the most appropriate market-implied spread measure is the **discount margin (DM)**, sometimes called the **floating-rate note spread**. It is preferable to a conventional fixed-rate Z-spread because: - Bond E’s coupons reset with SOFR. - Its primary rate exposure is to the future path of SOFR, not to the Tre…
Compares Bond E's effective spread (DM ~155bp) to the fixed-rate bonds' Z-spreads, while…earned
Autonomously identifies discount margin (DM) as the appropriate spread measure for a floa…earned
Computes the DM premium over the quoted margin: DM (~150-160bp) minus QM (125bp) = approx…earned
Does not report the quoted margin of 125bp as the bond's current spread. The quoted margi…earned
Analyzes DM sensitivity to SOFR: explains that for a short-dated FRN, DM is largely insen…earned
Item score100.0
Panel — 5 of 5 criteria earned, each verified against a quoted span of the answer.
haiku-4.5
FAIL
152s
# Comprehensive Spread Analysis: Vanguard Telecom Bond E (SOFR FRN) ## EXECUTIVE SUMMARY Bond E trades at **99.50 (discount)** despite offering current coupon of **5.55%** (SOFR + 125bp margin). This analysis reveals: - **OAS-equivalent spread:** ~127bp (accounting for the reset frequency and rate-path dependency) - **Discount drivers:** Market credit repricing (tightening from issuance), not margin inadequacy - *…
Compares Bond E's effective spread (DM ~155bp) to the fixed-rate bonds' Z-spreads, while…missed
Autonomously identifies discount margin (DM) as the appropriate spread measure for a floa…missed
Computes the DM premium over the quoted margin: DM (~150-160bp) minus QM (125bp) = approx…missed
Does not report the quoted margin of 125bp as the bond's current spread. The quoted margi…missed
Analyzes DM sensitivity to SOFR: explains that for a short-dated FRN, DM is largely insen…missed
Item score-11.1
Panel — 0 of 5 criteria earned, each verified against a quoted span of the answer.
Points by stage
synthesis9
calculation3
convention2

What this snapshot rests on

This snapshot was published from 6 runs holding 2,993 scored task-and-model pairs and 8,973 judge verdicts.

Every number on this site comes from that one document. Nothing is recomputed in the browser, and publishing a new snapshot is what changes them.

← Back to leaderboard