FIBRE

apodex-1.1-mini

Apodex · apodex-1.1-mini · unpinned

Ranked 4 of 5 overall on 202 tasks. Strongest on monitoring at 84.7, weakest on credit analysis at 65.2.

Score
75.5
rank 4 of 5
Latency
34.7s
median per task
Cost
$0.024
per task at list prices

Capability profile

Mean score per category on question-answering, 0–100.
CreditMonitorStructuredResearchDiscovery

Score by category

5 categories, 152 tasks on question-answering.
Credit Analysis65.2 · 62
Monitoring84.7 · 20
Structured Products80.8 · 18
Research Intelligence76.9 · 10
Discovery & Prediction81.6 · 2

Where it sits

apodex-1.1-mini against the rest of the board. Everything else is drawn back so the one model reads.
No model here.Better on both axes at once.$0.01$0.02$0.05$0.1$0.2$0.5$165758595Cost per task · better ←Score · better ↑Best per taskapodex-1.1-mini
apodex-1.1-mini beaten on both by gpt-5.6-lunaFrontier · 3 models nothing beats on both

Best per task is the diamond: for each of the 200 tasks every model attempted, the highest score any of them got, averaged — 95.0 against 91.3 for opus-5 used on everything. Nobody can build it. The choice is made with the result in hand, which a router does not have when it has to choose, so it marks the edge of what having 5 different models is worth rather than a score anyone can reach. It is also flattered by noise — the highest of 5 measurements rises with the count even when the models behind them do not improve — so it will drift up as models are added and is not comparable between boards with different numbers of them. 2 tasks are left out of it, for not having been attempted by every model. Its cost is measured the way the board measures a model’s, on the question-answering track that holds more of the tasks.

How it works when given tools

Behaviour across 150 agentic-research runs.
Delivered an answer
79%
of runs produced a deliverable
Mean steps
20.7
tool calls per run
Tools used per run
fetch_url 14.9sec_search 3.8python_exec 1.2web_search 0.8
How runs ended
final answer 62%max steps 38%

Where it lost the most

The lowest-scoring tasks on question-answering, with the rubric gates that fired where they are known.
PARTIAL
Relative Value
Credit Analysis · RV, REIT, recommendation
8.9
PARTIAL
Spread Analysis
Credit Analysis · spread, accrued interest, dirty price
10.5
PARTIAL
New Issue Evaluation
Credit Analysis · new issue, high yield, fair value
penalty: Does not apply IG concession norms (10-15bp) to a BB- high-yield deal. HY and IG new issu…
17.9
PARTIAL
Spread Analysis
Credit Analysis · spread, discount margin, floating rate
penalty: Does not report the quoted margin of 125bp as the bond's current spread. The quoted margi…
19.4
PARTIAL
Curve Trade Analysis
Credit Analysis · rates, curve trade, steepener
26.6