A Chartnaut benchmark

NautilusBench v1

AI models, ranked on real trading research. We give AI models 77 closed research tasks in Chartnaut, then check every number they report against an answer we can prove. Here is how each model performed.

First published runUpdated 29 Sep 202677 tasks12 models

Best solve rate
96%
Opus 5.5 · Anthropic
Best value
$0.008
per solve · GPT-6 Luna, 65%
Hardest area
Performance
Performance and scale, average 69% across entries
Most common cause
Wrong logic
48% of all failures

Leaderboard

A model solves a task when its saved script, re-run on hidden windows, matches the certified answer and its report matches its own run. The range under each score is its 95% interval.

#ModelRun
1tied
Opus 5.5Anthropic
96%91 to 100
94.382.3
$0.45$0.43/task · $33.18 run
494k3.3 min27 Sep 2026
1tied
GPT-6 SolOpenAI
96%91 to 100
95.084.0
$0.15$0.15/task · $11.41 run
318k3.9 min27 to 28 Sep 2026
1tied
GPT-6 AstraOpenAI
96%91 to 100
95.691.1
$0.53$0.51/task · $39.29 run
235k2.2 min27 to 29 Sep 2026
2tied
GPT-5.6 SolOpenAI
95%90 to 99
94.377.1
$0.32$0.30/task · $23.19 run
300k3.9 min28 Sep 2026
2tied
GPT-5.6 TerraOpenAI
95%88 to 100
93.790.8
$0.17$0.16/task · $12.42 run
294k2.2 min28 Sep 2026
3
Sonnet 5Anthropic
90%81 to 96
86.885.1
$0.53$0.47/task · $37.69 run
753k5.5 min26 to 27 Sep 2026
4
GPT-5.6 LunaOpenAI
84%77 to 92
83.668.2
$0.021$0.018/task · $1.39 run
299k3.2 min28 Sep 2026
5
DeepSeek V4.1 FlashDeepSeek
81%71 to 88
74.277.8
$0.050$0.040/task · $3.07 run
1.20M9.4 min28 Sep 2026
6
Fable 5.1Anthropic
79%70 to 88
74.270.0
$1.63$1.29/task · $99.69 run
889k4.0 min27 Sep 2026
7
GLM 5.3 FlashZhipu AI
73%62 to 81
64.869.4
$0.12$0.088/task · $6.80 run
1.73M9.7 min28 to 29 Sep 2026
8
GPT-6 LunaOpenAI
65%53 to 75
57.251.5
$0.008$0.005/task · $0.42 run
236k2.3 min28 to 29 Sep 2026
9
Haiku 4.5Anthropic
48%38 to 58
40.345.5
$0.38$0.18/task · $13.99 run
908k2.6 min27 Sep 2026

The range under each score is a 95% bootstrap interval over tasks. Tier-weighted counts Foundation tasks once, Hard twice and Expert three times, so easy tasks cannot carry a score. Weighted also counts code efficiency on solved tasks (see How well the code runs). Tokens include input, output and cache. Cost is the model's recorded tokens priced at its lab's public API list price (standard tier, cache reads and writes at their own rates): cost per solve is the whole run's cost divided by the tasks solved, so a model pays for its failures too. Every model runs in the same harness, with the same tools and limits. 77 tasks in bench v1; an entry is scored on the tasks it was graded on.

Solve rate

Share of the 77 tasks each model solved. The line across each bar is its 95% interval over tasks.

  1. Opus 5.596%
  2. GPT-6 Sol96%
  3. GPT-6 Astra96%
  4. GPT-5.6 Sol95%
  5. GPT-5.6 Terra95%
  6. Sonnet 590%
  7. GPT-5.6 Luna84%
  8. DeepSeek V4.1 Flash81%
  9. Fable 5.179%
  10. GLM 5.3 Flash73%
  11. GPT-6 Luna65%
  12. Haiku 4.548%

Which model should I use?

Set a budget per task and see the best model within it, or set the solve rate you need and see the cheapest model that gets there. Cost is the list price of one task, failed attempts included.

Up to $0.11 per task
Best within budget
GPT-5.6 Luna
84% of tasks solved · $0.018 per task · $0.021 per solve
  1. Opus 5.596%$0.43/task
  2. GPT-6 Sol96%$0.15/task
  3. GPT-6 Astra96%$0.51/task
  4. GPT-5.6 Sol95%$0.30/task
  5. GPT-5.6 Terra95%$0.16/task
  6. Sonnet 590%$0.47/task
  7. GPT-5.6 Luna84%$0.018/taskBest fit
  8. DeepSeek V4.1 Flash81%$0.040/taskFits
  9. Fable 5.179%$1.29/task
  10. GLM 5.3 Flash73%$0.088/taskFits
  11. GPT-6 Luna65%$0.005/taskFits
  12. Haiku 4.548%$0.18/task

What a solved task costs

Up is more tasks solved, left is cheaper. The line joins the entries nothing beats on both.

30%40%50%60%70%80%90%100%$0.005$0.010$0.020$0.050$0.10$0.20$0.50$1.00$2.00cost per solved task, log scale →Opus 5.5GPT-6 SolGPT-6 AstraGPT-5.6 SolGPT-5.6 TerraSonnet 5GPT-5.6 LunaDeepSeek V4.1 FlashFable 5.1GLM 5.3 FlashGPT-6 LunaHaiku 4.5

What the top entry spends

Opus 5.5 solves 96% of tasks. It uses 494k tokens and 3.3 min per task on average, and each solved task costs $0.45.

The cheapest entry

GPT-6 Luna solves 65% at $0.008 per solved task.

Why tokens are high

A task is a working session, not one question. The model reads Chartnaut's docs, writes scripts, runs them, reads the results and fixes what is wrong. Most of the tokens are usually cached context, read again on each turn.

By tier, area and category

Share of tasks solved by the kind of work the task asks for. Numbers are percentages.

Indicators11 tasksDefinitions19 tasksResearch28 tasksDebugging6 tasksReuse2 tasksPerformance2 tasksCorrectness1 tasksHonesty8 tasks
Opus 5.5 · Anthropic1009593100100100100100
GPT-6 Sol · OpenAI10010089100100100100100
GPT-6 Astra · OpenAI1001009310010010010088
GPT-5.6 Sol · OpenAI100959310010010010088
GPT-5.6 Terra · OpenAI919593100100100100100
Sonnet 5 · Anthropic828986100100100100100
GPT-5.6 Luna · OpenAI8295821005010010063
DeepSeek V4.1 Flash · DeepSeek10095578310010010088
Fable 5.1 · Anthropic911005783501000100
GLM 5.3 Flash · Zhipu AI100744610010010010088
GPT-6 Luna · OpenAI737943675010010088
Haiku 4.5 · Anthropic7363215050100063
Average91907190831008389

How agents fail

Every failed attempt gets one cause. Most fall into five broad causes, with a small remainder we keep as Other. The causes are the same for every model. What changes is the mix. In this run, causes are assigned from which grading check the attempt failed.

Wrong logicData and time handlingDidn't flag a limitCouldn't finishMisreported the resultOther
Opus 5.5 · Anthropic
3 failed
GPT-6 Sol · OpenAI
3 failed
GPT-6 Astra · OpenAI
3 failed
GPT-5.6 Sol · OpenAI
4 failed
GPT-5.6 Terra · OpenAI
4 failed
Sonnet 5 · Anthropic
9 failed
GPT-5.6 Luna · OpenAI
12 failed
DeepSeek V4.1 Flash · DeepSeek
15 failed
Fable 5.1 · Anthropic
16 failed
GLM 5.3 Flash · Zhipu AI
21 failed
GPT-6 Luna · OpenAI
27 failed
Haiku 4.5 · Anthropic
40 failed

Wrong logic

The script runs, but it computes something other than what was asked.

  • Lookahead: a high that includes the bar being tested, or a daily value read before the day closed
  • Long-side rules applied to a short trade
  • Off-by-one windows, or a different formula from the one asked for

Why it is common. Nothing errors. The number is plausible, often better than the truth, so nothing prompts a second look.

Range across models
33% (Opus 5.5) to 75% (GPT-5.6 Terra)

Data and time handling

The logic is right, but it runs on the wrong slice of the market.

  • Warmup too short: a smoothed indicator that has not settled
  • Sessions and time zones: an exchange open hardcoded in UTC
  • Timeframes: mixed bar grids, seconds read as milliseconds

Why it is common. These pass on the window the model tested and fail on the hidden ones, which start on different dates.

Range across models
0% (Opus 5.5) to 0% (Opus 5.5)

Didn't flag a limit

The data cannot fully answer the question, and the model answered anyway.

  • A record high asked for on history that starts part way through
  • A volume answer on an instrument that carries no volume
  • A rate from a handful of events, presented like one from hundreds

Why it is common. The run succeeds. The reason not to trust it is one warning line in the summary.

Range across models
0% (Opus 5.5) to 33% (GPT-6 Astra)

Couldn't finish

The model ran out of time or budget, or never produced a script that runs.

  • Retrying the same fix against the same error
  • Rebuilding work that was already correct
  • Leaving validation errors unfixed

Why it happens. Every task has the same time, turn and usage limits. Long chains of work leave less room for a wrong turn early on.

Range across models
0% (GPT-6 Astra) to 67% (Opus 5.5)

Misreported the result

The final answer is not what Chartnaut returned.

  • Zero events from a bug, reported as "this never happens"
  • A subgroup's rate reported as the whole
  • A number from a different run, or rounded into a different answer

Why it matters most. This is the failure that would mislead a trader, because the report sounds confident and specific.

Range across models
0% (Opus 5.5) to 41% (GPT-6 Luna)

Other

Failures that fit none of the above.

  • Misread the task and answered a different question
  • A partial answer: one of the two numbers asked for
  • A tool or network error the model didn't retry

Why we show it. A list of causes with no remainder hides the cases that don't fit. We keep it.

Range across models
0% (Opus 5.5) to 25% (Fable 5.1)

How well the code runs

A right answer from a script that takes a minute per chart is not much use to a trader. So we run every script the model saved, not just check its answer, and measure the code itself the same way on every model.

  • SpeedRun time of the model's own script on the hidden windows, against our reference script for the same task. 1.0× means as fast as ours.
  • Budget warningsHow often the script hits Chartnaut's processing limits, such as more than 5 seconds in one call. On a live chart that is the script that gets paused.
  • LintEvery script the model saves must pass Chartnaut's validator. Warnings are allowed, errors are not.

Why it only counts when the answer is right

Fast code that gives the wrong answer is worth nothing, so efficiency only counts on solved tasks. A wrong answer scores 0 on that task however quickly it ran. A right answer scores between 0.5 and 1, depending on how efficient the code is.

task score = solved ? efficiency : 0
efficiency is 1.0 at or under the reference's run time, falls to 0.5 at 8 times slower, and loses 0.1 for each budget warning. It never goes below 0.5.
weighted is the average task score, shown in the leaderboard's Weighted column.

Solving still matters most: the weighting can never lift a wrong answer.

Right, and fast

Up is more tasks solved; left is faster code. The dashed line at 1.0× is our own reference scripts. The line joins the entries nothing beats on both.

30%40%50%60%70%80%90%100%1×1.25×1.5×2×2.5×3×code run time against our reference, log scale →Opus 5.5GPT-6 SolGPT-6 AstraGPT-5.6 SolGPT-5.6 TerraSonnet 5GPT-5.6 LunaDeepSeek V4.1 FlashFable 5.1GLM 5.3 FlashGPT-6 LunaHaiku 4.5
ModelSolvedSpeed vs referenceEfficiencyWeighted
GPT-6 Astra · OpenAI96%1.04×0.9491.1▲2
GPT-5.6 Terra · OpenAI95%1.05×0.9590.8▲3
Sonnet 5 · Anthropic90%1.08×0.9585.1▲3
GPT-6 Sol · OpenAI96%1.51×0.8684.0▼2
Opus 5.5 · Anthropic96%1.62×0.8482.3▼4
DeepSeek V4.1 Flash · DeepSeek81%1.09×0.9677.8▲2
GPT-5.6 Sol · OpenAI95%1.87×0.8077.1▼3
Fable 5.1 · Anthropic79%1.47×0.8770.0▲1
GLM 5.3 Flash · Zhipu AI73%1.11×0.9569.4▲1
GPT-5.6 Luna · OpenAI84%1.96×0.7968.2▼3
GPT-6 Luna · OpenAI65%2.72×0.7651.5·
Haiku 4.5 · Anthropic48%1.06×0.9445.5·

Speed is the median over solved attempts; above 1.0× is slower than the reference. Efficiency is the average over solved attempts only, so a model is not rewarded for fast code on answers it got wrong. The arrow shows places gained or lost against the Solved ranking.

Redoes all of history on every bar

const all = ctx.accum("h", [], (p) => [...p, ctx.high]);
const hi = Math.max(...all.slice(-21, -1));   // copies every bar so far, every bar

Keeps only what it needs

const last = ctx.accum("h", [], (p) => [...p, ctx.high].slice(-21));
const hi = Math.max(...last.slice(0, -1));    // 21 values, whatever the chart length

An illustration, not a task. Both give the same answer, so both pass the correctness checks. The first gets slower the longer the chart, and on long histories of minute bars it runs into the processing limits.

What we test

Every task sits in a tier, and every task above Foundation in a capability area. The tasks themselves stay closed. Below is what each tier and area asks of a model, with a generic example of the kind of work involved.

Foundation26 tasks · counts 1×

Separates working models from broken ones. One script, a precise spec, one answer.

Hard21 tasks · counts 2×

Separates good models from average ones. Each task hits at least two things that make trading code go wrong quietly.

Expert31 tasks · counts 3×

Separates the best models from each other. Several artifacts, long chains, and answers that must hold on data the model never saw.

Inside each area

What the tasks ask

Bar times arrive in UTC. Exchange clocks, daylight saving changes, weekends and session boundaries are the model's job.

The kind of work involved

  • An opening range defined in a named exchange time zone, over a stretch of dates that crosses a clock change.

A generic illustration. Real tasks give exact instruments, windows and output shapes, and are re-run on windows the model never sees.

Tasks
5
Average
87%
Top to bottom
40 pts

Solved, by entry

GPT-6 Sol100%
GPT-6 Astra100%
GPT-5.6 Terra100%
GPT-5.6 Luna100%
DeepSeek V4.1 Flash100%
Fable 5.1100%
Opus 5.580%
GPT-5.6 Sol80%
Sonnet 580%
GLM 5.3 Flash80%
GPT-6 Luna60%
Haiku 4.560%

Also in the framework, with no graded tasks in this run: layers and rendering, apps, long-horizon builds.

Why we built NautilusBench

Chartnaut is a trading platform where AI writes the research. That only works if the numbers are right.

In Chartnaut, a trader describes an idea and an AI model builds it: the indicator that draws it, the definition that marks every time it happened, the study that measures what came next. The trader then makes decisions from those numbers. A model that writes code which runs but counts the wrong thing is worse than no model at all, because the answer looks trustworthy.

General coding benchmarks do not tell us how often that happens. A model can be excellent at writing software and still read tomorrow's high into today's bar. So we built a benchmark out of the work Chartnaut asks models to do, graded against answers we can prove.

Choosing which models to offer

We use these results when deciding which models Chartnaut offers for each kind of work. A model has to do well on the areas a job needs, not just on the headline number.

Tracking capability over time

Every entry runs on the same bench version, with its run date on the table. As new models come out we run them the same way, so the table shows how model capability on trading research changes over time.

Making Chartnaut easier to get right

Every failure gets a cause. When many attempts stumble in the same place, we make that part of Chartnaut clearer, in the docs, the error messages or the tools, then measure again on the next version.

What this means for traders. You can choose a model on evidence from work like yours, and you know where to double-check: a result that looks too clean, a count of zero, an answer about history Chartnaut does not have.

How it works

Think of it as an exam where we mark the working as well as the answer, then hand the student a fresh set of numbers to see if the working still holds.

  1. 1

    The model gets a job

    A task in plain English, the kind a trader would type, with the exact result shape to hand back. It works in our harness: an empty Chartnaut project, the Chartnaut CLI, tools to read and edit files, 20 minutes and a fixed usage budget. Nothing else.

  2. 2

    It builds and runs scripts

    The model writes indicators, definitions and studies, and runs them on Chartnaut's servers. It never sees raw prices, so it cannot work the answer out in its head or look it up. The only way to the number is a script that is right.

  3. 3

    It hands back an answer

    A short result: a number, a count, a yes or no, or "this can't be answered, because...". It also names the run the answer came from.

  4. 4

    We check it three ways

    Does the named run exist, from the model's own saved script, on the window we asked for? Does the answer match what that run returned? And when we re-run the script on hidden windows it never saw, does it still match the certified answer? All three must pass.

Where the correct answers come from

Each task's answer is worked out twice: once by a Chartnaut script we wrote, once by a separate program written from scratch from raw bars, sharing no code with Chartnaut. A task enters the set only when the two agree on every window, public and hidden.

Why hidden windows

A script can be right on the dates it was tested on and wrong everywhere else: a warmup that is too short, a session hardcoded for one season. Re-running the model's own script on dates it never saw catches that. The result must match the model's run, and the run must hold up on the hidden windows.

What the range means

The range under each score is a 95% bootstrap interval: we resample the tasks thousands of times and see how far the score moves. When two entries' ranges overlap, we cannot honestly say one is better, so they share a rank.

Same conditions for every model

Every model runs in one harness we control, driven through the Chartnaut CLI. Same system prompt, same tools, same time, turn and usage limits. No lab's own coding app is involved, so a score reflects the model, not the wrapper around it.

How the headline is weighted

Tier-weighted counts Foundation 1×, Hard 2×, Expert 3×, so a model cannot reach the top on easy tasks alone. Solved is the plain share of tasks solved, and it is what the ranks and ranges use.

What it does not measure

Whether a strategy makes money. How nice the charts look. It measures one thing: given a research question, does the model get the right number and tell you the truth about it.

Why can't I see the tasks?

Once tasks are public, models get trained on them and the scores stop meaning anything. The task set stays closed. We describe what each area and tier tests, and add new tasks on market data that did not exist when the listed models were trained.

Why not run each model in its own lab's coding app?

Because then you are comparing the apps as much as the models. One app retries more, another reads more files first. Every model here runs in the same harness, with the same tools and limits, so the only thing that changes between rows is the model.

Is Chartnaut's own agent on the table?

No. The table ranks models from AI labs, run the same way. Ranking our own product here would mean marking our own homework.

Does a higher score mean better trades?

No. It means the model is more likely to give a correct number about the past, and to say when it cannot. It does not measure whether a strategy makes money.

Why is there only a small number of models?

This is the first published run. More models go on the table as they are run on the current bench version, each with its run date.

Key facts

Name
NautilusBench, by Chartnaut
Published by
Chartnaut, a trading platform for manual traders with AI research agents
What it measures
How accurately AI models do trading research: building indicators, definitions and studies, running them, and reporting correct numbers
Version
Bench v1, first published run, last updated 29 September 2026
Tasks
77 closed tasks in 8 categories and three tiers (26 Foundation, 21 Hard, 31 Expert)
Models
12 models from Anthropic, OpenAI, DeepSeek, Zhipu AI
How models run
One controlled harness for every model, with the Chartnaut CLI and file tools, 20 minutes and a fixed usage budget per task
Grading
Exact answers checked three ways, including a re-run on hidden windows; answers certified by two independent implementations; 95% bootstrap intervals; 1 run per task
Code quality
Every saved script is re-run and measured for speed and budget warnings. Efficiency (0.5 to 1.0) counts on solved tasks only; a wrong answer scores 0
Task set
Closed. Tasks, data, windows and reference code are never published
Top model
Opus 5.5 (Anthropic), 96% solved
v1 · Sep 2026First published run. 77 tasks, 12 entries.