Wrong logic
48%The script runs, but it computes something other than what was asked.
- Lookahead: a high that includes the bar being tested, or a daily value read before the day closed
- Long-side rules applied to a short trade
- Off-by-one windows, or a different formula from the one asked for
Why it is common. Nothing errors. The number is plausible, often better than the truth, so nothing prompts a second look.
- Range across models
- 33% (Opus 5.5) to 75% (GPT-5.6 Terra)
Data and time handling
0%The logic is right, but it runs on the wrong slice of the market.
- Warmup too short: a smoothed indicator that has not settled
- Sessions and time zones: an exchange open hardcoded in UTC
- Timeframes: mixed bar grids, seconds read as milliseconds
Why it is common. These pass on the window the model tested and fail on the hidden ones, which start on different dates.
- Range across models
- 0% (Opus 5.5) to 0% (Opus 5.5)
Didn't flag a limit
6%The data cannot fully answer the question, and the model answered anyway.
- A record high asked for on history that starts part way through
- A volume answer on an instrument that carries no volume
- A rate from a handful of events, presented like one from hundreds
Why it is common. The run succeeds. The reason not to trust it is one warning line in the summary.
- Range across models
- 0% (Opus 5.5) to 33% (GPT-6 Astra)
Couldn't finish
28%The model ran out of time or budget, or never produced a script that runs.
- Retrying the same fix against the same error
- Rebuilding work that was already correct
- Leaving validation errors unfixed
Why it happens. Every task has the same time, turn and usage limits. Long chains of work leave less room for a wrong turn early on.
- Range across models
- 0% (GPT-6 Astra) to 67% (Opus 5.5)
Misreported the result
12%The final answer is not what Chartnaut returned.
- Zero events from a bug, reported as "this never happens"
- A subgroup's rate reported as the whole
- A number from a different run, or rounded into a different answer
Why it matters most. This is the failure that would mislead a trader, because the report sounds confident and specific.
- Range across models
- 0% (Opus 5.5) to 41% (GPT-6 Luna)
Other
5%Failures that fit none of the above.
- Misread the task and answered a different question
- A partial answer: one of the two numbers asked for
- A tool or network error the model didn't retry
Why we show it. A list of causes with no remainder hides the cases that don't fit. We keep it.
- Range across models
- 0% (Opus 5.5) to 25% (Fable 5.1)