AA-AnalystAgent
Whether an agent can answer the kind of quantitative question a business or data analyst faces day to day — working from a folder of real spreadsheets and documents, running Python in a sandbox, and returning the one right value on all five of five attempts.
What does an AnalystAgent task look like?
One task drops the agent into a workspace holding a folder of reference spreadsheets and documents (.xlsx and .docx, in some cases zipped) — real-world material such as a government expenditure report, a commodity trade release, a hydrology dataset, an energy cost model or a company valuation sheet — and asks one quantitative question about them: locate the right source figure and diagnose a discrepancy, filter and total a category, compute a ratio, project a trend or sensitivity, or work a P&L, cash-flow, balance-sheet or valuation calculation. The agent gets sandboxed Python 3.12 with pandas, openpyxl and the usual data libraries preinstalled, URL fetching, image viewing for vision-capable models, and a submit tool, under a 100-turn cap. It must submit only the answer value — a number, a percentage or a label, no explanation — and is graded correct or incorrect against a human-authored reference answer it never sees. Each question is run five times, and the headline metric counts it solved only if all five submissions are right. Artificial Analysis publishes example tasks (prompt and reference answer) on the leaderboard page; none is reproduced here, in keeping with this site's policy of never republishing a real test item.
How many points on AnalystAgent is real?
On AnalystAgent, a gap smaller than 11.2 points is treated as noise — with 80 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √80 ≈ 11.2, two standard errors on a benchmark this size.
Grade B: No major known issues on AnalystAgent, and the results here hold up — but with 80 items, close calls still deserve an independent check.
12 independent AnalystAgent scores — dots inside one dashed band are closer than AnalystAgent’s 11.2-point noise band, statistically indistinguishable.
Can you trust AnalystAgent? Caveats
AnalystAgent is actively discriminating between frontier models — a gap of at least 11.2 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.
Dataset and maintainer
AA-AnalystAgent is built, held and run by Artificial Analysis. Per its methodology page (checked 2026-09-29) the set is 80 quantitative questions across 14 business and scientific domains, including environmental reporting, trade and commodity statistics, healthcare expenditure reports, hydrology and weather data, government appropriations, energy cost models, financial models and project schedules, each paired with a folder of reference spreadsheets and documents (xlsx, docx) uploaded into the agent's workspace, and each with a human-authored reference answer that Artificial Analysis says it validated independently and that the agent never sees. The questions, answers and source files are privately held to limit contamination; only example tasks (prompt and reference answer) are published. It is reported as a standalone leaderboard and is not part of the Artificial Analysis Intelligence Index. Artificial Analysis's launch article is dated 2026-08-10 (26 models, Claude Opus 5 leading at 54%); Gemini 3.7 Flash took the lead in a results post three days later. The benchmark is seven weeks old at the time of writing.
Scoring mechanics
Every model is run as an agent through Stirrup, Artificial Analysis's open-source (MIT) harness, with a 100-turn cap per task and a small toolset — Python 3.12 in an isolated Linux sandbox with the question's files mounted and a pinned set of data libraries, URL fetching, image viewing for vision-capable models, and a final-answer tool. Models are told to submit only the value, no explanation. Each question is run five independent times. The leaderboard score is pass^5, the share of questions answered correctly on all five attempts; Artificial Analysis also publishes pass@1 (mean success per attempt) and pass@5 (solved at least once). Grading is binary against the held-out reference: every cell goes to an LLM judge, Gemini 3 Flash (Reasoning), but a deterministic numeric-equivalence pre-check overrides the judge whenever the answer equals the reference in the same unit convention at the asked precision. The pre-check is one-sided — it can only pass an answer, never fail one — so judge discretion is confined to the cells it cannot settle.
What pass^5 does to the numbers
The metric is built to punish inconsistency, and the tracked rows show how much. Gemini 3.7 Flash (high) leads at 60.0 on pass^5 against 70.5 on pass@1 and 77.5 on pass@5; MiniMax M3 sits at 10.0 on pass^5 against 44.0 on pass@1 and 73.75 on pass@5, the widest spread on the tracked set — it solves 73.75% of questions at least once, gets 44% of individual attempts right, and almost never gets all five. Claude Opus 4.8's 78.75 pass@5 is the highest ceiling among the tracked Claude rows while its 45.0 pass^5 is the lowest, so a reader who cares about one-shot reliability and one who cares about best-of-several attempts are reading two different boards. This site tracks pass^5 only, because it is the headline metric and the one every row on the board is ranked by.
Coverage and exclusions
12 of the 32 models tracked here have a row, every one run by Artificial Analysis under the same effort tier this site's other AA rows use for that model (Grok 4.6 at high, GPT-5.6 Sol and GPT-6 Astra at max, Gemini 3.7 Flash at high, Claude Fable 5.1 at max with default fallback). Two more tracked models appear on the board under a different build: Artificial Analysis's rows read "DeepSeek V4 Pro 0424" and "DeepSeek V4 Flash 0420", earlier checkpoints than the 0813 and 0731 builds this site's model records describe, so those scores are excluded rather than attributed — the same rule that drops MathArena's DeepSeek rows from the HMMT record. The other 18 tracked models, including Claude Opus 5.5, GPT-6 Sol and Luna, Grok 4.7, Gemini 3.8 Flash and GLM-5.3, did not appear in the board's 33-model selector on 2026-09-29: not yet run, not excluded. Values are stored exactly in this site's data (every pass^5 score is a multiple of 1.25, one question in 80); both Artificial Analysis's board and the table below display them rounded to one decimal.
Noise band
11.2 points: this site's standard 100/√n, rounded up to one decimal, with n = 80. The five repeats per question do not tighten it — the unit of the headline metric is the question, 80 all-or-nothing outcomes, and the 400 attempt-level cells feed each question's pass^5 decision rather than adding independent items. The band is the widest of any benchmark this site lists as current, and it decides most of the board: of the 66 pairs among the 12 tracked rows, 39 are ties and 27 are real gaps, and 18 of the 27 real gaps involve the leader or MiniMax M3; the other nine are Claude Fable 5.1, Claude Opus 5 and GPT-6 Astra over the bottom half of the board. Two calls sit on the edge of the band — Gemini 3.7 Flash over Claude Fable 5 and Claude Fable 5.1 over Claude Sonnet 5 are both 11.25-point gaps, real by 0.05 of a point — and four ties sit at a 10.0-point gap, one question (1.25 points) short of the band, so a single question changing hands would flip them.
Trust grade
B. Why not A: n = 80 is the smallest sample among the current-lifecycle benchmarks tracked here; the set is private, so nobody outside Artificial Analysis can audit a question, a reference answer or a run; and the LLM judge is Gemini 3 Flash while the board leader is Gemini 3.7 Flash — a same-vendor judge-and-leader overlap, the kind of conflict this site treats as a reason to withhold an A. The overlap is mitigated here, not absent: answers are bare values, and the deterministic pre-check overrides the judge on unambiguous matches, so judge discretion decides only the cells the pre-check cannot settle — and it can only pass, never fail. Why not C: there is no contamination signal (private, held-out, with reference answers the maintainer says it validated), one uniform harness (each row at the effort tier this site's other AA rows use for that model), an open-source harness anyone can read, and not a single vendor self-report on the board.
Sources · 7
- Artificial Analysis – AA-AnalystAgent leaderboard (the board this page tracks; its embedded payload carries 10 of the 12 values, read 2026-09-29)
- Artificial Analysis – "Announcing AA-AnalystAgent" launch article (2026-08-10; 26 models, Claude Opus 5 leading at 54%)
- Artificial Analysis – per-model pages (e.g. Grok 4.6), where all 12 values were re-read on 2026-09-29
- Artificial Analysis – Intelligence Benchmarking Methodology, AA-AnalystAgent section (dataset, pass^5 definition, grading, agent prompt; checked 2026-09-29)
- ArtificialAnalysis/Stirrup on GitHub – the open-source (MIT) agent harness every model on the board is run through
- AlphaSignal – "Artificial Analysis's AA-AnalystAgent Benchmark Reveals Claude Opus 5 Beats GPT-5.5 on Reliability" (2026-08-11; the earliest dated third-party coverage this site found)
- Artificial Analysis on X – results update (2026-08-13): Gemini 3.7 Flash (high) takes the lead at 60%, ahead of Claude Opus 5 at 54% and Fable 5 at 49%
Benchmarks related to AnalystAgent
- Toolathlon-Verified — Multi-tool chores, trust grade B
- Agents' Last Exam — Professional work, trust grade A
Is the current lead on AnalystAgent real?
TieA 2.5-point gap on AA-AnalystAgent (n=80) is inside the noise band (±11.2) — call it a tie.
Written up in full, for models on this AnalystAgent board: Claude Opus 5 vs Claude Sonnet 5 · Claude Opus 4.8 vs Gemini 3.1 Pro Preview · Gemini 3.1 Pro Preview vs Kimi K3 — every benchmark each pair shares, not just AnalystAgent.
AnalystAgent leaderboard: current scores
AnalystAgent’s 12 tracked scores were each verified against their source on 2026-09-29.
| Model | AnalystAgent score |
|---|---|
| Gemini 3.7 Flash | 60.0independent |
| Claude Fable 5.1 | 57.5independent |
| Claude Opus 5 | 53.8independent |
| GPT-6 Astra | 51.3independent |
| Claude Fable 5 | 48.8independent |
| GPT-5.6 Sol | 47.5independent |
| Claude Sonnet 5 | 46.3independent |
| Claude Opus 4.8 | 45.0independent |
| Gemini 3.1 Pro Preview | 41.3independent |
| Grok 4.6 | 41.3independent |
| Kimi K3 | 38.8independent |
| MiniMax M3 | 10.0independent |
AnalystAgent FAQ
What does AnalystAgent measure?
Whether an agent can answer the kind of quantitative question a business or data analyst faces day to day — working from a folder of real spreadsheets and documents, running Python in a sandbox, and returning the one right value on all five of five attempts. The benchmark has 80 items.
How big does a gap on AnalystAgent have to be to mean anything?
At least 11.2 points. Across 80 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.
Can you trust AnalystAgent scores?
We grade it B and list it as current — it still tells frontier models apart across 80 items. A gap of at least 11.2 points clears that noise band, but it only counts as a real gap once both scores come from independent runs.
What does pass^5 mean on AA-AnalystAgent?
Every one of the 80 questions is run five separate times, and a question counts as solved only if all five submissions are correct. The headline score is the share of questions solved that way. Artificial Analysis also publishes pass@1 (the average success rate per attempt) and pass@5 (solved at least once in five), which show the gap between a model's reliability and its ceiling: Gemini 3.7 Flash reads 60.0 on pass^5 against 77.5 on pass@5, while MiniMax M3 reads 10.0 against 73.75 — it reaches the right answer often, but rarely five times running.
Why are DeepSeek V4 Pro and V4 Flash missing here when Artificial Analysis lists them?
Because the rows on Artificial Analysis's board are labelled DeepSeek V4 Pro 0424 and DeepSeek V4 Flash 0420 — earlier checkpoints than the 0813 and 0731 builds this site's model records describe. A score earned by a different build is not attributed to the model page for a later one, the same rule that keeps MathArena's DeepSeek rows off this site's HMMT record. If Artificial Analysis runs the dated builds, the rows will be added.
Is the AA-AnalystAgent question set public, and who runs the scores?
No and one party. Artificial Analysis built the 80 questions, holds the questions, reference answers and source files privately to limit contamination, publishes only example tasks, and runs every model itself through its open-source Stirrup harness. That gives the board one uniform harness and no vendor self-reports, but it also means no one outside Artificial Analysis can audit an item or replicate a run — which is part of why this site grades it B rather than A.
Source: https://artificialanalysis.ai/evaluations/aa-analyst-agent