HMMT Feb 2026

Whether a model can solve HMMT Feb 2026's short-answer competition math problems — pre-college-level algebra, number theory, combinatorics, and geometry, each with one objectively-checkable numeric or symbolic answer.

Items
33
Trust grade
C
Status
contaminated

What does a HMMT task look like?

One task presents a single HMMT-style problem stated in a few sentences of prose (any needed diagram is described rather than rendered as a complex image), drawn from algebra, number theory, combinatorics, or geometry. The model must work through to one specific final value — an integer, a fraction, or a simple closed-form expression — with no multiple-choice options offered and no partial credit for correct intermediate steps. Grading is exact match against the competition's official answer key, not an LLM judge, so a right method with a slipped final value still scores zero. No official HMMT item is reproduced here: the problems are the competition's own published material, and this site's standing policy is to never republish a real test item regardless of a benchmark's specific copyright terms — only the format is described above.

How many points on HMMT is real?

On HMMT the 17.5-point noise floor is beside the point: the benchmark is contaminated, so every comparison across its 33 items is labeled tainted instead of being measured against a threshold.

Grade C: Usable, but HMMT's 33-item sample is small or the benchmark is nearing saturation — treat close results with extra skepticism.

4 independent HMMT scores, no noise band drawn: the benchmark is contaminated, so every HMMT head-to-head is labelled tainted rather than measured against its 17.5-point threshold. How tightly the HMMTdots bunch above is the reason for that call.

Trust caveats

HMMT has known data leakage or gaming issues — treat its scores with real skepticism.

Dataset and maintainer

HMMT Feb 2026 is the February edition of the Harvard-MIT Mathematics Tournament (hmmt.org), a student-run competition-math contest for pre-college (high school) students, organized by students at Harvard and MIT. MathArena (SRI Lab at ETH Zurich and INSAIT — matharena.ai, arXiv:2605.00674 and its predecessor arXiv:2505.23281, presented at NeurIPS 2025's Datasets & Benchmarks track) evaluates model performance against each competition's actual problems shortly after release and publishes accuracy, cost, and token usage on a public leaderboard. The tracked page is Final-Answer Comps → HMMT Feb 2026 (verified live 2026-08-25 via matharena.ai/?comp=hmmt--hmmt_feb_2026), listing 33 problems evaluated across 32 models per matharena.ai/competitions — up from the 30-problem HMMT Feb 2025 edition tracked as a now-deprecated competition on the same site, so raw item counts aren't directly comparable year over year.

Scoring mechanics

Each problem carries one fixed correct answer (numeric or symbolic), graded by exact match against the official key — no LLM judge, no partial credit, no multiple-choice guess floor. Per MathArena's own harness documentation (eth-sri/matharena on GitHub, `--n` flag: "Number of runs per problem (default: 4)"), each model is run 4 times on every problem by default, with the leaderboard's reported accuracy and cost being the average across those runs — consistent with MathArena's own stated methodology of running each model repeatedly per problem and averaging score and cost.

Contamination caveat — read this before trusting any score on this page: every one of this site's 4 tracked scores on this benchmark carries MathArena's own disclosed contamination-warning flag, shown directly on the leaderboard as "Model was released after competition release." Verified live on 2026-08-25, each flagged individually: Claude-Opus-4.8 (max) 95.45%, Gemini 3.1 Pro Preview 94.70%, Kimi K3 (Think) 96.97%, GLM 5.2 92.42%. This isn't a suspicion — it's MathArena's own dating logic: every model this site tracks was released after HMMT Feb 2026 took place, so none of them can be assumed to have had no exposure to these exact problems (or writeups of them) during training or post-training. There is no clean, unflagged subset of tracked rows to fall back on for this benchmark.

DeepSeek V4 Pro (0813) and DeepSeek V4 Flash (0731) are deliberately excluded from this record, not merely unmatched. MathArena's own evaluation config for both models (eth-sri/matharena GitHub, configs/models/deepseek/deepseek_v4_pro.yaml and deepseek_v4_flash.yaml, checked live 2026-08-25) hardcodes `date: "2026-04-24"` and points to the unversioned Hugging Face repos deepseek-ai/DeepSeek-V4-Pro and deepseek-ai/DeepSeek-V4-Flash — not the dated -0813/-0731 repos this site's own model records cite as the source of their specs. Pulling each repo's live config.json (checked 2026-08-25) confirms this is a materially different checkpoint, not just a naming gap: the unversioned Pro and Flash repos (Hugging Face API `lastModified` 2026-06-22 for both) lack the `dspark_block_size`/`dspark_target_layer_ids` fields present in the -0813 and -0731 repos' configs — the DSpark speculative-decoding module documented as new to those GA builds. MathArena benchmarked an earlier, architecturally distinct build of both models, so its HMMT Feb 2026 scores for "DeepSeek-v4-Pro (Max)" (93.94%) and "DeepSeek-v4-Flash (Max)" (93.94%) are dropped from this record rather than attributed to deepseek-v4-pro or deepseek-v4-flash.

Trust grade

C. The grade lands here because of a specific, disclosed mechanism rather than a vague worry: MathArena flags the contamination risk itself, per row, on its own public leaderboard — the source is surfacing the problem, not hiding it, which is what keeps this off a D. It doesn't go higher than C because the risk isn't confined to a minority of rows the way it is on HLE's chem/bio subset dispute (FutureHouse's audit found contested answer keys in roughly 18-29% of a 321-item subset out of HLE's 2,500 total items, per this site's own hle notes) — here it affects 100% of the 4 rows tracked, with no unflagged control group to point to. USAMO 2026 coverage on MathArena is out of scope for this record; its thinner sample size and proof-graded (rather than numeric-answer) format don't bundle cleanly with HMMT's short-answer format and would need its own benchmark record if added later.

Is the current lead on HMMT real?

TaintedHMMT Feb 2026 is contaminated — rankings here are unreliable regardless of the gap.

Current scores

HMMT’s 4 tracked scores were verified against their sources on or after 2026-08-25.

Kimi K3
Claude Opus 4.8
Gemini 3.1 Pro Preview
GLM-5.2

FAQ

What does HMMT measure?

Whether a model can solve HMMT Feb 2026's short-answer competition math problems — pre-college-level algebra, number theory, combinatorics, and geometry, each with one objectively-checkable numeric or symbolic answer. The benchmark has 33 items.

How big does a gap on HMMT have to be to mean anything?

No gap on HMMT qualifies. Because the benchmark is contaminated, we label every head-to-head here tainted and never apply the 17.5-point threshold its 33 items would otherwise imply.

Can you trust HMMT scores?

We grade it C and list it as contaminated — known leakage means a high HMMT score can reflect memorised answers rather than ability, so every head-to-head on its 33 items is labelled tainted.

Source: https://matharena.ai/?comp=hmmt--hmmt_feb_2026