Tencent

Tencent Hy4 preview

Tencent's open-weight 770B flagship under Apache 2.0, at $0.834/$2.501 per million tokens, flat across a 1M window. All six tracked scores are Tencent's own, and none has an independent run a month after release. The predecessor offers some reassurance: Tencent's current Hy3 figures sit within about a point of independent runs on three benchmarks, and inside Terminal-Bench 2.1's noise band on the fourth.

Tencent Hy4 preview benchmarks and pricing, every number sourced: 6 tracked Tencent Hy4 preview benchmark scores (0 independently run, 6 still resting on a vendor’s own claim), priced at $0.83 per million input tokens and $2.50 per million output.

Tencent Hy4 preview architecture: Mixture-of-Experts; 770B total parameters (49B activated per token); 1M-token context window.

Tencent Hy4 preview’s 6 benchmark scores on this page were verified against their source on 2026-09-27.

Released
2026-08-28
License
open-weights
Context window
1M tokens
Knowledge cutoff
Not disclosed
Parameters
770B (49B active)
Architecture
Mixture-of-Experts

Tencent Hy4 preview’s verified record

Tencent Hy4 preview’s most-compared rival is DeepSeek V4 Pro (0813): 6 not callable across their 6 shared comparisons. Tencent Hy4 preview is priced at $0.83/$2.50 per 1M tokens (in/out) vs DeepSeek V4 Pro (0813)’s $1.32/$3.96.

Against the 130 head-to-head comparisons Tencent Hy4 preview shares with other tracked models: 0 real gaps, 0 inside the noise band, and 130 we will not call.

A gap counts for Tencent Hy4 preview only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Tencent Hy4 preview trails on 0 of them.

No verdict for Tencent Hy4 preview anywhere on DeepSWE, HLE · no tools, HLE · with tools, Terminal-Bench 2.1, Toolathlon-Verified (nothing independently confirmed on both sides); GPQA Diamond (saturated).

  • None of Tencent Hy4 preview’s reasoning comparisons are independently confirmed on both sides yet.
  • None of Tencent Hy4 preview’s coding comparisons are independently confirmed on both sides yet.
  • None of Tencent Hy4 preview’s agentic comparisons are independently confirmed on both sides yet.

Tencent Hy4 preview API pricing

$0.83 in / $2.50 out per 1M tokens — official pricing

What Tencent Hy4 preview costs per job

Tencent Hy4 preview cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.108
A codebase review1,000K / 100K$1.08
A day of agent work10,000K / 1,000K$10.84

Computed from Tencent Hy4 preview’s list rates above — cache discounts and batch tiers are not applied.

Tencent Hy4 preview benchmark scores

Tencent Hy4 preview benchmark scores, provenance, and source links
BenchmarkScore
Terminal-Bench 2.1[1]
Terminal ops · ±10.6 is noise
DeepSWE[2]
Long-horizon coding · ±9.5 is noise
Toolathlon-Verified[3]
Multi-tool chores · ±9.7 is noise
GPQA Diamondsaturated[4]
Expert science Q&A — not ranked at any gap size
HLE(no tools)[5]
Reasoning · ±2 is noise
HLE(with tools)[6]
Reasoning · ±2 is noise

Who ran these numbers: 0 of 6 independent; vendor self-reported (6).

  1. Terminal-Bench 2.1: Tencent's model-card appendix figure at the highest reasoning setting, run on a Claude Code harness with up to 500 turns and a 12-hour timeout per trial. For the predecessor, Tencent's re-run figure (70.8) sits 6.4 points above Artificial Analysis's independent run (64.4) — inside this benchmark's 10.6-point noise band, with different harnesses. tbench.ai's 2.1 board did not list Hy4 on 2026-09-27.
  2. DeepSWE: Tencent's model-card appendix figure at the highest reasoning setting, run on mini-swe-agent (8 CPUs, 16 GB per task) — the same harness the independent deepswe.datacurve.ai board uses, but Tencent's own run. That board did not list Hy4 on 2026-09-27.
  3. Toolathlon-Verified: Tencent's model-card appendix figure at the highest reasoning setting, run on Tencent's internal agent scaffold with some MCP tools reimplemented and a 2-hour timeout instead of the official 5,400 seconds. toolathlon.xyz lists the predecessor Hy3 (56.8) but not Hy4 as of 2026-09-27.
  4. GPQA Diamond: Tencent's model-card appendix figure at the highest reasoning setting; harness not stated. GPQA Diamond is graded saturated on this site. Artificial Analysis had no Hy4 entry on 2026-09-27.
  5. HLE: Tencent's model-card appendix figure (text-only subset, no tools) at the highest reasoning setting. Artificial Analysis had no Hy4 entry on 2026-09-27; for the predecessor, Tencent's re-run figure (34.4) sat within a point of AA's independent run (33.5).
  6. HLE: Tencent's model-card appendix figure (text-only subset, with tools) at the highest reasoning setting. No independent with-tools run found.

Notes on the record

Pricing per Tencent's announcement: $0.834 input / $2.501 output per million tokens, $0.042 cached. Tencent Cloud's TokenHub lists the model at ¥6 / ¥18 (¥0.3 cache hit) with no length condition and no peak/off-peak rate — one flat rate across the window (checked 2026-09-27). No API launch promotion.

Weights are on Hugging Face (tencent/Hy4-preview) under Apache 2.0: a 770B-parameter mixture-of-experts with 49B active, a 1M-token context (TokenHub caps input at 960K), 64K max output, text in and text out. Two reasoning modes, no_think and high (the default). Knowledge cutoff not disclosed. Tencent calls it an early version shipped with known issues, including "spending longer than necessary reasoning" and "a tendency to over-verify its own work."

All six tracked scores come from the benchmark appendix image on Tencent's model card, each at the model's highest reasoning setting. The harnesses matter: Terminal-Bench 2.1 ran on Claude Code with up to 500 turns and a 12-hour timeout; Toolathlon-Verified on Tencent's own scaffold, with some MCP tools reimplemented and a longer timeout; DeepSWE on mini-swe-agent, the harness the independent board uses. Tencent also reports Agents' Last Exam on the 105-task ALE-CLI split (22.8); this site tracks the full split, so that figure is not recorded. Artificial Analysis, tbench.ai, Toolathlon, DeepSWE, ARC Prize, LiveBench, Agents' Last Exam and MathArena had no Hy4 row on 2026-09-27.

The predecessor gives a rough calibration point. The appendix includes Tencent's re-run of Hy3 — its notes say some Hy3 scores changed after harness, judge and anti-hacking updates — at 34.4 on HLE without tools, 90.9 on GPQA Diamond, 56.2 on Toolathlon-Verified and 70.8 on Terminal-Bench 2.1. Independent runs land within about a point on the first three — 33.5 and 89.7 (Artificial Analysis), 56.8 (Toolathlon's own board) — and 6.4 points lower on Terminal-Bench 2.1 (Artificial Analysis, 64.4), inside that benchmark's 10.6-point noise band and across different harnesses. Hy3 is not tracked on this site.

Compare with

FAQ

Has Tencent Hy4 preview been independently benchmarked?

Not yet, a month after its 2026-08-28 release. Artificial Analysis, tbench.ai, Toolathlon, the DeepSWE board, ARC Prize, LiveBench, Snorkel's Agents' Last Exam and MathArena had no Hy4 row when checked on 2026-09-27, and vals.ai no longer runs new models on GPQA Diamond, SWE-bench Verified or LiveCodeBench. Every score on this page is Tencent's own, read from its model card.

How much does Tencent Hy4 preview cost?

$0.834 per million input tokens and $2.501 per million output tokens, with cached input at $0.042, per Tencent's announcement. Tencent Cloud's TokenHub lists ¥6 and ¥18 per million tokens with no length-based tier and no peak/off-peak pricing, so the rate is flat across the 1M-token window. It is also sold through OpenRouter at the same USD rates.

Is Tencent Hy4 preview open source?

The weights are. Tencent published Hy4 preview and an FP8 version on Hugging Face under the Apache 2.0 license on 2026-08-28, alongside hosted access through Tencent Cloud's TokenHub and OpenRouter. The model is text-only, with two reasoning modes: no_think and high.

How far can Tencent's own benchmark numbers be trusted?

The predecessor is the best evidence available, with a caveat: the Hy3 figures in Hy4's appendix are Tencent's re-runs, not its original launch claims. Those re-runs sit within about a point of independent results on Humanity's Last Exam, GPQA Diamond and Toolathlon-Verified, and 6.4 points above Artificial Analysis on Terminal-Bench 2.1 — inside that benchmark's 10.6-point noise band, with different harnesses on each side. That is a reasonable record, but it is not a substitute for an independent run of Hy4.

How does Tencent Hy4 preview compare with Hy3?

On Tencent's own table the jump is large: DeepSWE from 28.0 to 64.3, Humanity's Last Exam without tools from 34.4 to 43.4, Terminal-Bench 2.1 from 70.8 to 85.4. Tencent calls it the largest generation-over-generation gain it has measured. Hy3 is still sold on TokenHub at a lower price, and no Tencent page describes Hy4 as replacing it.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.