Hy4 preview
Tencent / 腾讯混元 · 2026-08-28 · 类别未确认
Hy4 preview
发布文将 Hy4 preview 定位为 770B 总参/49B 激活、1M 上下文、Apache 2.0 开源的旗舰预览版(各模型均按最高推理档评测)。46 项评测集中于智能体编码与搜索、工作智能体及 STEM 推理,亮点为 Terminal-Bench 2.1 85.4 与 SWE-bench Multilingual 82.9。
- 输入模态
- 文本
- 上下文
- 1M
- 参数
- 770B-A49B MoE
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: swebench-multilingual · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png
{
"harness": "swe-agent scaffold",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: swebench-pro · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png
{
"harness": "swe-agent scaffold",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: deepswe · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png
{
"harness": "mini-swe-agent; sandboxed 8 CPU / 16 GB per task",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: swe-atlas-codebase-qna · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png
{
"harness": "Claude Code, 256-turn budget; rubric judge parsing fixed + network allowlist anti-hacking",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: swe-atlas-test-writing · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png
{
"harness": "Claude Code, 256-turn budget",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: swe-atlas-refactoring · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png
{
"harness": "Claude Code, 256-turn budget",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swe-atlas-refactoring not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/01.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: swe-marathon · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png
{
"harness": "official settings except agent timeout at 2x official; each task repeated until 8 valid scores",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: terminalbench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png
{
"harness": "Claude Code harness, up to 500 turns, 12h timeout per trial, 16 CPUs / 32 GB RAM",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: nl2repo · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png
{
"harness": "Claude Code, 1000-turn budget per task, anti-hacking prompt constraints + tool-call monitoring",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: cybergym · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: program-bench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png
{
"harness": "Claude Code, 2000-turn budget per task (DSV4-Pro-0803 switched to mini-swe-agent)",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: posttrain-bench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: harbor-index · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: harbor-index not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-backend-2-0 · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-backend-2-0 not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/02.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-swe-max · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png
{
"harness": "Claude Code scaffold, 200-turn budget, 16 CPU / 32 GB, web access disabled text-only, 3-run mean, pass/fail hidden verifiers",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": null,
"judge": null
}new-benchmark: hy-swe-max not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-companybench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-companybench not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: widesearch · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png
{
"harness": "internal harness (agentic search)",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: one-million-bench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: draco · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-lifesearch · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-lifesearch not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-browsecomp-pro2 · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-browsecomp-pro2 not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: officeqa-pro · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png
{
"harness": "Claude Code (GPT series evaluated with Codex CLI)",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/03.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: mcp-atlas · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png
{
"harness": "Scale April 2026 methodology, 100 tool-call budget retained",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Gemini 3.1 Pro Preview"
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: toolathlon · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png
{
"harness": "internal agent scaffold with re-implemented MCP tools, 2 CPU / 10 GB, 2h timeout, 100-step budget, 3-run pass@1",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: apex-agents · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png
{
"harness": "official ReAct Toolbelt harness, max 250 steps per trial, 4 CPUs / 16 GB",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: skillsbench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png
{
"harness": "Claude Code",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: jobbench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: workspace-bench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: agents-last-exam · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png
{
"harness": "official evaluation protocol, Claude Code harness, up to 12h, 8 CPU / 32 GB, official ALE evaluators",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: gdpval-aa · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/04.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: automationbench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: bankertoolbench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png
{
"harness": "OpenCode scaffold",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Gemini 3 flash"
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: e-bench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: e-bench not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: e-bench-code · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: e-bench-code not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-finagentbench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-finagentbench not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-finmodelbench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-finmodelbench not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: biomysterybench · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png
{
"harness": "Claude Code, 2h / 250 turns per task (gpt-5.6-luna/sol via Codex)",
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Kimi-K3"
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hlehle · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/05.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: critpt · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: gpqa · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hlehle · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: superchem · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: superchem not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: arxivmath · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: arxivmath not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: horizonmath · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: matharena-apex · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: matharena-apex not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: brokenarxiv · figure: data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest available reasoning setting (per appendix note)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: brokenarxiv not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy4-preview/images/split/06.png). Page prints the row under the Hy4 preview column; protocol fields transcribed from the notes block printed beneath the same table.