Hy3
Tencent / 腾讯混元 · 2026-07-06 · 类别未确认
Hy3
腾讯发布混元 Hy3:295B 总参/21B 激活 MoE、256K 上下文、三种思考模式,Apache 2.0 开源。评测覆盖编码、智能体搜索与推理,亮点如 SWE-bench Verified 78%、BrowseComp 84.2%。
- 输入模态
- 文本
- 上下文
- 256K
- 参数
- 295B-A21B MoE
- 价格(每百万 tokens)
- CNY 输入 1 / 输出 4 · 页面原文为 1/4/0.25 元/百万 tokens 三档报价,第三档含义未标注
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 更可靠的产品体验 → 复杂上下文承接与多轮意图保持能力 · row: MRCR · quote_snippet: 长对话理解基准中取得显著跨越(如 MRCR 从 42.9% 升至 75.1%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page prose value (DOM text). The 42.9% in the same sentence refers to Hy3 preview, not a competitor. Long-context understanding benchmark family; page does not print the MRCR variant/needle count — variant unspecified.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 更可靠的产品体验 → 输出格式和工具调用稳定性 · row: SWE Bench Verified · quote_snippet: 在 SWE Bench Verified 上的分数标准差控制在 4 个百分点以内
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose mentions SWE Bench Verified by name but reports no score for Hy3; the disclosure is a cross-scaffold robustness claim (std < 4pp across Codebuddy / Cline / KiloCode scaffolds). Score value appears only in the appendix image (see the separate pending image row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: swebench-multilingual · figure: data/model-releases/official/_archive/tencent/hy3/images/split/01.png
{
"harness": "SWE-agent scaffold",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/01.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: swebench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/01.png
{
"harness": "SWE-agent scaffold",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/01.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: swebench-pro · figure: data/model-releases/official/_archive/tencent/hy3/images/split/01.png
{
"harness": "SWE-agent scaffold",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/01.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: terminalbench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/01.png
{
"harness": "Terminus-2 scaffold, xml parser, 4h agent timeout, 16 cores / 32 GB, max 500 episodes",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/01.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: nl2repo · figure: data/model-releases/official/_archive/tencent/hy3/images/split/01.png
{
"harness": "Claude Code, 250-turn budget, 12000s timeout, 4 CPU / 32 GB, anti-hacking constraints",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/01.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: deepswe · figure: data/model-releases/official/_archive/tencent/hy3/images/split/01.png
{
"harness": "mini-swe-agent, 2h timeout per task, 2 CPU / 8 GB, strict network isolation",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/01.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-backend-2-0 · figure: data/model-releases/official/_archive/tencent/hy3/images/split/01.png
{
"harness": "Claude Code scaffold",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-backend-2-0 not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/01.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-swe-max · figure: data/model-releases/official/_archive/tencent/hy3/images/split/01.png
{
"harness": "Claude Code scaffold",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-swe-max not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/01.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-companybench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": "Claude Code scaffold",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-companybench not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: browsecomp · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": "internal harness; self-summary context management (agentic search)",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: widesearch · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": "internal harness (agentic search)",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: deepsearchqa · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": "internal harness (agentic search)",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: mcp-atlas · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": "Scale April 2026 methodology, 100 tool-call budget per task",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Gemini 2.5 Pro"
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: toolathlon · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: apex-agents · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: claw-eval · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": "internal harness, 20260325 version (105 queries)",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Gemini-3.5-flash"
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: wildclawbench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": "OpenClaw Harness",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: skillsbench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/02.png
{
"harness": "Claude Code",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/02.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: e-bench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: e-bench not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-finmodelbench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-finmodelbench not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: prodbench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": "OpenClaw Harness",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: prodbench not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-skillsworld · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-skillsworld not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hlehle · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-euler-pro · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-euler-pro not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: gpqa · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hlehle · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: frontier-science-research · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: frontier-science-olympiad · figure: data/model-releases/official/_archive/tencent/hy3/images/split/03.png
{
"harness": "self-evaluated with OpenAI FrontierScience paper judge prompts",
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "gpt-oss-120b (high reasoning effort)"
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/03.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: usamo-2026 · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: matharena-apex · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: matharena-apex not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: arxivmath · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: arxivmath not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: horizonmath · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: hy-math · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hy-math not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: phybench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: cmt-benchmark · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cmt-benchmark not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: imo-answerbench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: superchem · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: superchem not yet in data/benchmarks/. Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: cl-bench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/04.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/04.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: cl-bench · figure: data/model-releases/official/_archive/tencent/hy3/images/split/05.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/05.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 附录:模型得分 · row: aa-lcr · figure: data/model-releases/official/_archive/tencent/hy3/images/split/05.png
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest tier (per appendix note: reasoning effort set to the highest tier for all models)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value read from the page appendix image and confirmed by re-reading the archived image with the Read tool on 2026-09-01 (data/model-releases/official/_archive/tencent/hy3/images/split/05.png). Page prints the row under the Hy3 column; protocol fields transcribed from the notes block printed beneath the same table.