Kimi K2.6
Moonshot AI / Kimi · 2026-04-20 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Kimi K2.6
Kimi K2.6 是月之暗面的开源编码旗舰(Advancing Open-Source Coding),主打长程编码与代理蜂群(300 子代理 / 4000 步)。评测横跨编码、代理、数学与视觉:SWE-bench Verified 80.2、SWE-Bench Pro 58.6、Terminal-Bench 2.0(Terminus-2)66.7、HLE-Full w/ tools 54.0。
- 输入模态
- 文本 / 图像
- 上下文
- 256K
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Table (DOM summary table) / Footnotes · row: Humanity's Last Exam · quote_snippet: Humanity's Last Exam | 58.7 | 54.0 | 52.1 | 53.0 | 51.4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Default reporting condition is the HLE full set per footnote 2. Competitor cells: Kimi K3 58.7, GPT-5.4 52.1, Opus 4.6 53.0, Gemini 3.1 Pro 51.4, Model A 49.8, Model B 48.3. Tool condition resolved (2026-09-01): the appendix 'Benchmark table' prints HLE-Full w/ tools = 54.0 for K2.6, identical to this summary row, so the summary row is the w/ tools condition; the appendix no-tools row (34.7) is recorded separately.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Table (DOM summary table) / Footnotes · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: K3 91.2, GPT-5.4 89.6, Opus 4.6 88.8, Gemini 3.1 Pro 89.1, Model A 86.3, Model B 84.9. Page-internal divergence: the page's appendix 'Benchmark table' prints K2.6 = 90.5 for GPQA-Diamond; both values kept as separate evidence rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Table (DOM summary table) / Footnotes · row: AIME 2026
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aime-26 (AIME 2026) not yet in data/benchmarks/. Competitor cells: K3 96.7, GPT-5.4 95.0, Opus 4.6 92.8, Gemini 3.1 Pro 94.2, Model A 90.1, Model B 88.6. Page-internal divergence: the page's appendix 'Benchmark table' prints K2.6 = 96.4 for AIME 2026; both values kept as separate evidence rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Table (DOM summary table) / Footnotes · row: SWE-Bench Pro
{
"harness": "in-house SWE-agent-derived framework (bash/createfile/insert/view/strreplace/submit tools)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 10,
"aggregation": "mean over 10 independent runs",
"judge": null
}swebench-pro id already introduced by prior batches, still not in data/benchmarks/. Competitor cells: K3 63.4, GPT-5.4 57.7, Opus 4.6 53.4, Gemini 3.1 Pro 54.2, Model A 51.0, Model B 49.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Table (DOM summary table) / Footnotes · row: Terminal-Bench 2.0
{
"harness": "Terminus-2 with JSON parser, preserve thinking mode",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 10,
"aggregation": "mean over 10 independent runs",
"judge": null
}Competitor cells: K3 71.8, GPT-5.4 65.4, Opus 4.6 65.4, Gemini 3.1 Pro 68.5, Model A 61.9, Model B 60.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: Humanity's Last Exam (Full) w/ tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Section heading confirms the w/tools condition is reported on the page; numeric value renders in the non-DOM benchmark component. Footnote 3a: search + code-interpreter + web-browsing tools; 262,144-token budget with 49,152 per step; retain-latest-round context management. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 xhigh 52.1, Claude Opus 4.6 (max) 53.0, Gemini 3.1 Pro (thinking high) 51.4, K2.5 50.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Footnote: discard-all context management strategy, same as Kimi K2.5 and DeepSeek-V3.2. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 82.7, Claude Opus 4.6 83.7, Gemini 3.1 Pro 85.9, K2.5 74.9;agent swarm 行另列 K2.6 86.3 / K2.5 78.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: DeepSearchQA (f1-score)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}deepsearchqa id already introduced by prior batches, still not in data/benchmarks/. Metric is f1-score per section heading. Footnote: no context management for K2.6; over-context tasks counted as failed; competitor scores cited from Claude Opus 4.7 System Card. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:f1 行:GPT-5.4 78.6, Claude Opus 4.6 91.3, Gemini 3.1 Pro 81.9, K2.5 89.0;accuracy 行另列 K2.6 83.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}toolathlon id already introduced by prior batches, still not in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 54.6, Claude Opus 4.6 47.2, Gemini 3.1 Pro 48.8, K2.5 27.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OSWorld-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Value in non-DOM component. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 75.0, Claude Opus 4.6 72.7, Gemini 3.1 Pro '-', K2.5 63.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SWE-Multilingual
{
"harness": "in-house SWE-agent-derived framework",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 10,
"aggregation": "mean over 10 independent runs",
"judge": null
}new-benchmark: swebench-multilingual already introduced by prior batches, still not in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 '-', Claude Opus 4.6 77.8, Gemini 3.1 Pro 76.9*, K2.5 73.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MathVision w/ python
{
"harness": null,
"tools": [
"python"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "avg@3",
"judge": null
}mathvision id already introduced by prior batches, still not in data/benchmarks/. Footnote 5: python settings max-tokens-per-step 65,536, max-steps 50. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:w/ python 行:GPT-5.4 96.1*, Claude 84.6*, Gemini 95.7*, K2.5 85.0;无 python 行 K2.6 87.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: V* w/ python
{
"harness": null,
"tools": [
"python"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "avg@3",
"judge": null
}new-benchmark: vstar (V*) not yet in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 98.4*, Claude 86.4*, Gemini 96.9*, K2.5 86.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HLE-Full
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "max generation 98,304; context 262,144",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Appendix no-tools full-set row. Competitor cells: GPT-5.4 (xhigh) 39.8, Claude Opus 4.6 (max effort) 40.0, Gemini 3.1 Pro (thinking high) 44.4, Kimi K2.5 30.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / General Testing Details · row: Humanity's Last Exam (text-only subset) · quote_snippet: For the text-only subset, Kimi K2.6 achieves 36.4% accuracy without tools and 55.5% with tools.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "max generation 98,304",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose-printed subset values - distinct construct from HLE-Full rows; never average text-only with full set.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / General Testing Details · row: Humanity's Last Exam (text-only subset) · quote_snippet: 36.4% accuracy without tools and 55.5% with tools
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "max generation 98,304",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose-printed subset values - distinct construct from HLE-Full rows; never merge.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BrowseComp (agent swarm)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Swarm row - distinct from single-agent BrowseComp 83.2; only K2.6 has a value (all competitor cells '-'), Kimi K2.5 baseline 78.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: DeepSearchQA (accuracy)
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Accuracy companion metric - never merge with the f1-score row (K2.6 92.5). Competitor cells: GPT-5.4 63.7, Claude Opus 4.6 80.6, Gemini 3.1 Pro 60.2, Kimi K2.5 77.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: WideSearch (item-f1)
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}K2.6 value 80.8; all competitor columns '-'. The K2.5 baseline printed here (72.7) equals K2.5's single-agent base row on the K2.5 page, so this reads as the same single-agent protocol; swarm-vs-single is not labelled on this row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MCPMark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mcpmark not yet in data/benchmarks/. Competitor cells: GPT-5.4 62.5*, Claude Opus 4.6 56.7*, Gemini 3.1 Pro 55.9*, Kimi K2.5 29.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: Claw Eval (pass^3)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}pass^3 (all 3 attempts succeed) - distinct aggregation from pass@3 row; never merge. Competitor cells: GPT-5.4 60.3, Claude Opus 4.6 70.4, Gemini 3.1 Pro 57.8, Kimi K2.5 52.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: Claw Eval (pass@3)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}pass@3 (at least one of 3 succeeds) - distinct aggregation from pass^3 row; never merge. Competitor cells: GPT-5.4 78.4, Claude Opus 4.6 82.4, Gemini 3.1 Pro 82.9, Kimi K2.5 75.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: APEX-Agents
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote: evaluated on 452 of 480 public tasks. Competitor cells: GPT-5.4 33.3, Claude Opus 4.6 33.0, Gemini 3.1 Pro 32.0, Kimi K2.5 11.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SWE-Bench Verified
{
"harness": "in-house framework adapted from SWE-agent (minimal tool set)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 10,
"aggregation": "mean over 10 independent runs",
"judge": null
}Footnote 4: coding scores averaged over 10 independent runs; SWE series on in-house framework adapted from SWE-agent with bash/createfile/insert/view/strreplace/submit tools. Competitor cells: Claude Opus 4.6 80.8, Gemini 3.1 Pro 80.6, Kimi K2.5 76.8; GPT-5.4 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SciCode
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 10,
"aggregation": "mean over 10 independent runs",
"judge": null
}Footnote 4: coding scores averaged over 10 independent runs; SWE series on in-house framework adapted from SWE-agent with bash/createfile/insert/view/strreplace/submit tools. Distinct from K2 Thinking's SciCode protocol (temp 0.0, no tools) - never merge. Competitor cells: GPT-5.4 56.6, Claude Opus 4.6 51.9, Gemini 3.1 Pro 58.9, Kimi K2.5 48.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OJBench (python)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 10,
"aggregation": "mean over 10 independent runs",
"judge": null
}Footnote 4: coding scores averaged over 10 independent runs; SWE series on in-house framework adapted from SWE-agent with bash/createfile/insert/view/strreplace/submit tools. Python split - K2.5 page printed the cpp split (54.7 there); different language subset, never merge. Competitor cells: Claude Opus 4.6 60.3, Gemini 3.1 Pro 70.7, Kimi K2.5 54.7; GPT-5.4 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: LiveCodeBench (v6)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 10,
"aggregation": "mean over 10 independent runs",
"judge": null
}Footnote 4: coding scores averaged over 10 independent runs; SWE series on in-house framework adapted from SWE-agent with bash/createfile/insert/view/strreplace/submit tools. Competitor cells: Claude Opus 4.6 88.8, Gemini 3.1 Pro 91.7, Kimi K2.5 85.0; GPT-5.4 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Reasoning section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: AIME 2026
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page-internal divergence: this appendix row prints K2.6 = 96.4 while the page's summary DOM table prints 93.3 for the same benchmark+model; the page prints both without reconciliation, so both are kept as separate evidence rows. Competitor cells: GPT-5.4 99.2, Claude Opus 4.6 96.7, Gemini 3.1 Pro 98.3, Kimi K2.5 95.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Reasoning section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HMMT 2026 (Feb)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GPT-5.4 97.7, Claude Opus 4.6 96.2, Gemini 3.1 Pro 94.7, Kimi K2.5 87.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Reasoning section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: IMO-AnswerBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote: GPT-5.4 and Claude Opus 4.6 values cited from z.ai/blog/glm-5.1. Competitor cells: GPT-5.4 91.4 (cited), Claude Opus 4.6 75.3 (cited), Gemini 3.1 Pro 91.0*, Kimi K2.5 81.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Reasoning section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: GPQA-Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page-internal divergence: this appendix row prints K2.6 = 90.5 while the page's summary DOM table prints 88.4 for the same benchmark+model; both kept as separate evidence rows. Competitor cells: GPT-5.4 92.8, Claude Opus 4.6 91.3, Gemini 3.1 Pro 94.3, Kimi K2.5 87.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MMMU-Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 5: vision protocol max-tokens 98,304, avg@3. MMMU-Pro follows the official protocol (input order preserved, images prepended). Competitor cells: GPT-5.4 81.2, Claude Opus 4.6 73.9, Gemini 3.1 Pro 83.0*, Kimi K2.5 78.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MMMU-Pro w/ python
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 5: vision protocol max-tokens 98,304, avg@3. Python-tool setting: max-tokens-per-step 65,536, max-steps 50 - never merge with no-python row. Competitor cells: GPT-5.4 82.1, Claude Opus 4.6 77.3, Gemini 3.1 Pro 85.3*, Kimi K2.5 77.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: CharXiv (RQ)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 5: vision protocol max-tokens 98,304, avg@3. Competitor cells: GPT-5.4 82.8*, Claude Opus 4.6 69.1, Gemini 3.1 Pro 80.2*, Kimi K2.5 77.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: CharXiv (RQ) w/ python
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 5: vision protocol max-tokens 98,304, avg@3. Python-tool setting: max-tokens-per-step 65,536, max-steps 50 - never merge with no-python row. Competitor cells: GPT-5.4 90.0*, Claude Opus 4.6 84.7, Gemini 3.1 Pro 89.9*, Kimi K2.5 78.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MathVision
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 5: vision protocol max-tokens 98,304, avg@3. Distinct from the w/ python row (93.2); never merge. Competitor cells: GPT-5.4 92.0*, Claude Opus 4.6 71.2*, Gemini 3.1 Pro 89.8*, Kimi K2.5 84.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BabyVision
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 5: vision protocol max-tokens 98,304, avg@3. Competitor cells: GPT-5.4 49.7, Claude Opus 4.6 14.8, Gemini 3.1 Pro 51.6, Kimi K2.5 36.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BabyVision w/ python
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 5: vision protocol max-tokens 98,304, avg@3. Python-tool setting: max-tokens-per-step 65,536, max-steps 50 - never merge with no-python row. Competitor cells: GPT-5.4 80.2*, Claude Opus 4.6 38.4*, Gemini 3.1 Pro 68.3*, Kimi K2.5 40.5.