← 模型目录

Kimi K2.6

Moonshot AI / Kimi · 2026-04-20 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Kimi K2.6

Kimi K2.6 是月之暗面的开源编码旗舰(Advancing Open-Source Coding),主打长程编码与代理蜂群(300 子代理 / 4000 步)。评测横跨编码、代理、数学与视觉:SWE-bench Verified 80.2、SWE-Bench Pro 58.6、Terminal-Bench 2.0(Terminus-2)66.7、HLE-Full w/ tools 54.0。

输入模态
文本 / 图像
上下文
256K
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 54.0 模型 kimi-k2-6 · 版本 Full set · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Table (DOM summary table) / Footnotes · row: Humanity's Last Exam · quote_snippet: Humanity's Last Exam | 58.7 | 54.0 | 52.1 | 53.0 | 51.4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Default reporting condition is the HLE full set per footnote 2. Competitor cells: Kimi K3 58.7, GPT-5.4 52.1, Opus 4.6 53.0, Gemini 3.1 Pro 51.4, Model A 49.8, Model B 48.3. Tool condition resolved (2026-09-01): the appendix 'Benchmark table' prints HLE-Full w/ tools = 54.0 for K2.6, identical to this summary row, so the summary row is the w/ tools condition; the appendix no-tools row (34.7) is recorded separately.

打开官方来源

gpqa 88.4 模型 kimi-k2-6 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Table (DOM summary table) / Footnotes · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: K3 91.2, GPT-5.4 89.6, Opus 4.6 88.8, Gemini 3.1 Pro 89.1, Model A 86.3, Model B 84.9. Page-internal divergence: the page's appendix 'Benchmark table' prints K2.6 = 90.5 for GPQA-Diamond; both values kept as separate evidence rows.

打开官方来源

aime-26 93.3 模型 kimi-k2-6 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Table (DOM summary table) / Footnotes · row: AIME 2026

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: aime-26 (AIME 2026) not yet in data/benchmarks/. Competitor cells: K3 96.7, GPT-5.4 95.0, Opus 4.6 92.8, Gemini 3.1 Pro 94.2, Model A 90.1, Model B 88.6. Page-internal divergence: the page's appendix 'Benchmark table' prints K2.6 = 96.4 for AIME 2026; both values kept as separate evidence rows.

打开官方来源

swebench-pro 58.6 模型 kimi-k2-6 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Table (DOM summary table) / Footnotes · row: SWE-Bench Pro

{
  "harness": "in-house SWE-agent-derived framework (bash/createfile/insert/view/strreplace/submit tools)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 10,
  "aggregation": "mean over 10 independent runs",
  "judge": null
}

swebench-pro id already introduced by prior batches, still not in data/benchmarks/. Competitor cells: K3 63.4, GPT-5.4 57.7, Opus 4.6 53.4, Gemini 3.1 Pro 54.2, Model A 51.0, Model B 49.5.

打开官方来源

terminalbench 66.7 模型 kimi-k2-6 · 版本 2.0, Terminus-2 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Table (DOM summary table) / Footnotes · row: Terminal-Bench 2.0

{
  "harness": "Terminus-2 with JSON parser, preserve thinking mode",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 10,
  "aggregation": "mean over 10 independent runs",
  "judge": null
}

Competitor cells: K3 71.8, GPT-5.4 65.4, Opus 4.6 65.4, Gemini 3.1 Pro 68.5, Model A 61.9, Model B 60.2.

打开官方来源

hlehle 54.0 模型 kimi-k2-6 · 版本 Full set, w/ tools (section) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: Humanity's Last Exam (Full) w/ tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Section heading confirms the w/tools condition is reported on the page; numeric value renders in the non-DOM benchmark component. Footnote 3a: search + code-interpreter + web-browsing tools; 262,144-token budget with 49,152 per step; retain-latest-round context management. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 xhigh 52.1, Claude Opus 4.6 (max) 53.0, Gemini 3.1 Pro (thinking high) 51.4, K2.5 50.2。

打开官方来源

browsecomp 83.2 模型 kimi-k2-6 · 版本 w/ tools (section) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Footnote: discard-all context management strategy, same as Kimi K2.5 and DeepSeek-V3.2. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 82.7, Claude Opus 4.6 83.7, Gemini 3.1 Pro 85.9, K2.5 74.9;agent swarm 行另列 K2.6 86.3 / K2.5 78.4。

打开官方来源

deepsearchqa 92.5 模型 kimi-k2-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: DeepSearchQA (f1-score)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

deepsearchqa id already introduced by prior batches, still not in data/benchmarks/. Metric is f1-score per section heading. Footnote: no context management for K2.6; over-context tasks counted as failed; competitor scores cited from Claude Opus 4.7 System Card. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:f1 行:GPT-5.4 78.6, Claude Opus 4.6 91.3, Gemini 3.1 Pro 81.9, K2.5 89.0;accuracy 行另列 K2.6 83.0。

打开官方来源

toolathlon 50.0 模型 kimi-k2-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

toolathlon id already introduced by prior batches, still not in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 54.6, Claude Opus 4.6 47.2, Gemini 3.1 Pro 48.8, K2.5 27.8。

打开官方来源

osworld 73.1 模型 kimi-k2-6 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OSWorld-Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Value in non-DOM component. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 75.0, Claude Opus 4.6 72.7, Gemini 3.1 Pro '-', K2.5 63.3。

打开官方来源

swebench-multilingual 76.7 模型 kimi-k2-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SWE-Multilingual

{
  "harness": "in-house SWE-agent-derived framework",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 10,
  "aggregation": "mean over 10 independent runs",
  "judge": null
}

new-benchmark: swebench-multilingual already introduced by prior batches, still not in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 '-', Claude Opus 4.6 77.8, Gemini 3.1 Pro 76.9*, K2.5 73.0。

打开官方来源

mathvision 93.2 模型 kimi-k2-6 · 版本 w/ python · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MathVision w/ python

{
  "harness": null,
  "tools": [
    "python"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "avg@3",
  "judge": null
}

mathvision id already introduced by prior batches, still not in data/benchmarks/. Footnote 5: python settings max-tokens-per-step 65,536, max-steps 50. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:w/ python 行:GPT-5.4 96.1*, Claude 84.6*, Gemini 95.7*, K2.5 85.0;无 python 行 K2.6 87.4。

打开官方来源

vstar 96.9 模型 kimi-k2-6 · 版本 w/ python · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Agents section heading · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: V* w/ python

{
  "harness": null,
  "tools": [
    "python"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "avg@3",
  "judge": null
}

new-benchmark: vstar (V*) not yet in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列:GPT-5.4 98.4*, Claude 86.4*, Gemini 96.9*, K2.5 86.9。

打开官方来源

hlehle 34.7 模型 kimi-k2-6 · 版本 Full set, no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HLE-Full

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "max generation 98,304; context 262,144",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Appendix no-tools full-set row. Competitor cells: GPT-5.4 (xhigh) 39.8, Claude Opus 4.6 (max effort) 40.0, Gemini 3.1 Pro (thinking high) 44.4, Kimi K2.5 30.1.

打开官方来源

hlehle 36.4 模型 kimi-k2-6 · 版本 Text-only subset, no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Footnotes / General Testing Details · row: Humanity's Last Exam (text-only subset) · quote_snippet: For the text-only subset, Kimi K2.6 achieves 36.4% accuracy without tools and 55.5% with tools.

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "max generation 98,304",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose-printed subset values - distinct construct from HLE-Full rows; never average text-only with full set.

打开官方来源

hlehle 55.5 模型 kimi-k2-6 · 版本 Text-only subset, w/ tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Footnotes / General Testing Details · row: Humanity's Last Exam (text-only subset) · quote_snippet: 36.4% accuracy without tools and 55.5% with tools

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "max generation 98,304",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose-printed subset values - distinct construct from HLE-Full rows; never merge.

打开官方来源

browsecomp 86.3 模型 kimi-k2-6 · 版本 Agent Swarm mode · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BrowseComp (agent swarm)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Swarm row - distinct from single-agent BrowseComp 83.2; only K2.6 has a value (all competitor cells '-'), Kimi K2.5 baseline 78.4.

打开官方来源

deepsearchqa 83.0 模型 kimi-k2-6 · 版本 accuracy · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: DeepSearchQA (accuracy)

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Accuracy companion metric - never merge with the f1-score row (K2.6 92.5). Competitor cells: GPT-5.4 63.7, Claude Opus 4.6 80.6, Gemini 3.1 Pro 60.2, Kimi K2.5 77.1.

打开官方来源

widesearch 80.8 模型 kimi-k2-6 · 版本 item-f1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: WideSearch (item-f1)

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

K2.6 value 80.8; all competitor columns '-'. The K2.5 baseline printed here (72.7) equals K2.5's single-agent base row on the K2.5 page, so this reads as the same single-agent protocol; swarm-vs-single is not labelled on this row.

打开官方来源

mcpmark 55.9 模型 kimi-k2-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MCPMark

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mcpmark not yet in data/benchmarks/. Competitor cells: GPT-5.4 62.5*, Claude Opus 4.6 56.7*, Gemini 3.1 Pro 55.9*, Kimi K2.5 29.5.

打开官方来源

claw-eval 62.3 模型 kimi-k2-6 · 版本 pass^3 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: Claw Eval (pass^3)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

pass^3 (all 3 attempts succeed) - distinct aggregation from pass@3 row; never merge. Competitor cells: GPT-5.4 60.3, Claude Opus 4.6 70.4, Gemini 3.1 Pro 57.8, Kimi K2.5 52.3.

打开官方来源

claw-eval 80.9 模型 kimi-k2-6 · 版本 pass@3 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: Claw Eval (pass@3)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

pass@3 (at least one of 3 succeeds) - distinct aggregation from pass^3 row; never merge. Competitor cells: GPT-5.4 78.4, Claude Opus 4.6 82.4, Gemini 3.1 Pro 82.9, Kimi K2.5 75.4.

打开官方来源

apex-agents 27.9 模型 kimi-k2-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: APEX-Agents

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote: evaluated on 452 of 480 public tasks. Competitor cells: GPT-5.4 33.3, Claude Opus 4.6 33.0, Gemini 3.1 Pro 32.0, Kimi K2.5 11.5.

打开官方来源

swebench 80.2 模型 kimi-k2-6 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SWE-Bench Verified

{
  "harness": "in-house framework adapted from SWE-agent (minimal tool set)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 10,
  "aggregation": "mean over 10 independent runs",
  "judge": null
}

Footnote 4: coding scores averaged over 10 independent runs; SWE series on in-house framework adapted from SWE-agent with bash/createfile/insert/view/strreplace/submit tools. Competitor cells: Claude Opus 4.6 80.8, Gemini 3.1 Pro 80.6, Kimi K2.5 76.8; GPT-5.4 not reported.

打开官方来源

scicode 52.2 模型 kimi-k2-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SciCode

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 10,
  "aggregation": "mean over 10 independent runs",
  "judge": null
}

Footnote 4: coding scores averaged over 10 independent runs; SWE series on in-house framework adapted from SWE-agent with bash/createfile/insert/view/strreplace/submit tools. Distinct from K2 Thinking's SciCode protocol (temp 0.0, no tools) - never merge. Competitor cells: GPT-5.4 56.6, Claude Opus 4.6 51.9, Gemini 3.1 Pro 58.9, Kimi K2.5 48.7.

打开官方来源

ojbench 60.6 模型 kimi-k2-6 · 版本 python · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OJBench (python)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 10,
  "aggregation": "mean over 10 independent runs",
  "judge": null
}

Footnote 4: coding scores averaged over 10 independent runs; SWE series on in-house framework adapted from SWE-agent with bash/createfile/insert/view/strreplace/submit tools. Python split - K2.5 page printed the cpp split (54.7 there); different language subset, never merge. Competitor cells: Claude Opus 4.6 60.3, Gemini 3.1 Pro 70.7, Kimi K2.5 54.7; GPT-5.4 not reported.

打开官方来源

lcb 89.6 模型 kimi-k2-6 · 版本 v6 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: LiveCodeBench (v6)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 10,
  "aggregation": "mean over 10 independent runs",
  "judge": null
}

Footnote 4: coding scores averaged over 10 independent runs; SWE series on in-house framework adapted from SWE-agent with bash/createfile/insert/view/strreplace/submit tools. Competitor cells: Claude Opus 4.6 88.8, Gemini 3.1 Pro 91.7, Kimi K2.5 85.0; GPT-5.4 not reported.

打开官方来源

aime-26 96.4 模型 kimi-k2-6 · 版本 Appendix 'Benchmark table' row · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Reasoning section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: AIME 2026

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Page-internal divergence: this appendix row prints K2.6 = 96.4 while the page's summary DOM table prints 93.3 for the same benchmark+model; the page prints both without reconciliation, so both are kept as separate evidence rows. Competitor cells: GPT-5.4 99.2, Claude Opus 4.6 96.7, Gemini 3.1 Pro 98.3, Kimi K2.5 95.8.

打开官方来源

hmmt-26 92.7 模型 kimi-k2-6 · 版本 February 2026 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Reasoning section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HMMT 2026 (Feb)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GPT-5.4 97.7, Claude Opus 4.6 96.2, Gemini 3.1 Pro 94.7, Kimi K2.5 87.1.

打开官方来源

imo-answerbench 86.0 模型 kimi-k2-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Reasoning section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: IMO-AnswerBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote: GPT-5.4 and Claude Opus 4.6 values cited from z.ai/blog/glm-5.1. Competitor cells: GPT-5.4 91.4 (cited), Claude Opus 4.6 75.3 (cited), Gemini 3.1 Pro 91.0*, Kimi K2.5 81.8.

打开官方来源

gpqa 90.5 模型 kimi-k2-6 · 版本 Diamond, Appendix 'Benchmark table' row · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Reasoning section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: GPQA-Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Page-internal divergence: this appendix row prints K2.6 = 90.5 while the page's summary DOM table prints 88.4 for the same benchmark+model; both kept as separate evidence rows. Competitor cells: GPT-5.4 92.8, Claude Opus 4.6 91.3, Gemini 3.1 Pro 94.3, Kimi K2.5 87.6.

打开官方来源

mmmu 79.4 模型 kimi-k2-6 · 版本 Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MMMU-Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 5: vision protocol max-tokens 98,304, avg@3. MMMU-Pro follows the official protocol (input order preserved, images prepended). Competitor cells: GPT-5.4 81.2, Claude Opus 4.6 73.9, Gemini 3.1 Pro 83.0*, Kimi K2.5 78.5.

打开官方来源

mmmu 80.1 模型 kimi-k2-6 · 版本 Pro, w/ python · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MMMU-Pro w/ python

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 5: vision protocol max-tokens 98,304, avg@3. Python-tool setting: max-tokens-per-step 65,536, max-steps 50 - never merge with no-python row. Competitor cells: GPT-5.4 82.1, Claude Opus 4.6 77.3, Gemini 3.1 Pro 85.3*, Kimi K2.5 77.7.

打开官方来源

charxiv-reasoning 80.4 模型 kimi-k2-6 · 版本 Reasoning query (RQ) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: CharXiv (RQ)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 5: vision protocol max-tokens 98,304, avg@3. Competitor cells: GPT-5.4 82.8*, Claude Opus 4.6 69.1, Gemini 3.1 Pro 80.2*, Kimi K2.5 77.5.

打开官方来源

charxiv-reasoning 86.7 模型 kimi-k2-6 · 版本 Reasoning query (RQ), w/ python · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: CharXiv (RQ) w/ python

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 5: vision protocol max-tokens 98,304, avg@3. Python-tool setting: max-tokens-per-step 65,536, max-steps 50 - never merge with no-python row. Competitor cells: GPT-5.4 90.0*, Claude Opus 4.6 84.7, Gemini 3.1 Pro 89.9*, Kimi K2.5 78.7.

打开官方来源

mathvision 87.4 模型 kimi-k2-6 · 版本 no python · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MathVision

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 5: vision protocol max-tokens 98,304, avg@3. Distinct from the w/ python row (93.2); never merge. Competitor cells: GPT-5.4 92.0*, Claude Opus 4.6 71.2*, Gemini 3.1 Pro 89.8*, Kimi K2.5 84.2.

打开官方来源

babyvision 39.8 模型 kimi-k2-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BabyVision

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 5: vision protocol max-tokens 98,304, avg@3. Competitor cells: GPT-5.4 49.7, Claude Opus 4.6 14.8, Gemini 3.1 Pro 51.6, Kimi K2.5 36.5.

打开官方来源

babyvision 68.5 模型 kimi-k2-6 · 版本 w/ python · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Agents section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BabyVision w/ python

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 5: vision protocol max-tokens 98,304, avg@3. Python-tool setting: max-tokens-per-step 65,536, max-steps 50 - never merge with no-python row. Competitor cells: GPT-5.4 80.2*, Claude Opus 4.6 38.4*, Gemini 3.1 Pro 68.3*, Kimi K2.5 40.5.

打开官方来源