Kimi K3
Moonshot AI / Kimi · 2026-07(仅精确到月) · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Kimi K3
Kimi K3 以「Open Frontier Intelligence」为定位发布(权重承诺 2026-07-27 前开源),评测统一 reasoning effort=max、temperature=1.0。已收录 35 项以上评测覆盖编码、Agent、搜索与多模态:亮点 Terminal-Bench 2.1 88.3、BrowseComp(1M 上下文)90.4。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Coding benchmarks · row: DeepSWE · quote_snippet: Kimi K3 attains 67.3 with the mini-SWE-agent harness
{
"harness": "mini-SWE-agent (official DeepSWE leaderboard)",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote states K3 is evaluated with the Kimi Code harness, and that the printed DeepSWE score comes from the official DeepSWE leaderboard (https://deepswe.datacurve.ai/) under mini-SWE-agent. DeepSWE v1.1 tasks. Global footnote: reasoning effort=max, temperature=1.0, top-p=1.0. Notes field new-benchmark: deepswe is not yet in data/benchmarks.json.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Coding benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: DeepSWE · figure: Benchmark comparison chart image 'Kimi K3 benchmark comparison' (coding section, thinking effort max/xhigh), kimi-file.kimi.ai asset 1d9chl6mdcmosb3rnlehg; the 'Full Benchmark Table' heading has no DOM table - values render as images
{
"harness": "Kimi Code",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deepswe not yet in data/benchmarks/. Score confirmed machine-readably (2026-09-01): the appendix 'Full Benchmark Table' embedded JSON in the archived page prints 67.5 for Kimi K3 (Kimi Code), matching the earlier vision read of the chart. Distinct protocol from the 67.3 mini-SWE-agent leaderboard row (footnote) - never merge. Different protocol from the 67.3 mini-SWE-agent leaderboard value in the footnote; do not merge. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):70.0 / 73.0 / 59.0 / 67.0 / 46.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Footnotes / Coding benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: Terminal-Bench 2.1 · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading; section contains no DOM table · quote_snippet: Terminal-Bench 2.1. Kimi K3 is evaluated with the Kimi Code harness.
{
"harness": "Kimi Code",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Protocol verified from footnote; score value is in image only. Footnote also states competitor rows report best-across-harness scores (GLM-5.2 Claude Code; Opus 4.8 / Fable 5 Terminus 2; GPT 5.5 / GPT 5.6 Sol Codex) - competitor cells are not same-protocol. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):84.6 / 88.8 / 84.6 / 83.4 / 82.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Coding benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: Program Bench · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: Program Bench. Kimi K3 is evaluated with the Kimi Code harness.
{
"harness": "Kimi Code",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: program-bench (Program Bench / ProgramBench) not yet in data/benchmarks.json. Competitor scores cited from vals.ai ProgramBench leaderboard per footnote. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):76.8 / 77.6 / 71.9 / 70.8 / 63.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Coding benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: SWE Marathon · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: H20-calibrated branch of the official v1.1 tasks
{
"harness": "Claude Code",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swe-marathon (swe-marathon.org) not yet in data/benchmarks.json. Footnote discloses Docker images, performance gates and reference oracles for GPU tasks recalibrated for H20 hardware; variant deviates from official v1.1. GPT-5.6 Sol column uses Codex harness. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):35.0 / 39.0 / 40.0 / 14.0 / 13.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Coding benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: FrontierSWE · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: Dominance scores are recomputed from the raw scores
{
"harness": "Kimi Code",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: frontierswe (frontierswe.com) not yet in data/benchmarks.json. Metric is a dominance score recomputed via the official evaluation script, current as of July 16, 2026. Other-model rows cited from frontierswe.com. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):86.6 / 71.3 / 66.7 / 64.9 / 67.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Coding benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: PostTrain Bench · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: official Harbor implementation at maximum reasoning effort, averaged over three runs on H20
{
"harness": "Claude Code (official Harbor implementation)",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "mean over 3 runs",
"judge": null
}new-benchmark: posttrain-bench (posttrainbench.com) not yet in data/benchmarks.json. Footnote: run on H20 GPU instead of H100 used in the official setting; GLM-5.2, GPT-5.5 and Opus 4.8 rows adopted from official PostTrainBench results. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):41.4 / 34.6 / 34.1 / 28.4 / 34.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Coding benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: MLS Bench Lite · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: MLS Bench Lite. Kimi K3 is evaluated with the Kimi Code harness
{
"harness": "Kimi Code",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mls-bench-lite not yet in data/benchmarks.json. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):49.9 / 46.2 / 42.8 / 35.5 / 40.4。 附录表该行名为 'MLS Bench'(JSON 行记为 Lite,附注一致对待)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Coding benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: KCB 2.0 · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading; chart axis label reads 'Kimi Code Bench 2.0 (Internal)' · quote_snippet: on this in-house benchmark, 10% of the tasks entered GPT-5.6 Sol's cyber guard
{
"harness": "Kimi Code and Claude Code (both reported for K3)",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: kcb (KCB / Kimi Code Bench, Kimi in-house) not yet in data/benchmarks.json. Footnote: GPT-5.5 uses the 'xhigh' effort setting on this row; competitor rows are not same-effort. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):76.9 / 64.8 / 71.7 / 69.0 / 64.2。 附录表该行注释为空,K3 仅一列 72.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: OfficeQA Pro · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: all PDFs rendered as images and no machine-readable text available
{
"harness": "Claude Code",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: officeqa-pro not yet in data/benchmarks.json. Footnote: each test case provides the entire PDF corpus rendered as images with no machine-readable text. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):69.9* / 63.2* / 63.9* / 60.9* / 41.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: SpreadsheetBench 2 · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: OfficeQA Pro and SpreadsheetBench 2 ... evaluated with the Claude Code harness
{
"harness": "Claude Code",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: spreadsheetbench not yet in data/benchmarks.json. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):34.7* / 32.4* / 31.55* / 29.05* / 28.12。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: MCP Atlas · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: 500-task public subset with a 100-turn limit, using Gemini 3.1 Pro as the judge
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": 100,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Gemini 3.1 Pro"
}new-benchmark: mcp-atlas not yet in data/benchmarks.json. LLM-judged protocol disclosed in footnote. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):84.7 / 83.6 / 83.6 / 82.8 / 82.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: AutomationBench · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: 600-task public subset, following the official GitHub setup
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: automationbench not yet in data/benchmarks.json. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):29.1 / 29.7 / 27.2 / 22.7 / 12.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: BrowseComp · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: context-compaction strategy used in the Claude model cards, triggered at 300K tokens
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": "context compaction triggered at 300K tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks.json. Main-table variant; the un-compacted 1M-context condition is a separate evidence row. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):88.0 / 90.4 / 84.3 / 84.4 / '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · row: BrowseComp · quote_snippet: 1M-token context window and no context management, Kimi K3 achieves a score of 90.4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": "1M-token context window, no context management",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks/. Alternative-protocol disclosure printed in the footnote, distinct from the main-table 300K-compaction variant; the two rows must not be merged or compared across variants.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: GDPval-AA · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: GDPval-AA, AA-Briefcase, and APEX-Agents scores are cited from artificialanalysis.ai
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: gdpval-aa not yet in data/benchmarks.json. Footnote states scores (including Kimi K3's own row) are cited from Artificial Analysis, not vendor-run. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):1760 / 1748 / 1600 / 1494 / 1514。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: AA-Briefcase · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: GDPval-AA, AA-Briefcase, and APEX-Agents scores are cited from artificialanalysis.ai
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aa-briefcase not yet in data/benchmarks.json. Scores cited from Artificial Analysis per footnote. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):1583 / 1495 / 1354 / 1158 / 1260。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: APEX-Agents · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: GDPval-AA, AA-Briefcase, and APEX-Agents scores are cited from artificialanalysis.ai
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: apex-agents not yet in data/benchmarks.json. Scores cited from Artificial Analysis per footnote. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):43.3 / 39.9 / 39.4 / 38.5 / 35.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Multimodal benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: ZeroBench · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: Except for ZeroBench, which follows the official setting and is run five times
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": null,
"judge": null
}new-benchmark: zerobench not yet in data/benchmarks.json. Only multimodal row run five times; all other multimodal rows averaged over three runs. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):23.0 / 17.0 / 17.0 / 22.0 / '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Multimodal benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: MMMU-Pro · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: following the official protocol, preserving the original input order and prepending images
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "mean over 3 runs",
"judge": null
}Mapped to existing benchmark mmmu with variant Pro; MMMU-Pro is a derived variant with its own protocol (vision-only input). Footnote discloses official protocol with images prepended to text input. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):81.2 / 83.0 / 78.9 / 81.2 / '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Multimodal benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: PerceptionBench · figure: Benchmark comparison chart images under 'Full Benchmark Table' heading · quote_snippet: PerceptionBench ... focuses on atomic visual perception capabilities
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "mean over 3 runs",
"judge": null
}new-benchmark: perception-bench (Kimi-published benchmark, kimi.com blog) not yet in data/benchmarks.json. Vendor-owned benchmark; treat as in-house like KCB. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取,与先前视觉读数一致)。同表竞品列(Fable5 / GPT-5.6 Sol / Opus4.8 / GPT-5.5 / GLM-5.2):57.2 / 59.7 / 47.2 / 55.8 / '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: DeepSearchQA (f1-score)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):94.2 / not reported / 93.1 / not reported / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: Toolathlon-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):77.9 / 74.9 / 76.2 / 73.5 / 59.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: Job Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: job-bench not yet in data/benchmarks/. 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):57.4 / 46.5 / 48.4 / 38.3 / 43.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: DECK-Bench (Internal)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deck-bench not yet in data/benchmarks/. Kimi-internal deck/slides benchmark. 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):73.0 / 74.7 / 66.9 / 68.2 / 68.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Reasoning & Knowledge benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: GPQA-Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):92.6 / 94.1 / 91.0 / 93.5 / 91.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Reasoning & Knowledge benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: HLE-Full
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": "per global footnote; HLE full-set default reporting",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Full set (text & image) no-tools row - distinct from the w/ tools row and from K2.6's text-only subset rows; never merge. 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):53.3 / 44.5 / 49.8* / 41.4* / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Reasoning & Knowledge benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: HLE-Full w/ tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Full set (text & image) w/ tools row - never merge with no-tools 43.5. 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):63.0 / 58.0 / 57.9* / 52.2* / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Vision benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: MMMU-Pro w/ python
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Python-tool setting - never merge with no-python 81.6. 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):86.5 / 84.6 / 82.7 / 83.2 / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Vision benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: CharXiv (RQ)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):88.9 / 84.6 / 80.5 / 84.1 / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Vision benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: CharXiv (RQ) w/ python
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Python-tool setting - never merge with no-python 84.8. 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):93.5 / 89.1 / 89.9 / 89.0 / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Vision benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: MathVision
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):94.8 / 95.8 / 86.7 / 92.2 / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Vision benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: MathVision w/ python
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Python-tool setting - never merge with no-python 94.3. 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):98.6 / 97.8 / 97.1 / 96.8 / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Vision benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: BabyVision w/ python
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):90.5 / 88.9 / 81.2 / 83.6 / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Vision benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: ZeroBench_main w/ python (pass@5)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@5",
"judge": null
}pass@5 aggregation - never merge with pass@1-style rows or the no-python 23.0 row. 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):46.0 / 35.0 / 34.0 / 41.0 / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Vision benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: WorldVQA ForceAnswer
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}ForceAnswer condition - distinct from WorldVQA rows in other releases; never merge. 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):56.7 / 41.8 / 39.1 / 38.5 / not reported。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table / Vision benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: OmniDocBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Version not printed on this page (K2.5 page printed 1.5). 同表竞品列(Fable 5 / GPT-5.6 Sol / Opus 4.8 / GPT-5.5 / GLM-5.2):89.8 / 85.8 / 87.9 / 89.4 / not reported。
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: BrowseComp · quote_snippet: results of Claude Fable 5 ... are cited from anthropic.com and openai.com
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks/. Competitor score transcribed by Kimi; footnote cites https://www.anthropic.com/news/claude-fable-5-mythos-5 as the origin. Score value appears only in image; verify against the Anthropic page before use. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(2026-09-01 提取)。本行为 K3 表中 Fable 5 列值 该值仍为 Kimi 转引口径(comparison_cited),如需对外引用请回溯厂商原文。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: BrowseComp · quote_snippet: results of ... Claude Opus 4.8 ... are cited from anthropic.com and openai.com
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks/. Competitor score; origin cited as the Anthropic Fable 5 / Mythos 5 announcement page per footnote. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(2026-09-01 提取)。本行为表中 Opus 4.8 列值 该值仍为 Kimi 转引口径(comparison_cited),如需对外引用请回溯厂商原文。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: BrowseComp · quote_snippet: results of ... GPT 5.6 Sol ... are cited from anthropic.com and openai.com
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks/. Competitor score; footnote cites https://openai.com/index/gpt-5-6/ as origin. GPT-5.6 page reports 92.2% for BrowseComp (see openai/gpt-5-6.json); cross-check before merging. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(2026-09-01 提取)。本行为表中 GPT 5.6 Sol 列值 该值仍为 Kimi 转引口径(comparison_cited),如需对外引用请回溯厂商原文。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Footnotes / Productivity and agentic benchmarks · table: Appendix 'Full Benchmark Table'(页面内嵌 JSON 数据,归档 page.html payload 机器可读;正文以图表图片渲染) · row: BrowseComp · quote_snippet: results of ... GPT 5.5 are cited from anthropic.com and openai.com
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks/. Competitor score; footnote cites the OpenAI GPT-5.6 page as origin. 数值取自归档 page.html 附录 Full Benchmark Table 内嵌 JSON(2026-09-01 提取)。本行为表中 GPT 5.5 列值 该值仍为 Kimi 转引口径(comparison_cited),如需对外引用请回溯厂商原文。