← 模型目录

Kimi K2.5

Moonshot AI / Kimi · 2026-01-27 · 产品更新

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Kimi K2.5

发布文以“视觉智能体智能”为题,将 K2.5 定位为在 K2 基础上经 15T 视觉+文本混合语料继续预训练的原生多模态模型,Agent Swarm 最多 100 个子智能体/1500 步。45 项评测覆盖视觉理解、视频、智能体搜索与编码,亮点为 OCRBench 92.3 与 BrowseComp(Agent Swarm 模式)78.4。

输入模态
文本 / 图像 / 视频
上下文
256K
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 31.5 模型 kimi-k2-5 · 版本 Full set, text subset, no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · row: HLE · quote_snippet: Kimi K2.5 scores 31.5 (text) and 21.3 (image) without tools, and 51.8 (text) and 39.8 (image) with tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "96k completion budget; 256k context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

HuggingFace access blocked to prevent data leakage (footnote 2). Full-set default reporting; text and image subsets split out.

打开官方来源

hlehle 21.3 模型 kimi-k2-5 · 版本 Full set, image subset, no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · row: HLE · quote_snippet: Kimi K2.5 scores 31.5 (text) and 21.3 (image) without tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "96k completion budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Image-subset score; distinct construct from text subset - never average the two.

打开官方来源

hlehle 51.8 模型 kimi-k2-5 · 版本 Full set, text subset, w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · row: HLE · quote_snippet: 51.8 (text) and 39.8 (image) with tools

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "context management: retain only latest round of tool messages beyond threshold",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tools per footnote 3. Context management disclosed.

打开官方来源

hlehle 39.8 模型 kimi-k2-5 · 版本 Full set, image subset, w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · row: HLE · quote_snippet: 51.8 (text) and 39.8 (image) with tools

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Image-subset, tools-on.

打开官方来源

aime-25 96.1 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: AIME 2025

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 2 gives protocol: 96k completion budget, avg@32. Value in non-DOM appendix table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 xhigh 100, Claude 4.5 Opus (Extend) 92.8, Gemini 3 Pro (High) 95.0, DeepSeek V3.2 (Thinking) 93.1。

打开官方来源

hmmt25 95.4 模型 kimi-k2-5 · 版本 February 2025 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HMMT 2025 (Feb)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 32,
  "aggregation": "avg@32",
  "judge": null
}

hmmt-25 id already introduced by prior batches, still not in data/benchmarks/. Value in non-DOM table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 99.4, Claude 4.5 Opus 92.9*, Gemini 3 Pro 97.3*, DeepSeek V3.2 92.5。

打开官方来源

gpqa 87.6 模型 kimi-k2-5 · 版本 Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: gpqa-diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 8,
  "aggregation": "avg@8",
  "judge": null
}

Value in non-DOM table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 92.4, Claude 4.5 Opus 87, Gemini 3 Pro 91.9, DeepSeek V3.2 82.4。

打开官方来源

imo-answerbench 81.8 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: imo-answerbench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

imo-answerbench id already introduced by prior batches. Value in non-DOM table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 86.3, Claude 4.5 Opus 78.5*, Gemini 3 Pro 83.1*, DeepSeek V3.2 78.3。

打开官方来源

swebench 76.8 模型 kimi-k2-5 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: swebench-verified

{
  "harness": "internally developed SWE-series framework; highest scores under non-thinking mode",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Footnote 5: SWE series run on in-house framework with minimal tool set; non-thinking mode achieved highest scores (inverted vs expectations - disclosed). Token-cost chart shows strong performance at fraction of cost vs competitors. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 80.0, Claude 4.5 Opus 80.9, Gemini 3 Pro 76.2, DeepSeek V3.2 73.1。

打开官方来源

swebench-multilingual 73.0 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: swebench-multilingual

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

new-benchmark: swebench-multilingual already introduced by prior batches, still not in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 72.0, Claude 4.5 Opus 77.5, Gemini 3 Pro 65.0, DeepSeek V3.2 70.2。

打开官方来源

terminalbench 50.8 模型 kimi-k2-5 · 版本 2.0, non-thinking mode · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: terminalbench-2-0

{
  "harness": "Terminus-2 default agent framework + provided JSON parser",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 5: evaluated under NON-thinking mode because the thinking-mode context management is incompatible with Terminus-2 - protocol deviation disclosed by vendor. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 54.0, Claude 4.5 Opus 59.3, Gemini 3 Pro 54.2, DeepSeek V3.2 46.4。

打开官方来源

seal-0 57.4 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: seal-0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 4,
  "aggregation": "avg@4",
  "judge": null
}

new-benchmark: seal-0 (also in kimi-k2-thinking.json). Value in non-DOM table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 45.0, Claude 4.5 Opus 47.7*, Gemini 3 Pro 45.5*, DeepSeek V3.2 49.5*。

打开官方来源

widesearch 72.7 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: widesearch

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 4,
  "aggregation": "avg@4",
  "judge": null
}

new-benchmark: widesearch not yet in data/benchmarks/. Appears in Agent Swarm chart vs Claude Opus 4.5 and footnotes. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。Claude 4.5 Opus 76.2*, Gemini 3 Pro 57.0, DeepSeek V3.2 32.5*;K2.5 Agent Swarm 模式另列 79.0。

打开官方来源

mmmu 78.5 模型 kimi-k2-5 · 版本 Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: mmmu-pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "avg@3",
  "judge": null
}

Footnote 4: official protocol, input order preserved, images prepended. Vision max-tokens 64k. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 79.5, Claude 4.5 Opus 74.0, Gemini 3 Pro 81.0, Qwen3-VL-235B-A22B 69.3。

打开官方来源

mathvision 84.2 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Image section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: mathvision

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "avg@3",
  "judge": null
}

mathvision id already introduced by prior batches. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 83, Claude 4.5 Opus 77.1*, Gemini 3 Pro 86.1, Qwen3-VL 74.6。

打开官方来源

zerobench 11 模型 kimi-k2-5 · 版本 w/ tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: zerobench-w-tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "max-tokens-per-step 24k, max-steps 30",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

zerobench id already introduced by prior batches. Multi-step tool protocol from footnote 4. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。同表 ZeroBench 无工具 K2.5=9;ZeroBench w/ tools 行 GPT-5.2 7*, Claude 9*, Gemini 12*。

打开官方来源

omnidocbench 88.8 模型 kimi-k2-5 · 版本 1.5 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Image section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: omnidocbench-1-5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: omnidocbench not yet in data/benchmarks/. Score = (1 - normalized Levenshtein distance) x 100 per footnote 4. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 85.7, Claude 4.5 Opus 87.7*, Gemini 3 Pro 88.5, Qwen3-VL 82.0*。

打开官方来源

video-mmmu 86.6 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Video section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: videommmu

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

video-mmmu id already introduced by prior batches. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 85.9, Claude 4.5 Opus 84.4*, Gemini 3 Pro 87.6, Qwen3-VL 80.0。

打开官方来源

longvideobench 79.8 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Video section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: longvideobench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: longvideobench not yet in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 76.5, Claude 4.5 Opus 67.2, Gemini 3 Pro 77.7*, Qwen3-VL 65.6*。

打开官方来源

worldvqa 46.3 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: worldvqa

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: worldvqa not yet in data/benchmarks/ (Kimi-published, github.com/MoonshotAI/WorldVQA). Atomic vision-centric world knowledge. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 28.0, Claude 4.5 Opus 36.8, Gemini 3 Pro 47.4, Qwen3-VL 23.5。

打开官方来源

aa-lcr 70.0 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: aa-lcr

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "avg@3",
  "judge": null
}

new-benchmark: aa-lcr (Artificial Analysis long-context reasoning) not yet in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 72.3*, Claude 4.5 Opus 71.3*, Gemini 3 Pro 65.3*, DeepSeek V3.2 64.3*。

打开官方来源

longbench 61.0 模型 kimi-k2-5 · 版本 v2, ~128k standardized inputs · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: longbench-v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

longbench id exists in data/benchmarks/ (v1); v2 variant with standardized ~128k inputs per footnote 6. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 54.5*, Claude 4.5 Opus 64.4*, Gemini 3 Pro 68.2*, DeepSeek V3.2 59.8*。

打开官方来源

browsecomp 78.4 模型 kimi-k2-5 · 版本 Agent Swarm mode · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: browsecomp-swarm · figure: Chart image https://statics.kimi.ai/blogs/k2-5/20260127-131347.jpeg (K2.5 Agent Swarm vs Claude Opus 4.5 on BrowseComp, Wide Search, In-house Bench)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": "main agent 15 steps; sub-agents 100 steps",
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp already introduced by prior batches. Swarm-mode protocol differs from single-agent BrowseComp rows everywhere else - never merge. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。附录同 benchmark 另有两行:默认 60.6、w/ctx mgm 74.9;Agent Swarm 行仅 K2.5 有值。

打开官方来源

ai-office-bench 官方未公布数值 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Office Productivity section · row: ai-office-bench · quote_snippet: K2.5 shows 59.3% and 24.3% improvements over K2 Thinking

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ai-office-bench (internal) not yet in data/benchmarks/. Relative-improvement claim only (59.3% on AI Office Benchmark, 24.3% on General Agent Benchmark vs K2 Thinking); the companion image alt text reads 71.2%/39.0% - conflicting with prose, prose taken as authoritative, discrepancy flagged for the technical report (arXiv 2602.02276).

打开官方来源

hlehle 30.1 模型 kimi-k2-5 · 版本 Full set (text & image), no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HLE-Full

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "96k completion budget; 256k context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Appendix full-set row - distinct from the four footnote-2 subset rows (text 31.5 / image 21.3); never average or merge. DeepSeek V3.2 cell 25.1* is its text-only subset (footnote 2 dagger); HuggingFace access blocked to prevent leakage. Competitor cells: GPT-5.2 (xhigh) 34.5, Claude 4.5 Opus (Extend Thinking) 30.8, Gemini 3 Pro (High) 37.5, Qwen3-VL not reported.

打开官方来源

hlehle 50.2 模型 kimi-k2-5 · 版本 Full set (text & image), w/ tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HLE-Full w/ tools

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "context management: once context exceeds threshold, only latest round of tool messages retained",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Appendix full-set w/ tools row - distinct from footnote-2 subset rows (text 51.8 / image 39.8); never merge. DeepSeek V3.2 cell 40.8* is its text-only subset (dagger). Competitor cells: GPT-5.2 45.5, Claude 4.5 Opus 43.2, Gemini 3 Pro 45.8, Qwen3-VL not reported.

打开官方来源

mmlu-pro 87.1 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MMLU-Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Asterisked competitor cells re-tested by Kimi. Competitor cells: GPT-5.2 86.7*, Claude 4.5 Opus 89.3*, Gemini 3 Pro 90.1, DeepSeek V3.2 85.0.

打开官方来源

charxiv-reasoning 77.5 模型 kimi-k2-5 · 版本 Reasoning query (RQ) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: CharXiv (RQ)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 4: vision protocol max-tokens 64k, avg@3. Asterisked cells re-tested by Kimi. Competitor cells: GPT-5.2 82.1, Claude 4.5 Opus 67.2*, Gemini 3 Pro 81.4, Qwen3-VL 66.1; DeepSeek V3.2 not reported (text-only model on vision rows).

打开官方来源

mathvista 90.1 模型 kimi-k2-5 · 版本 mini · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MathVista (mini)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 82.8*, Claude 4.5 Opus 80.2*, Gemini 3 Pro 89.8*, Qwen3-VL 85.8; DeepSeek V3.2 not reported.

打开官方来源

zerobench 9 模型 kimi-k2-5 · 版本 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: ZeroBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 4: vision protocol max-tokens 64k, avg@3. No-tools row - distinct from ZeroBench (w/ tools) 11 row; never merge. Competitor cells: GPT-5.2 9*, Claude 4.5 Opus 3*, Gemini 3 Pro 8*, Qwen3-VL 4*; DeepSeek V3.2 not reported.

打开官方来源

ocrbench 92.3 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OCRBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ocrbench not yet in data/benchmarks/. Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 80.7*, Claude 4.5 Opus 86.5*, Gemini 3 Pro 90.3*, Qwen3-VL 87.5; DeepSeek V3.2 not reported.

打开官方来源

infovqa 92.6 模型 kimi-k2-5 · 版本 test · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: InfoVQA (test)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: infovqa not yet in data/benchmarks/. Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 84*, Claude 4.5 Opus 76.9*, Gemini 3 Pro 57.2*, Qwen3-VL 89.5; DeepSeek V3.2 not reported.

打开官方来源

simplevqa 71.2 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SimpleVQA

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 55.8*, Claude 4.5 Opus 69.7*, Gemini 3 Pro 69.7*, Qwen3-VL 56.8*; DeepSeek V3.2 not reported.

打开官方来源

mmvu 80.4 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MMVU

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 80.8*, Claude 4.5 Opus 77.3, Gemini 3 Pro 77.5, Qwen3-VL 71.1; DeepSeek V3.2 not reported.

打开官方来源

motionbench 70.4 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MotionBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 64.8, Claude 4.5 Opus 60.3, Gemini 3 Pro 70.3; DeepSeek V3.2 and Qwen3-VL not reported.

打开官方来源

video-mme 87.4 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: VideoMME

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 86.0, Gemini 3 Pro 88.4*, Qwen3-VL 79.0; Claude 4.5 Opus and DeepSeek V3.2 not reported.

打开官方来源

lvbench 75.9 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: LVBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: Gemini 3 Pro 73.5*, Qwen3-VL 63.6; GPT-5.2, Claude 4.5 Opus and DeepSeek V3.2 not reported.

打开官方来源

swebench-pro 50.7 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SWE-Bench Pro

{
  "harness": "in-house SWE-series framework (minimal tool set), non-thinking mode scored highest",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Competitor cells: GPT-5.2 55.6, Claude 4.5 Opus 55.4*; Gemini 3 Pro, DeepSeek V3.2 and Qwen3-VL not reported.

打开官方来源

paperbench 63.5 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: PaperBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Competitor cells: GPT-5.2 63.7*, Claude 4.5 Opus 72.9*, DeepSeek V3.2 47.1; Gemini 3 Pro and Qwen3-VL not reported.

打开官方来源

cybergym 41.3 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: CyberGym

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Claude Opus 4.5 cell 50.6 reported by Kimi under the non-thinking setting (footnote 5 disclosure). Competitor cells: Gemini 3 Pro 39.9*, DeepSeek V3.2 17.3*; GPT-5.2 and Qwen3-VL not reported.

打开官方来源

scicode 48.7 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SciCode

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Distinct from K2 Thinking's SciCode run (temp 0.0, no tools) - different protocol, never merge. Competitor cells: GPT-5.2 52.1, Claude 4.5 Opus 49.5, Gemini 3 Pro 56.1, DeepSeek V3.2 38.9; Qwen3-VL not reported.

打开官方来源

ojbench 57.4 模型 kimi-k2-5 · 版本 cpp · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OJBench (cpp)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Competitor cells: Claude 4.5 Opus 54.6*, Gemini 3 Pro 68.5*, DeepSeek V3.2 54.7*; GPT-5.2 and Qwen3-VL not reported.

打开官方来源

lcb 85.0 模型 kimi-k2-5 · 版本 v6 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: LiveCodeBench (v6)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Competitor cells: Claude 4.5 Opus 82.2*, Gemini 3 Pro 87.4*, DeepSeek V3.2 83.3; GPT-5.2 and Qwen3-VL not reported.

打开官方来源

browsecomp 60.6 模型 kimi-k2-5 · 版本 single-agent, discard-all context strategy · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BrowseComp

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 3: agentic-search tool set (search, code-interpreter, web-browsing); for BrowseComp K2.5 and DeepSeek V3.2 used the discard-all context strategy. One of three BrowseComp protocols on this page (default / w/ ctx mgm / Agent Swarm) - never merge. Competitor cells: Claude 4.5 Opus 37.0, Gemini 3 Pro 37.8, DeepSeek V3.2 51.4; GPT-5.2 and Qwen3-VL not reported.

打开官方来源

browsecomp 74.9 模型 kimi-k2-5 · 版本 single-agent, w/ context management · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BrowseComp (w/ctx mgm)

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

One of three BrowseComp protocols on this page; w/ context management row is a different protocol from the default 60.6 and Agent Swarm 78.4 rows - never merge. Competitor cells: GPT-5.2 65.8, Claude 4.5 Opus 57.8, Gemini 3 Pro 59.2, DeepSeek V3.2 67.6; Qwen3-VL not reported.

打开官方来源

widesearch 79.0 模型 kimi-k2-5 · 版本 Agent Swarm mode · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: WideSearch (item-f1) (Agent Swarm)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": "main and sub-agents max 100 steps",
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: widesearch registered. Footnote 7: WideSearch Swarm Mode main and sub-agents max 100 steps. Swarm row - distinct from single-agent WideSearch 72.7; only K2.5 has a value on this row (all competitor cells '-').

打开官方来源

deepsearchqa 77.1 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: DeepSearchQA

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 3 tool set. Competitor cells: GPT-5.2 71.3*, Claude 4.5 Opus 76.1*, Gemini 3 Pro 63.2*, DeepSeek V3.2 60.9*; Qwen3-VL not reported.

打开官方来源

finsearchcomp 67.8 模型 kimi-k2-5 · 版本 T2 & T3 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: FinSearchCompT2&T3

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote 3 tool set. Combined T2&T3 split - distinct from K2 Thinking's T3-only row. Competitor cells: Claude 4.5 Opus 66.2*, Gemini 3 Pro 49.9, DeepSeek V3.2 59.1*; GPT-5.2 and Qwen3-VL not reported.

打开官方来源

osworld 63.3 模型 kimi-k2-5 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OSWorld-Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Computer-use section. Competitor cells: GPT-5.2 8.6*, Claude 4.5 Opus 66.3, Gemini 3 Pro 20.7*, Qwen3-VL 38.1; DeepSeek V3.2 not reported.

打开官方来源

webarena 58.9 模型 kimi-k2-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: WebArena

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Computer-use section. Competitor cells: Claude 4.5 Opus 63.4*, Qwen3-VL 26.4*; GPT-5.2, Gemini 3 Pro and DeepSeek V3.2 not reported.

打开官方来源