Kimi K2.5
Moonshot AI / Kimi · 2026-01-27 · 产品更新
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Kimi K2.5
发布文以“视觉智能体智能”为题,将 K2.5 定位为在 K2 基础上经 15T 视觉+文本混合语料继续预训练的原生多模态模型,Agent Swarm 最多 100 个子智能体/1500 步。45 项评测覆盖视觉理解、视频、智能体搜索与编码,亮点为 OCRBench 92.3 与 BrowseComp(Agent Swarm 模式)78.4。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 256K
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · row: HLE · quote_snippet: Kimi K2.5 scores 31.5 (text) and 21.3 (image) without tools, and 51.8 (text) and 39.8 (image) with tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "96k completion budget; 256k context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}HuggingFace access blocked to prevent data leakage (footnote 2). Full-set default reporting; text and image subsets split out.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · row: HLE · quote_snippet: Kimi K2.5 scores 31.5 (text) and 21.3 (image) without tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "96k completion budget",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Image-subset score; distinct construct from text subset - never average the two.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · row: HLE · quote_snippet: 51.8 (text) and 39.8 (image) with tools
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "context management: retain only latest round of tool messages beyond threshold",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Tools per footnote 3. Context management disclosed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · row: HLE · quote_snippet: 51.8 (text) and 39.8 (image) with tools
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Image-subset, tools-on.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: AIME 2025
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 2 gives protocol: 96k completion budget, avg@32. Value in non-DOM appendix table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 xhigh 100, Claude 4.5 Opus (Extend) 92.8, Gemini 3 Pro (High) 95.0, DeepSeek V3.2 (Thinking) 93.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HMMT 2025 (Feb)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 32,
"aggregation": "avg@32",
"judge": null
}hmmt-25 id already introduced by prior batches, still not in data/benchmarks/. Value in non-DOM table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 99.4, Claude 4.5 Opus 92.9*, Gemini 3 Pro 97.3*, DeepSeek V3.2 92.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: gpqa-diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 8,
"aggregation": "avg@8",
"judge": null
}Value in non-DOM table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 92.4, Claude 4.5 Opus 87, Gemini 3 Pro 91.9, DeepSeek V3.2 82.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: imo-answerbench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}imo-answerbench id already introduced by prior batches. Value in non-DOM table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 86.3, Claude 4.5 Opus 78.5*, Gemini 3 Pro 83.1*, DeepSeek V3.2 78.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: swebench-verified
{
"harness": "internally developed SWE-series framework; highest scores under non-thinking mode",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Footnote 5: SWE series run on in-house framework with minimal tool set; non-thinking mode achieved highest scores (inverted vs expectations - disclosed). Token-cost chart shows strong performance at fraction of cost vs competitors. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 80.0, Claude 4.5 Opus 80.9, Gemini 3 Pro 76.2, DeepSeek V3.2 73.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: swebench-multilingual
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}new-benchmark: swebench-multilingual already introduced by prior batches, still not in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 72.0, Claude 4.5 Opus 77.5, Gemini 3 Pro 65.0, DeepSeek V3.2 70.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: terminalbench-2-0
{
"harness": "Terminus-2 default agent framework + provided JSON parser",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 5: evaluated under NON-thinking mode because the thinking-mode context management is incompatible with Terminus-2 - protocol deviation disclosed by vendor. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 54.0, Claude 4.5 Opus 59.3, Gemini 3 Pro 54.2, DeepSeek V3.2 46.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: seal-0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 4,
"aggregation": "avg@4",
"judge": null
}new-benchmark: seal-0 (also in kimi-k2-thinking.json). Value in non-DOM table. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 45.0, Claude 4.5 Opus 47.7*, Gemini 3 Pro 45.5*, DeepSeek V3.2 49.5*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: widesearch
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 4,
"aggregation": "avg@4",
"judge": null
}new-benchmark: widesearch not yet in data/benchmarks/. Appears in Agent Swarm chart vs Claude Opus 4.5 and footnotes. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。Claude 4.5 Opus 76.2*, Gemini 3 Pro 57.0, DeepSeek V3.2 32.5*;K2.5 Agent Swarm 模式另列 79.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: mmmu-pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "avg@3",
"judge": null
}Footnote 4: official protocol, input order preserved, images prepended. Vision max-tokens 64k. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 79.5, Claude 4.5 Opus 74.0, Gemini 3 Pro 81.0, Qwen3-VL-235B-A22B 69.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Image section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: mathvision
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "avg@3",
"judge": null
}mathvision id already introduced by prior batches. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 83, Claude 4.5 Opus 77.1*, Gemini 3 Pro 86.1, Qwen3-VL 74.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: zerobench-w-tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "max-tokens-per-step 24k, max-steps 30",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}zerobench id already introduced by prior batches. Multi-step tool protocol from footnote 4. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。同表 ZeroBench 无工具 K2.5=9;ZeroBench w/ tools 行 GPT-5.2 7*, Claude 9*, Gemini 12*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Image section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: omnidocbench-1-5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: omnidocbench not yet in data/benchmarks/. Score = (1 - normalized Levenshtein distance) x 100 per footnote 4. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 85.7, Claude 4.5 Opus 87.7*, Gemini 3 Pro 88.5, Qwen3-VL 82.0*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Video section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: videommmu
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}video-mmmu id already introduced by prior batches. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 85.9, Claude 4.5 Opus 84.4*, Gemini 3 Pro 87.6, Qwen3-VL 80.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Video section · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: longvideobench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: longvideobench not yet in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 76.5, Claude 4.5 Opus 67.2, Gemini 3 Pro 77.7*, Qwen3-VL 65.6*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: worldvqa
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: worldvqa not yet in data/benchmarks/ (Kimi-published, github.com/MoonshotAI/WorldVQA). Atomic vision-centric world knowledge. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 28.0, Claude 4.5 Opus 36.8, Gemini 3 Pro 47.4, Qwen3-VL 23.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: aa-lcr
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "avg@3",
"judge": null
}new-benchmark: aa-lcr (Artificial Analysis long-context reasoning) not yet in data/benchmarks/. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 72.3*, Claude 4.5 Opus 71.3*, Gemini 3 Pro 65.3*, DeepSeek V3.2 64.3*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: longbench-v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}longbench id exists in data/benchmarks/ (v1); v2 variant with standardized ~128k inputs per footnote 6. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。GPT-5.2 54.5*, Claude 4.5 Opus 64.4*, Gemini 3 Pro 68.2*, DeepSeek V3.2 59.8*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Footnotes · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: browsecomp-swarm · figure: Chart image https://statics.kimi.ai/blogs/k2-5/20260127-131347.jpeg (K2.5 Agent Swarm vs Claude Opus 4.5 on BrowseComp, Wide Search, In-house Bench)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": "main agent 15 steps; sub-agents 100 steps",
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp already introduced by prior batches. Swarm-mode protocol differs from single-agent BrowseComp rows everywhere else - never merge. 数值取自归档 page.html 附录 'Benchmark table' 内嵌 JSON(机器可读文本,非视觉转写;2026-09-01 提取)。同表竞品列(K2.5 列为 Thinking):温 * 为 Kimi 转引标注。附录同 benchmark 另有两行:默认 60.6、w/ctx mgm 74.9;Agent Swarm 行仅 K2.5 有值。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Office Productivity section · row: ai-office-bench · quote_snippet: K2.5 shows 59.3% and 24.3% improvements over K2 Thinking
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ai-office-bench (internal) not yet in data/benchmarks/. Relative-improvement claim only (59.3% on AI Office Benchmark, 24.3% on General Agent Benchmark vs K2 Thinking); the companion image alt text reads 71.2%/39.0% - conflicting with prose, prose taken as authoritative, discrepancy flagged for the technical report (arXiv 2602.02276).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HLE-Full
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "96k completion budget; 256k context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Appendix full-set row - distinct from the four footnote-2 subset rows (text 31.5 / image 21.3); never average or merge. DeepSeek V3.2 cell 25.1* is its text-only subset (footnote 2 dagger); HuggingFace access blocked to prevent leakage. Competitor cells: GPT-5.2 (xhigh) 34.5, Claude 4.5 Opus (Extend Thinking) 30.8, Gemini 3 Pro (High) 37.5, Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: HLE-Full w/ tools
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "context management: once context exceeds threshold, only latest round of tool messages retained",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Appendix full-set w/ tools row - distinct from footnote-2 subset rows (text 51.8 / image 39.8); never merge. DeepSeek V3.2 cell 40.8* is its text-only subset (dagger). Competitor cells: GPT-5.2 45.5, Claude 4.5 Opus 43.2, Gemini 3 Pro 45.8, Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MMLU-Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Asterisked competitor cells re-tested by Kimi. Competitor cells: GPT-5.2 86.7*, Claude 4.5 Opus 89.3*, Gemini 3 Pro 90.1, DeepSeek V3.2 85.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: CharXiv (RQ)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 4: vision protocol max-tokens 64k, avg@3. Asterisked cells re-tested by Kimi. Competitor cells: GPT-5.2 82.1, Claude 4.5 Opus 67.2*, Gemini 3 Pro 81.4, Qwen3-VL 66.1; DeepSeek V3.2 not reported (text-only model on vision rows).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MathVista (mini)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 82.8*, Claude 4.5 Opus 80.2*, Gemini 3 Pro 89.8*, Qwen3-VL 85.8; DeepSeek V3.2 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: ZeroBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 4: vision protocol max-tokens 64k, avg@3. No-tools row - distinct from ZeroBench (w/ tools) 11 row; never merge. Competitor cells: GPT-5.2 9*, Claude 4.5 Opus 3*, Gemini 3 Pro 8*, Qwen3-VL 4*; DeepSeek V3.2 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OCRBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ocrbench not yet in data/benchmarks/. Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 80.7*, Claude 4.5 Opus 86.5*, Gemini 3 Pro 90.3*, Qwen3-VL 87.5; DeepSeek V3.2 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: InfoVQA (test)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: infovqa not yet in data/benchmarks/. Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 84*, Claude 4.5 Opus 76.9*, Gemini 3 Pro 57.2*, Qwen3-VL 89.5; DeepSeek V3.2 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SimpleVQA
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 55.8*, Claude 4.5 Opus 69.7*, Gemini 3 Pro 69.7*, Qwen3-VL 56.8*; DeepSeek V3.2 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MMVU
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 80.8*, Claude 4.5 Opus 77.3, Gemini 3 Pro 77.5, Qwen3-VL 71.1; DeepSeek V3.2 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: MotionBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 64.8, Claude 4.5 Opus 60.3, Gemini 3 Pro 70.3; DeepSeek V3.2 and Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: VideoMME
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: GPT-5.2 86.0, Gemini 3 Pro 88.4*, Qwen3-VL 79.0; Claude 4.5 Opus and DeepSeek V3.2 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: LVBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 4: vision protocol max-tokens 64k, avg@3. Competitor cells: Gemini 3 Pro 73.5*, Qwen3-VL 63.6; GPT-5.2, Claude 4.5 Opus and DeepSeek V3.2 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SWE-Bench Pro
{
"harness": "in-house SWE-series framework (minimal tool set), non-thinking mode scored highest",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Competitor cells: GPT-5.2 55.6, Claude 4.5 Opus 55.4*; Gemini 3 Pro, DeepSeek V3.2 and Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: PaperBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Competitor cells: GPT-5.2 63.7*, Claude 4.5 Opus 72.9*, DeepSeek V3.2 47.1; Gemini 3 Pro and Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: CyberGym
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Claude Opus 4.5 cell 50.6 reported by Kimi under the non-thinking setting (footnote 5 disclosure). Competitor cells: Gemini 3 Pro 39.9*, DeepSeek V3.2 17.3*; GPT-5.2 and Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: SciCode
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Distinct from K2 Thinking's SciCode run (temp 0.0, no tools) - different protocol, never merge. Competitor cells: GPT-5.2 52.1, Claude 4.5 Opus 49.5, Gemini 3 Pro 56.1, DeepSeek V3.2 38.9; Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OJBench (cpp)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Competitor cells: Claude 4.5 Opus 54.6*, Gemini 3 Pro 68.5*, DeepSeek V3.2 54.7*; GPT-5.2 and Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: LiveCodeBench (v6)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Footnote 5: in-house SWE-series framework (bash/createfile/insert/view/strreplace/submit tools) with tailored system prompts; highest scores under non-thinking mode; all coding tasks averaged over 5 independent runs. Competitor cells: Claude 4.5 Opus 82.2*, Gemini 3 Pro 87.4*, DeepSeek V3.2 83.3; GPT-5.2 and Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BrowseComp
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 3: agentic-search tool set (search, code-interpreter, web-browsing); for BrowseComp K2.5 and DeepSeek V3.2 used the discard-all context strategy. One of three BrowseComp protocols on this page (default / w/ ctx mgm / Agent Swarm) - never merge. Competitor cells: Claude 4.5 Opus 37.0, Gemini 3 Pro 37.8, DeepSeek V3.2 51.4; GPT-5.2 and Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: BrowseComp (w/ctx mgm)
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}One of three BrowseComp protocols on this page; w/ context management row is a different protocol from the default 60.6 and Agent Swarm 78.4 rows - never merge. Competitor cells: GPT-5.2 65.8, Claude 4.5 Opus 57.8, Gemini 3 Pro 59.2, DeepSeek V3.2 67.6; Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: WideSearch (item-f1) (Agent Swarm)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": "main and sub-agents max 100 steps",
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: widesearch registered. Footnote 7: WideSearch Swarm Mode main and sub-agents max 100 steps. Swarm row - distinct from single-agent WideSearch 72.7; only K2.5 has a value on this row (all competitor cells '-').
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: DeepSearchQA
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 3 tool set. Competitor cells: GPT-5.2 71.3*, Claude 4.5 Opus 76.1*, Gemini 3 Pro 63.2*, DeepSeek V3.2 60.9*; Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: FinSearchCompT2&T3
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote 3 tool set. Combined T2&T3 split - distinct from K2 Thinking's T3-only row. Competitor cells: Claude 4.5 Opus 66.2*, Gemini 3 Pro 49.9, DeepSeek V3.2 59.1*; GPT-5.2 and Qwen3-VL not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: OSWorld-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Computer-use section. Competitor cells: GPT-5.2 8.6*, Claude 4.5 Opus 66.3, Gemini 3 Pro 20.7*, Qwen3-VL 38.1; DeepSeek V3.2 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix / Benchmark table · table: Appendix 'Benchmark table'(页面内嵌 JSON 数据,归档 page.html payload 中机器可读;页面正文以图片渲染) · row: WebArena
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Computer-use section. Competitor cells: Claude 4.5 Opus 63.4*, Qwen3-VL 26.4*; GPT-5.2, Gemini 3 Pro and DeepSeek V3.2 not reported.