← 模型目录

Kimi K2 Thinking

Moonshot AI / Kimi · 2025-11-06 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Kimi K2 Thinking

月之暗面将 Kimi K2 Thinking 定位为思考型智能体(thinking agent),通过同时扩展思考 token 与工具调用步数推进测试时扩展前沿,可跨数百步规划、推理、执行与调整,全部评测结果以 INT4 QAT 精度报告。已收录评测覆盖通用推理、数学、智能体检索、编码与终端等领域,HLE 带工具 44.9(Heavy Mode 51.0)、SWE-bench Verified 71.3。

输入模态
文本
上下文
256K
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 23.9 模型 kimi-k2-thinking · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' (single table on page) · row: Humanity's Last Exam / no tools · quote_snippet: Humanity's Last Exam | no tools | 23.9

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "96k thinking-token budget; 256k context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Three HLE conditions printed as distinct sub-rows; do not merge. Competitor cells same row: GPT-5 (High) 26.3, Claude Sonnet 4.5 (Thinking) 19.8*, K2 0905 7.9, DeepSeek-V3.2 19.8, Grok-4 25.4. Global footnote 2a: all benchmarks at temperature=1.0, 256k context (SciCode exception temp 0.0).

打开官方来源

hlehle 44.9 模型 kimi-k2-thinking · 版本 w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks + prose section Agentic Reasoning · table: DOM table under 'Full Evaluations [2]' · row: Humanity's Last Exam / w/ tools [4] · quote_snippet: K2 Thinking achieves 44.9% on HLE with tools

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "120 max steps, 48k-token reasoning budget per step",
  "turn_limit": 120,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "o3-mini (official HLE setting, judge prompts verbatim from official repository)"
}

Prose cross-check: '44.9% on HLE with tools' in Evaluations summary. Footnote 4f: Hugging Face access blocked during testing; without blocking K2 Thinking scores 51.3 (contamination disclosure). Competitor cells: GPT-5 41.7, Sonnet 4.5 32.0*, K2 0905 21.7, DeepSeek-V3.2 20.3*, Grok-4 41.0.

打开官方来源

hlehle 51.0 模型 kimi-k2-thinking · 版本 heavy mode, w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: Humanity's Last Exam / heavy [6] · quote_snippet: heavy [6] | 51.0

{
  "harness": "parallel rollout strategy: 8 simultaneous trajectories + reflective aggregation",
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": 120,
  "time_limit": null,
  "run_count": 8,
  "aggregation": "reflective aggregation over 8 trajectories",
  "judge": "o3-mini"
}

Footnote 6: Heavy Mode is a different aggregation protocol; the GPT-5 heavy column denotes the official GPT-5 Pro score (Grok-4 heavy 50.7). Not comparable to single-trajectory rows.

打开官方来源

aime-25 94.5 模型 kimi-k2-thinking · 版本 no tools · 指标 avg@32 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: AIME 2025 / no tools · quote_snippet: AIME 2025 | no tools | 94.5

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "96k thinking-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": 32,
  "aggregation": "avg@32",
  "judge": null
}

Footnote 2c: AIME/HMMT no-tools reported as average of 32 runs. Competitor cells: GPT-5 94.6, Sonnet 4.5 87.0, K2 0905 51.0, DeepSeek-V3.2 89.3, Grok-4 91.7.

打开官方来源

aime-25 99.1 模型 kimi-k2-thinking · 版本 w/ python · 指标 avg@16 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: AIME 2025 / w/ python · quote_snippet: AIME 2025 | w/ python | 99.1

{
  "harness": null,
  "tools": [
    "python"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "96k thinking-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": 16,
  "aggregation": "avg@16",
  "judge": null
}

Footnote 2c: w/-python rows averaged over 16 runs. Competitor cells: GPT-5 99.6, Sonnet 4.5 100.0, K2 0905 75.2, DeepSeek-V3.2 58.1*, Grok-4 98.8.

打开官方来源

aime-25 100.0 模型 kimi-k2-thinking · 版本 heavy mode, w/ python · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: AIME 2025 / heavy [6] · quote_snippet: AIME 2025 | heavy [6] | 100.0

{
  "harness": "8 simultaneous trajectories + reflective aggregation",
  "tools": [
    "python"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 8,
  "aggregation": "reflective aggregation over 8 trajectories",
  "judge": null
}

GPT-5 heavy (= official GPT-5 Pro) 100.0, Grok-4 heavy 100.0. Not comparable to avg@32/avg@16 rows.

打开官方来源

hmmt25 89.4 模型 kimi-k2-thinking · 版本 no tools · 指标 avg@32 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: HMMT 2025 / no tools · quote_snippet: HMMT 2025 | no tools | 89.4

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "96k thinking-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": 32,
  "aggregation": "avg@32",
  "judge": null
}

new-benchmark marker retained from prior batches: hmmt-25 still not in data/benchmarks/. Competitor cells: GPT-5 93.3, Sonnet 4.5 74.6*, K2 0905 38.8, DeepSeek-V3.2 83.6, Grok-4 90.0.

打开官方来源

hmmt25 95.1 模型 kimi-k2-thinking · 版本 w/ python · 指标 avg@16 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: HMMT 2025 / w/ python · quote_snippet: HMMT 2025 | w/ python | 95.1

{
  "harness": null,
  "tools": [
    "python"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "96k thinking-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": 16,
  "aggregation": "avg@16",
  "judge": null
}

Competitor cells: GPT-5 96.7, Sonnet 4.5 88.8*, K2 0905 70.4, DeepSeek-V3.2 49.5*, Grok-4 93.9.

打开官方来源

hmmt25 97.5 模型 kimi-k2-thinking · 版本 heavy mode, w/ python · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: HMMT 2025 / heavy [6] · quote_snippet: HMMT 2025 | heavy [6] | 97.5

{
  "harness": "8 simultaneous trajectories + reflective aggregation",
  "tools": [
    "python"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 8,
  "aggregation": "reflective aggregation over 8 trajectories",
  "judge": null
}

GPT-5 heavy (= official GPT-5 Pro) 100.0, Grok-4 heavy 96.7.

打开官方来源

imo-answerbench 78.6 模型 kimi-k2-thinking · 版本 no tools · 指标 avg@8 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: IMO-AnswerBench · quote_snippet: IMO-AnswerBench | no tools | 78.6

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "128k thinking-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": 8,
  "aggregation": "avg@8",
  "judge": null
}

new-benchmark: imo-answerbench already introduced by prior batches, still not in data/benchmarks/. Footnote 3c: GPT-5 scored 65.6 in the benchmark paper; Kimi re-evaluated GPT-5 with official API obtaining 76.0 (printed cell). Sonnet 4.5 65.9*, K2 0905 45.8, DeepSeek-V3.2 76.0*, Grok-4 73.1.

打开官方来源

gpqa 84.5 模型 kimi-k2-thinking · 版本 Diamond, no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: GPQA-Diamond · quote_snippet: GPQA-Diamond | no tools | 84.5

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "96k thinking-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GPT-5 85.7, Sonnet 4.5 83.4, K2 0905 74.2, DeepSeek-V3.2 79.9, Grok-4 87.5.

打开官方来源

mmlu-pro 84.6 模型 kimi-k2-thinking · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / General Tasks · table: DOM table under 'Full Evaluations [2]' · row: MMLU-Pro · quote_snippet: MMLU-Pro | no tools | 84.6

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GPT-5 87.1, Sonnet 4.5 87.5, K2 0905 81.9, DeepSeek-V3.2 85.0, Grok-4 not reported.

打开官方来源

mmlu-redux 94.4 模型 kimi-k2-thinking · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / General Tasks · table: DOM table under 'Full Evaluations [2]' · row: MMLU-Redux · quote_snippet: MMLU-Redux | no tools | 94.4

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GPT-5 95.3, Sonnet 4.5 95.6, K2 0905 92.7, DeepSeek-V3.2 93.7, Grok-4 not reported.

打开官方来源

longform-writing 73.8 模型 kimi-k2-thinking · 版本 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / General Tasks · table: DOM table under 'Full Evaluations [2]' · row: Longform Writing · quote_snippet: Longform Writing | no tools | 73.8

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "32k completion-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: longform-writing not yet in data/benchmarks/. Metric not labeled beyond table Intro 'no tools'. Competitor cells: GPT-5 71.4, Sonnet 4.5 79.8, K2 0905 62.8, DeepSeek-V3.2 72.5.

打开官方来源

healthbench 58.0 模型 kimi-k2-thinking · 版本 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / General Tasks · table: DOM table under 'Full Evaluations [2]' · row: HealthBench · quote_snippet: HealthBench | no tools | 58.0

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: healthbench (OpenAI HealthBench) not yet in data/benchmarks/. Competitor cells: GPT-5 67.2, Sonnet 4.5 44.2, K2 0905 43.8, DeepSeek-V3.2 46.9.

打开官方来源

browsecomp 60.2 模型 kimi-k2-thinking · 版本 w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Agentic Search Tasks + prose Agentic Search and Browsing · table: DOM table under 'Full Evaluations [2]' · row: BrowseComp · quote_snippet: K2 Thinking achieved a score of 60.2%, significantly outperforming the human baseline of 29.2%

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "300 max steps, 24k-token reasoning budget per step; context management hides previous tool outputs when accumulated input exceeds 256k",
  "turn_limit": 300,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Prose discloses human baseline 29.2%. Competitor cells: GPT-5 54.9, Sonnet 4.5 24.1, K2 0905 7.4, DeepSeek-V3.2 40.1, Grok-4 not reported.

打开官方来源

browsecomp-zh 62.3 模型 kimi-k2-thinking · 版本 w/ tools · 指标 avg@4 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Agentic Search Tasks · table: DOM table under 'Full Evaluations [2]' · row: BrowseComp-ZH · quote_snippet: BrowseComp-ZH | w/ tools | 62.3

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "300 max steps, 24k-token reasoning budget per step",
  "turn_limit": 300,
  "time_limit": null,
  "run_count": 4,
  "aggregation": "avg@4 (4 independent runs)",
  "judge": null
}

new-benchmark: browsecomp-zh not yet in data/benchmarks/. Footnote 4b: run 4 times independently, average reported. Competitor cells: GPT-5 63.0*, Sonnet 4.5 42.4*, K2 0905 22.2, DeepSeek-V3.2 47.9.

打开官方来源

seal-0 56.3 模型 kimi-k2-thinking · 版本 w/ tools · 指标 avg@4 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Agentic Search Tasks · table: DOM table under 'Full Evaluations [2]' · row: Seal-0 · quote_snippet: Seal-0 | w/ tools | 56.3

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "300 max steps, 24k-token reasoning budget per step",
  "turn_limit": 300,
  "time_limit": null,
  "run_count": 4,
  "aggregation": "avg@4",
  "judge": null
}

new-benchmark: seal-0 not yet in data/benchmarks/. Competitor cells: GPT-5 51.4*, Sonnet 4.5 53.4*, K2 0905 25.2, DeepSeek-V3.2 38.5*.

打开官方来源

finsearchcomp 47.4 模型 kimi-k2-thinking · 版本 T3 · 指标 avg@4 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Agentic Search Tasks · table: DOM table under 'Full Evaluations [2]' · row: FinSearchComp-T3 · quote_snippet: FinSearchComp-T3 | w/ tools | 47.4

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "300 max steps, 24k-token reasoning budget per step",
  "turn_limit": 300,
  "time_limit": null,
  "run_count": 4,
  "aggregation": "avg@4",
  "judge": null
}

finsearchcomp id already introduced by prior batches (Seed 1.8), still not in data/benchmarks/. Competitor cells: GPT-5 48.5*, Sonnet 4.5 44.0*, K2 0905 10.4, DeepSeek-V3.2 27.0*.

打开官方来源

frames 87.0 模型 kimi-k2-thinking · 版本 w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Agentic Search Tasks · table: DOM table under 'Full Evaluations [2]' · row: Frames · quote_snippet: Frames | w/ tools | 87.0

{
  "harness": null,
  "tools": [
    "search",
    "code-interpreter",
    "web-browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "300 max steps, 24k-token reasoning budget per step",
  "turn_limit": 300,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: frames not yet in data/benchmarks/. Competitor cells: GPT-5 86.0*, Sonnet 4.5 85.0*, K2 0905 58.1, DeepSeek-V3.2 80.2*.

打开官方来源

swebench 71.3 模型 kimi-k2-thinking · 版本 Verified · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Coding Tasks + prose Agentic Coding · table: DOM table under 'Full Evaluations [2]' · row: SWE-bench Verified · quote_snippet: It achieves scores of 61.1% on SWE-Multilingual, 71.3% on SWE-Bench Verified, and 47.1% on Terminal-Bench

{
  "harness": "in-house harness derived from SWE-agent (Bash/Edit tool context windows clamped, system prompt rewritten to task semantics)",
  "tools": [
    "agentic coding tools"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Footnote 5c: all reported coding scores averaged over 5 independent runs. Competitor cells: GPT-5 74.9, Sonnet 4.5 77.2, K2 0905 69.2, DeepSeek-V3.2 67.8, Grok-4 not reported.

打开官方来源

swebench-multilingual 61.1 模型 kimi-k2-thinking · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: SWE-bench Multilingual · quote_snippet: SWE-bench Multilingual | w/ tools | 61.1

{
  "harness": "in-house harness derived from SWE-agent",
  "tools": [
    "agentic coding tools"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Competitor cells: GPT-5 55.3*, Sonnet 4.5 68.0, K2 0905 55.9, DeepSeek-V3.2 57.9, Grok-4 not reported.

打开官方来源

multi-swe-bench 41.9 模型 kimi-k2-thinking · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: Multi-SWE-bench · quote_snippet: Multi-SWE-bench | w/ tools | 41.9

{
  "harness": "in-house harness derived from SWE-agent",
  "tools": [
    "agentic coding tools"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

new-benchmark: multi-swe-bench already introduced by prior batches, still not in data/benchmarks/. Competitor cells: GPT-5 39.3*, Sonnet 4.5 44.3, K2 0905 33.5, DeepSeek-V3.2 30.6.

打开官方来源

scicode 44.8 模型 kimi-k2-thinking · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: SciCode · quote_snippet: SciCode | no tools | 44.8

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": null,
  "token_budget": "128k thinking-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

scicode id already introduced by prior batches, still not in data/benchmarks/. Footnote 2a: SciCode is the sole exception to the global temperature=1.0 (official setting 0.0). Competitor cells: GPT-5 42.9, Sonnet 4.5 44.7, K2 0905 30.7, DeepSeek-V3.2 37.7.

打开官方来源

lcb 83.1 模型 kimi-k2-thinking · 版本 v6 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: LiveCodeBench v6 · quote_snippet: LiveCodeBench v6 | no tools | 83.1

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "128k thinking-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GPT-5 87.0*, Sonnet 4.5 64.0*, K2 0905 56.1*, DeepSeek-V3.2 74.1.

打开官方来源

ojbench 48.7 模型 kimi-k2-thinking · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: OJ-Bench · quote_snippet: OJ-Bench | no tools | 48.7

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": "128k thinking-token budget",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ojbench (also introduced in kimi-k2.json this batch). Competitor cells: GPT-5 56.2*, Sonnet 4.5 30.4*, K2 0905 25.5*, DeepSeek-V3.2 38.2*.

打开官方来源

terminalbench 47.1 模型 kimi-k2-thinking · 版本 w/ simulated tools (JSON) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Evaluations / Coding Tasks + prose Agentic Coding · table: DOM table under 'Full Evaluations [2]' · row: Terminal-Bench · quote_snippet: Terminal-Bench | w/ simulated tools (JSON) | 47.1

{
  "harness": "Terminus-2 (default agent framework) with provided JSON parser",
  "tools": [
    "simulated terminal tools (JSON)"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 independent runs",
  "judge": null
}

Footnote 3a: GPT-5 Terminal-Bench value quoted from the public Terminal-Bench leaderboard (Terminus-2). Competitor cells: GPT-5 43.8, Sonnet 4.5 51.0, K2 0905 44.5, DeepSeek-V3.2 37.7.

打开官方来源