Kimi K2 Thinking
Moonshot AI / Kimi · 2025-11-06 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Kimi K2 Thinking
月之暗面将 Kimi K2 Thinking 定位为思考型智能体(thinking agent),通过同时扩展思考 token 与工具调用步数推进测试时扩展前沿,可跨数百步规划、推理、执行与调整,全部评测结果以 INT4 QAT 精度报告。已收录评测覆盖通用推理、数学、智能体检索、编码与终端等领域,HLE 带工具 44.9(Heavy Mode 51.0)、SWE-bench Verified 71.3。
- 输入模态
- 文本
- 上下文
- 256K
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' (single table on page) · row: Humanity's Last Exam / no tools · quote_snippet: Humanity's Last Exam | no tools | 23.9
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "96k thinking-token budget; 256k context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Three HLE conditions printed as distinct sub-rows; do not merge. Competitor cells same row: GPT-5 (High) 26.3, Claude Sonnet 4.5 (Thinking) 19.8*, K2 0905 7.9, DeepSeek-V3.2 19.8, Grok-4 25.4. Global footnote 2a: all benchmarks at temperature=1.0, 256k context (SciCode exception temp 0.0).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks + prose section Agentic Reasoning · table: DOM table under 'Full Evaluations [2]' · row: Humanity's Last Exam / w/ tools [4] · quote_snippet: K2 Thinking achieves 44.9% on HLE with tools
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "120 max steps, 48k-token reasoning budget per step",
"turn_limit": 120,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "o3-mini (official HLE setting, judge prompts verbatim from official repository)"
}Prose cross-check: '44.9% on HLE with tools' in Evaluations summary. Footnote 4f: Hugging Face access blocked during testing; without blocking K2 Thinking scores 51.3 (contamination disclosure). Competitor cells: GPT-5 41.7, Sonnet 4.5 32.0*, K2 0905 21.7, DeepSeek-V3.2 20.3*, Grok-4 41.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: Humanity's Last Exam / heavy [6] · quote_snippet: heavy [6] | 51.0
{
"harness": "parallel rollout strategy: 8 simultaneous trajectories + reflective aggregation",
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": 120,
"time_limit": null,
"run_count": 8,
"aggregation": "reflective aggregation over 8 trajectories",
"judge": "o3-mini"
}Footnote 6: Heavy Mode is a different aggregation protocol; the GPT-5 heavy column denotes the official GPT-5 Pro score (Grok-4 heavy 50.7). Not comparable to single-trajectory rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: AIME 2025 / no tools · quote_snippet: AIME 2025 | no tools | 94.5
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "96k thinking-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": 32,
"aggregation": "avg@32",
"judge": null
}Footnote 2c: AIME/HMMT no-tools reported as average of 32 runs. Competitor cells: GPT-5 94.6, Sonnet 4.5 87.0, K2 0905 51.0, DeepSeek-V3.2 89.3, Grok-4 91.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: AIME 2025 / w/ python · quote_snippet: AIME 2025 | w/ python | 99.1
{
"harness": null,
"tools": [
"python"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "96k thinking-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": 16,
"aggregation": "avg@16",
"judge": null
}Footnote 2c: w/-python rows averaged over 16 runs. Competitor cells: GPT-5 99.6, Sonnet 4.5 100.0, K2 0905 75.2, DeepSeek-V3.2 58.1*, Grok-4 98.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: AIME 2025 / heavy [6] · quote_snippet: AIME 2025 | heavy [6] | 100.0
{
"harness": "8 simultaneous trajectories + reflective aggregation",
"tools": [
"python"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 8,
"aggregation": "reflective aggregation over 8 trajectories",
"judge": null
}GPT-5 heavy (= official GPT-5 Pro) 100.0, Grok-4 heavy 100.0. Not comparable to avg@32/avg@16 rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: HMMT 2025 / no tools · quote_snippet: HMMT 2025 | no tools | 89.4
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "96k thinking-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": 32,
"aggregation": "avg@32",
"judge": null
}new-benchmark marker retained from prior batches: hmmt-25 still not in data/benchmarks/. Competitor cells: GPT-5 93.3, Sonnet 4.5 74.6*, K2 0905 38.8, DeepSeek-V3.2 83.6, Grok-4 90.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: HMMT 2025 / w/ python · quote_snippet: HMMT 2025 | w/ python | 95.1
{
"harness": null,
"tools": [
"python"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "96k thinking-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": 16,
"aggregation": "avg@16",
"judge": null
}Competitor cells: GPT-5 96.7, Sonnet 4.5 88.8*, K2 0905 70.4, DeepSeek-V3.2 49.5*, Grok-4 93.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: HMMT 2025 / heavy [6] · quote_snippet: HMMT 2025 | heavy [6] | 97.5
{
"harness": "8 simultaneous trajectories + reflective aggregation",
"tools": [
"python"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 8,
"aggregation": "reflective aggregation over 8 trajectories",
"judge": null
}GPT-5 heavy (= official GPT-5 Pro) 100.0, Grok-4 heavy 96.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: IMO-AnswerBench · quote_snippet: IMO-AnswerBench | no tools | 78.6
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "128k thinking-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": 8,
"aggregation": "avg@8",
"judge": null
}new-benchmark: imo-answerbench already introduced by prior batches, still not in data/benchmarks/. Footnote 3c: GPT-5 scored 65.6 in the benchmark paper; Kimi re-evaluated GPT-5 with official API obtaining 76.0 (printed cell). Sonnet 4.5 65.9*, K2 0905 45.8, DeepSeek-V3.2 76.0*, Grok-4 73.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Reasoning Tasks · table: DOM table under 'Full Evaluations [2]' · row: GPQA-Diamond · quote_snippet: GPQA-Diamond | no tools | 84.5
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "96k thinking-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GPT-5 85.7, Sonnet 4.5 83.4, K2 0905 74.2, DeepSeek-V3.2 79.9, Grok-4 87.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / General Tasks · table: DOM table under 'Full Evaluations [2]' · row: MMLU-Pro · quote_snippet: MMLU-Pro | no tools | 84.6
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GPT-5 87.1, Sonnet 4.5 87.5, K2 0905 81.9, DeepSeek-V3.2 85.0, Grok-4 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / General Tasks · table: DOM table under 'Full Evaluations [2]' · row: MMLU-Redux · quote_snippet: MMLU-Redux | no tools | 94.4
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GPT-5 95.3, Sonnet 4.5 95.6, K2 0905 92.7, DeepSeek-V3.2 93.7, Grok-4 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / General Tasks · table: DOM table under 'Full Evaluations [2]' · row: Longform Writing · quote_snippet: Longform Writing | no tools | 73.8
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "32k completion-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: longform-writing not yet in data/benchmarks/. Metric not labeled beyond table Intro 'no tools'. Competitor cells: GPT-5 71.4, Sonnet 4.5 79.8, K2 0905 62.8, DeepSeek-V3.2 72.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / General Tasks · table: DOM table under 'Full Evaluations [2]' · row: HealthBench · quote_snippet: HealthBench | no tools | 58.0
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: healthbench (OpenAI HealthBench) not yet in data/benchmarks/. Competitor cells: GPT-5 67.2, Sonnet 4.5 44.2, K2 0905 43.8, DeepSeek-V3.2 46.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Agentic Search Tasks + prose Agentic Search and Browsing · table: DOM table under 'Full Evaluations [2]' · row: BrowseComp · quote_snippet: K2 Thinking achieved a score of 60.2%, significantly outperforming the human baseline of 29.2%
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "300 max steps, 24k-token reasoning budget per step; context management hides previous tool outputs when accumulated input exceeds 256k",
"turn_limit": 300,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Prose discloses human baseline 29.2%. Competitor cells: GPT-5 54.9, Sonnet 4.5 24.1, K2 0905 7.4, DeepSeek-V3.2 40.1, Grok-4 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Agentic Search Tasks · table: DOM table under 'Full Evaluations [2]' · row: BrowseComp-ZH · quote_snippet: BrowseComp-ZH | w/ tools | 62.3
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "300 max steps, 24k-token reasoning budget per step",
"turn_limit": 300,
"time_limit": null,
"run_count": 4,
"aggregation": "avg@4 (4 independent runs)",
"judge": null
}new-benchmark: browsecomp-zh not yet in data/benchmarks/. Footnote 4b: run 4 times independently, average reported. Competitor cells: GPT-5 63.0*, Sonnet 4.5 42.4*, K2 0905 22.2, DeepSeek-V3.2 47.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Agentic Search Tasks · table: DOM table under 'Full Evaluations [2]' · row: Seal-0 · quote_snippet: Seal-0 | w/ tools | 56.3
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "300 max steps, 24k-token reasoning budget per step",
"turn_limit": 300,
"time_limit": null,
"run_count": 4,
"aggregation": "avg@4",
"judge": null
}new-benchmark: seal-0 not yet in data/benchmarks/. Competitor cells: GPT-5 51.4*, Sonnet 4.5 53.4*, K2 0905 25.2, DeepSeek-V3.2 38.5*.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Agentic Search Tasks · table: DOM table under 'Full Evaluations [2]' · row: FinSearchComp-T3 · quote_snippet: FinSearchComp-T3 | w/ tools | 47.4
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "300 max steps, 24k-token reasoning budget per step",
"turn_limit": 300,
"time_limit": null,
"run_count": 4,
"aggregation": "avg@4",
"judge": null
}finsearchcomp id already introduced by prior batches (Seed 1.8), still not in data/benchmarks/. Competitor cells: GPT-5 48.5*, Sonnet 4.5 44.0*, K2 0905 10.4, DeepSeek-V3.2 27.0*.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Agentic Search Tasks · table: DOM table under 'Full Evaluations [2]' · row: Frames · quote_snippet: Frames | w/ tools | 87.0
{
"harness": null,
"tools": [
"search",
"code-interpreter",
"web-browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "300 max steps, 24k-token reasoning budget per step",
"turn_limit": 300,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: frames not yet in data/benchmarks/. Competitor cells: GPT-5 86.0*, Sonnet 4.5 85.0*, K2 0905 58.1, DeepSeek-V3.2 80.2*.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Coding Tasks + prose Agentic Coding · table: DOM table under 'Full Evaluations [2]' · row: SWE-bench Verified · quote_snippet: It achieves scores of 61.1% on SWE-Multilingual, 71.3% on SWE-Bench Verified, and 47.1% on Terminal-Bench
{
"harness": "in-house harness derived from SWE-agent (Bash/Edit tool context windows clamped, system prompt rewritten to task semantics)",
"tools": [
"agentic coding tools"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Footnote 5c: all reported coding scores averaged over 5 independent runs. Competitor cells: GPT-5 74.9, Sonnet 4.5 77.2, K2 0905 69.2, DeepSeek-V3.2 67.8, Grok-4 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: SWE-bench Multilingual · quote_snippet: SWE-bench Multilingual | w/ tools | 61.1
{
"harness": "in-house harness derived from SWE-agent",
"tools": [
"agentic coding tools"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Competitor cells: GPT-5 55.3*, Sonnet 4.5 68.0, K2 0905 55.9, DeepSeek-V3.2 57.9, Grok-4 not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: Multi-SWE-bench · quote_snippet: Multi-SWE-bench | w/ tools | 41.9
{
"harness": "in-house harness derived from SWE-agent",
"tools": [
"agentic coding tools"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}new-benchmark: multi-swe-bench already introduced by prior batches, still not in data/benchmarks/. Competitor cells: GPT-5 39.3*, Sonnet 4.5 44.3, K2 0905 33.5, DeepSeek-V3.2 30.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: SciCode · quote_snippet: SciCode | no tools | 44.8
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": "128k thinking-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}scicode id already introduced by prior batches, still not in data/benchmarks/. Footnote 2a: SciCode is the sole exception to the global temperature=1.0 (official setting 0.0). Competitor cells: GPT-5 42.9, Sonnet 4.5 44.7, K2 0905 30.7, DeepSeek-V3.2 37.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: LiveCodeBench v6 · quote_snippet: LiveCodeBench v6 | no tools | 83.1
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "128k thinking-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GPT-5 87.0*, Sonnet 4.5 64.0*, K2 0905 56.1*, DeepSeek-V3.2 74.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Coding Tasks · table: DOM table under 'Full Evaluations [2]' · row: OJ-Bench · quote_snippet: OJ-Bench | no tools | 48.7
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": "128k thinking-token budget",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ojbench (also introduced in kimi-k2.json this batch). Competitor cells: GPT-5 56.2*, Sonnet 4.5 30.4*, K2 0905 25.5*, DeepSeek-V3.2 38.2*.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Evaluations / Coding Tasks + prose Agentic Coding · table: DOM table under 'Full Evaluations [2]' · row: Terminal-Bench · quote_snippet: Terminal-Bench | w/ simulated tools (JSON) | 47.1
{
"harness": "Terminus-2 (default agent framework) with provided JSON parser",
"tools": [
"simulated terminal tools (JSON)"
],
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 independent runs",
"judge": null
}Footnote 3a: GPT-5 Terminal-Bench value quoted from the public Terminal-Bench leaderboard (Terminus-2). Competitor cells: GPT-5 43.8, Sonnet 4.5 51.0, K2 0905 44.5, DeepSeek-V3.2 37.7.