← 模型目录

Kimi-K2-Instruct / Kimi-K2-Base

Moonshot AI / Kimi · 2025-07-11 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Kimi-K2-Instruct

Kimi K2 以「Open Agentic Intelligence」为定位发布,Kimi-K2-Instruct 为后训练的非思考(reflex-grade)MoE 模型(1T 总参/32B 激活)。已收录 32 项评测覆盖代码、数学、知识与 Agent 工具使用:亮点 SWE-bench Verified(Agentic Coding)65.8、MMLU 89.5。

输入模态
文本
上下文
官方资料未说明
参数
1T-A32B MoE
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

lcb 53.7 模型 kimi-k2-instruct · 版本 v6 (Aug 2024 - May 2025) · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table under 'Benchmarking Kimi K2' (only table on page), Coding Tasks section · row: LiveCodeBench v6(Aug 24-May 25) · quote_snippet: LiveCodeBench v6(Aug 24-May 25) | Pass@1 | 53.7

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Pass@1",
  "judge": null
}

Competitor cells same row: DeepSeek-V3-0324 46.9, Qwen3-235B-A22B (non-thinking) 37.0, Claude Sonnet 4 (w/o extended thinking) 48.5, Claude Opus 4 (w/o extended thinking) 47.4, GPT-4.1 44.7, Gemini 2.5 Flash Preview (05-20) 44.7.

打开官方来源

ojbench 27.1 模型 kimi-k2-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: OJBench · quote_snippet: OJBench | Pass@1 | 27.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Pass@1",
  "judge": null
}

new-benchmark: ojbench not yet in data/benchmarks/. Competitor cells: V3-0324 24.0, Qwen3-235B 11.3, Sonnet 4 15.3, Opus 4 19.6, GPT-4.1 19.5, Gemini 2.5 Flash 19.5.

打开官方来源

multipl-e 85.7 模型 kimi-k2-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: MultiPL-E · quote_snippet: MultiPL-E | Pass@1 | 85.7

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Pass@1",
  "judge": null
}

new-benchmark: multipl-e not yet in data/benchmarks/. Competitor cells: V3-0324 83.1, Qwen3-235B 78.2, Sonnet 4 88.6, Opus 4 89.6, GPT-4.1 86.7, Gemini 2.5 Flash 85.6.

打开官方来源

swebench 51.8 模型 kimi-k2-instruct · 版本 Verified (Agentless Coding) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: SWE-bench Verified (Agentless Coding) · quote_snippet: SWE-bench Verified (Agentless Coding) | Single Patch without Test (Acc) | 51.8

{
  "harness": "Agentless (single patch without test)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Single Patch without Test (Acc)",
  "judge": null
}

Three separate SWE-bench Verified protocols printed as distinct table rows; do not merge. Competitor cells: V3-0324 36.6, Qwen3-235B 39.4, Sonnet 4 50.2, Opus 4 53.0, GPT-4.1 40.8, Gemini 2.5 Flash 32.6.

打开官方来源

swebench 65.8 模型 kimi-k2-instruct · 版本 Verified (Agentic Coding) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: SWE-bench Verified (Agentic Coding) / Single Attempt (Acc) · quote_snippet: SWE-bench Verified (Agentic Coding) | Single Attempt (Acc) | 65.8

{
  "harness": "agentic (unspecified)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "Single Attempt (Acc)",
  "judge": null
}

Competitor cells carry asterisk (72.7* Sonnet 4, 72.5* Opus 4): the page prints no footnote explaining the asterisk semantics; treat those cells as vendor-reported-by-competitor values of unknown protocol. GPT-4.1 54.6, V3-0324 38.8, Qwen3-235B 34.4, Gemini 2.5 Flash not reported.

打开官方来源

swebench 71.6 模型 kimi-k2-instruct · 版本 Verified (Agentic Coding) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: SWE-bench Verified (Agentic Coding) / Multiple Attempts (Acc) · quote_snippet: SWE-bench Verified (Agentic Coding) | Multiple Attempts (Acc) | 71.6

{
  "harness": "agentic (unspecified)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Multiple Attempts (Acc)",
  "judge": null
}

Multiple-attempt aggregation is not pass@1 and is not comparable to single-attempt rows. Competitor cells (asterisked): Sonnet 4 80.2*, Opus 4 79.4*; all other columns not reported.

打开官方来源

swebench-multilingual 47.3 模型 kimi-k2-instruct · 版本 Agentic Coding, single attempt · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: SWE-bench Multilingual(Agentic Coding) · quote_snippet: SWE-bench Multilingual(Agentic Coding) | Single Attempt (Acc) | 47.3

{
  "harness": "agentic (unspecified)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "Single Attempt (Acc)",
  "judge": null
}

new-benchmark: swebench-multilingual not yet in data/benchmarks/ (id already used by prior batches' notes). Competitor cells: V3-0324 25.8, Qwen3-235B 20.9, Sonnet 4 51.0, GPT-4.1 31.5; Opus 4 and Gemini 2.5 Flash not reported.

打开官方来源

terminalbench 30.0 模型 kimi-k2-instruct · 版本 in-house framework · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: TerminalBench / Inhouse Framework (Acc) · quote_snippet: TerminalBench | Inhouse Framework (Acc) | 30.0

{
  "harness": "in-house framework (unspecified)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

Two TerminalBench harnesses printed as distinct rows; not comparable to each other. Competitor cells: Sonnet 4 35.5, Opus 4 43.2, GPT-4.1 8.3; V3-0324 / Qwen3-235B / Gemini 2.5 Flash not reported.

打开官方来源

terminalbench 25.0 模型 kimi-k2-instruct · 版本 Terminus harness · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: TerminalBench / Terminus (Acc) · quote_snippet: TerminalBench | Terminus (Acc) | 25.0

{
  "harness": "Terminus",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

Competitor cells: V3-0324 16.3, Qwen3-235B 6.6, GPT-4.1 30.3, Gemini 2.5 Flash 16.8; Claude columns not reported. Terminus is a third-party public harness, enabling cross-vendor comparison only when the harness version matches.

打开官方来源

aider 60.0 模型 kimi-k2-instruct · 版本 Polyglot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: Aider-Polyglot · quote_snippet: Aider-Polyglot | Acc | 60.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

Competitor cells: V3-0324 55.1, Qwen3-235B 61.8, Sonnet 4 56.4, Opus 4 70.7, GPT-4.1 52.4, Gemini 2.5 Flash 44.0.

打开官方来源

tau-bench 70.6 模型 kimi-k2-instruct · 版本 Tau2 retail · 指标 avg@4 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Tool Use Tasks section · row: Tau2 retail · quote_snippet: Tau2 retail | Avg@4 | 70.6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 4,
  "aggregation": "Avg@4",
  "judge": null
}

Competitor cells: V3-0324 69.1, Qwen3-235B 57.0, Sonnet 4 75.0, Opus 4 81.8, GPT-4.1 74.8, Gemini 2.5 Flash 64.3.

打开官方来源

tau-bench 56.5 模型 kimi-k2-instruct · 版本 Tau2 airline · 指标 avg@4 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Tool Use Tasks section · row: Tau2 airline · quote_snippet: Tau2 airline | Avg@4 | 56.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 4,
  "aggregation": "Avg@4",
  "judge": null
}

Competitor cells: V3-0324 39.0, Qwen3-235B 26.5, Sonnet 4 55.5, Opus 4 60.0, GPT-4.1 54.5, Gemini 2.5 Flash 42.5.

打开官方来源

tau-bench 65.8 模型 kimi-k2-instruct · 版本 Tau2 telecom · 指标 avg@4 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Tool Use Tasks section · row: Tau2 telecom · quote_snippet: Tau2 telecom | Avg@4 | 65.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 4,
  "aggregation": "Avg@4",
  "judge": null
}

Competitor cells: V3-0324 32.5, Qwen3-235B 22.1, Sonnet 4 45.2, Opus 4 57.0, GPT-4.1 38.6, Gemini 2.5 Flash 16.9.

打开官方来源

acebench 76.5 模型 kimi-k2-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Tool Use Tasks section · row: AceBench · quote_snippet: AceBench | Acc | 76.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

new-benchmark: acebench not yet in data/benchmarks/. Competitor cells: V3-0324 72.7, Qwen3-235B 70.5, Sonnet 4 76.2, Opus 4 75.6, GPT-4.1 80.1, Gemini 2.5 Flash 74.5.

打开官方来源

aime24 69.6 模型 kimi-k2-instruct · 版本 未说明 · 指标 avg@64 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: AIME 2024 · quote_snippet: AIME 2024 | Avg@64 | 69.6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 64,
  "aggregation": "Avg@64",
  "judge": null
}

Avg@64 aggregation differs from single-run AIME scores; not comparable. Competitor cells: V3-0324 59.4*, Qwen3-235B 40.1*, Sonnet 4 43.4, Opus 4 48.2, GPT-4.1 46.5, Gemini 2.5 Flash 61.3.

打开官方来源

aime-25 49.5 模型 kimi-k2-instruct · 版本 未说明 · 指标 avg@64 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: AIME 2025 · quote_snippet: AIME 2025 | Avg@64 | 49.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 64,
  "aggregation": "Avg@64",
  "judge": null
}

Competitor cells: V3-0324 46.7, Qwen3-235B 24.7*, Sonnet 4 33.1*, Opus 4 33.9*, GPT-4.1 37.0, Gemini 2.5 Flash 46.6.

打开官方来源

math500 97.4 模型 kimi-k2-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: MATH-500 · quote_snippet: MATH-500 | Acc | 97.4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

Competitor cells: V3-0324 94.0*, Qwen3-235B 91.2*, Sonnet 4 94.0, Opus 4 94.4, GPT-4.1 92.4, Gemini 2.5 Flash 95.4.

打开官方来源

hmmt25 38.8 模型 kimi-k2-instruct · 版本 未说明 · 指标 avg@32 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: HMMT 2025 · quote_snippet: HMMT 2025 | Avg@32 | 38.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 32,
  "aggregation": "Avg@32",
  "judge": null
}

new-benchmark: hmmt-25 already introduced by prior batches, still not in data/benchmarks/. Competitor cells: V3-0324 27.5, Qwen3-235B 11.9, Sonnet 4 15.9, Opus 4 15.9, GPT-4.1 19.4, Gemini 2.5 Flash 34.7.

打开官方来源

cnmo-2024 74.3 模型 kimi-k2-instruct · 版本 未说明 · 指标 avg@16 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: CNMO 2024 · quote_snippet: CNMO 2024 | Avg@16 | 74.3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 16,
  "aggregation": "Avg@16",
  "judge": null
}

new-benchmark: cnmo-2024 not yet in data/benchmarks/. Competitor cells: V3-0324 74.7, Qwen3-235B 48.6, Sonnet 4 60.4, Opus 4 57.6, GPT-4.1 56.6, Gemini 2.5 Flash 75.0.

打开官方来源

polymath-en 65.1 模型 kimi-k2-instruct · 版本 未说明 · 指标 avg@4 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: PolyMath-en · quote_snippet: PolyMath-en | Avg@4 | 65.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 4,
  "aggregation": "Avg@4",
  "judge": null
}

new-benchmark: polymath-en not yet in data/benchmarks/. Competitor cells: V3-0324 59.5, Qwen3-235B 51.9, Sonnet 4 52.8, Opus 4 49.8, GPT-4.1 54.0, Gemini 2.5 Flash 49.9.

打开官方来源

zebralogic 89.0 模型 kimi-k2-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: ZebraLogic · quote_snippet: ZebraLogic | Acc | 89.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

new-benchmark: zebralogic not yet in data/benchmarks/. Competitor cells: V3-0324 84.0, Qwen3-235B 37.7*, Sonnet 4 79.7, Opus 4 59.3, GPT-4.1 58.5, Gemini 2.5 Flash 57.9.

打开官方来源

autologi 89.5 模型 kimi-k2-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: AutoLogi · quote_snippet: AutoLogi | Acc | 89.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

new-benchmark: autologi not yet in data/benchmarks/. Competitor cells: V3-0324 88.9, Qwen3-235B 83.3*, Sonnet 4 89.8, Opus 4 86.1, GPT-4.1 88.2, Gemini 2.5 Flash 84.1.

打开官方来源

gpqa 75.1 模型 kimi-k2-instruct · 版本 Diamond · 指标 avg@8 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: GPQA-Diamond · quote_snippet: GPQA-Diamond | Avg@8 | 75.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 8,
  "aggregation": "Avg@8",
  "judge": null
}

Competitor cells: V3-0324 68.4*, Qwen3-235B 62.9*, Sonnet 4 70.0*, Opus 4 74.9*, GPT-4.1 66.3, Gemini 2.5 Flash 68.2.

打开官方来源

supergpqa 57.2 模型 kimi-k2-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: SuperGPQA · quote_snippet: SuperGPQA | Acc | 57.2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

new-benchmark: supergpqa already introduced by prior batches, still not in data/benchmarks/. Competitor cells: V3-0324 53.7, Qwen3-235B 50.2, Sonnet 4 55.7, Opus 4 56.5, GPT-4.1 50.8, Gemini 2.5 Flash 49.6.

打开官方来源

hlehle 4.7 模型 kimi-k2-instruct · 版本 Text Only · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: Humanity's Last Exam (Text Only) · quote_snippet: Humanity's Last Exam (Text Only) | Acc | 4.7

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

Text-only subset, no tools; not comparable to HLE w/ tools rows in other releases (e.g. K2 Thinking 44.9). Competitor cells: V3-0324 5.2, Qwen3-235B 5.7, Sonnet 4 5.8, Opus 4 7.1, GPT-4.1 3.7, Gemini 2.5 Flash 5.6.

打开官方来源

mmlu 89.5 模型 kimi-k2-instruct · 版本 未说明 · 指标 exact_match · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: MMLU · quote_snippet: MMLU | EM | 89.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "EM",
  "judge": null
}

Competitor cells: V3-0324 89.4, Qwen3-235B 87.0, Sonnet 4 91.5, Opus 4 92.9, GPT-4.1 90.4, Gemini 2.5 Flash 90.1.

打开官方来源

mmlu-redux 92.7 模型 kimi-k2-instruct · 版本 未说明 · 指标 exact_match · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: MMLU-Redux · quote_snippet: MMLU-Redux | EM | 92.7

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "EM",
  "judge": null
}

Competitor cells: V3-0324 90.5, Qwen3-235B 89.2*, Sonnet 4 93.6, Opus 4 94.2, GPT-4.1 92.4, Gemini 2.5 Flash 90.6.

打开官方来源

mmlu-pro 81.1 模型 kimi-k2-instruct · 版本 未说明 · 指标 exact_match · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: MMLU-Pro · quote_snippet: MMLU-Pro | EM | 81.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "EM",
  "judge": null
}

Competitor cells: V3-0324 81.2*, Qwen3-235B 77.3, Sonnet 4 83.7, Opus 4 86.6, GPT-4.1 81.8, Gemini 2.5 Flash 79.4.

打开官方来源

ifeval 89.8 模型 kimi-k2-instruct · 版本 strict prompt · 指标 prompt_strict_acc · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: IFEval · quote_snippet: IFEval | Prompt Strict | 89.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Prompt Strict",
  "judge": null
}

Competitor cells: V3-0324 81.1, Qwen3-235B 83.2*, Sonnet 4 87.6, Opus 4 87.4, GPT-4.1 88.0, Gemini 2.5 Flash 84.3.

打开官方来源

multichallenge 54.1 模型 kimi-k2-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: Multi-Challenge · quote_snippet: Multi-Challenge | Acc | 54.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Acc",
  "judge": null
}

new-benchmark: multichallenge already introduced by prior batches, still not in data/benchmarks/. Competitor cells: V3-0324 31.4, Qwen3-235B 34.0, Sonnet 4 46.8, Opus 4 49.0, GPT-4.1 36.4, Gemini 2.5 Flash 39.5.

打开官方来源

simpleqa 31.0 模型 kimi-k2-instruct · 版本 未说明 · 指标 correct_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: SimpleQA · quote_snippet: SimpleQA | Correct | 31.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Correct",
  "judge": null
}

Competitor cells: V3-0324 27.7, Qwen3-235B 13.2, Sonnet 4 15.9, Opus 4 22.8, GPT-4.1 42.3, Gemini 2.5 Flash 23.3.

打开官方来源

livebench 76.4 模型 kimi-k2-instruct · 版本 2024/11/25 snapshot · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: Livebench(2024/11/25) · quote_snippet: Livebench(2024/11/25) | Pass@1 | 76.4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Pass@1",
  "judge": null
}

LiveBench is continuously refreshed; the printed date-pinned snapshot (2024/11/25) is part of the protocol. Competitor cells: V3-0324 72.4, Qwen3-235B 67.6, Sonnet 4 74.8, Opus 4 74.6, GPT-4.1 69.8, Gemini 2.5 Flash 67.8.

打开官方来源

Kimi-K2-Base

Kimi-K2-Base 为 Kimi K2 发布中的基座模型条目。发布页基准表均以 Kimi-K2-Instruct 为评测对象,Base 无独立评测行。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。