Kimi-K2-Instruct / Kimi-K2-Base
Moonshot AI / Kimi · 2025-07-11 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Kimi-K2-Instruct
Kimi K2 以「Open Agentic Intelligence」为定位发布,Kimi-K2-Instruct 为后训练的非思考(reflex-grade)MoE 模型(1T 总参/32B 激活)。已收录 32 项评测覆盖代码、数学、知识与 Agent 工具使用:亮点 SWE-bench Verified(Agentic Coding)65.8、MMLU 89.5。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 1T-A32B MoE
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table under 'Benchmarking Kimi K2' (only table on page), Coding Tasks section · row: LiveCodeBench v6(Aug 24-May 25) · quote_snippet: LiveCodeBench v6(Aug 24-May 25) | Pass@1 | 53.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Pass@1",
"judge": null
}Competitor cells same row: DeepSeek-V3-0324 46.9, Qwen3-235B-A22B (non-thinking) 37.0, Claude Sonnet 4 (w/o extended thinking) 48.5, Claude Opus 4 (w/o extended thinking) 47.4, GPT-4.1 44.7, Gemini 2.5 Flash Preview (05-20) 44.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: OJBench · quote_snippet: OJBench | Pass@1 | 27.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Pass@1",
"judge": null
}new-benchmark: ojbench not yet in data/benchmarks/. Competitor cells: V3-0324 24.0, Qwen3-235B 11.3, Sonnet 4 15.3, Opus 4 19.6, GPT-4.1 19.5, Gemini 2.5 Flash 19.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: MultiPL-E · quote_snippet: MultiPL-E | Pass@1 | 85.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Pass@1",
"judge": null
}new-benchmark: multipl-e not yet in data/benchmarks/. Competitor cells: V3-0324 83.1, Qwen3-235B 78.2, Sonnet 4 88.6, Opus 4 89.6, GPT-4.1 86.7, Gemini 2.5 Flash 85.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: SWE-bench Verified (Agentless Coding) · quote_snippet: SWE-bench Verified (Agentless Coding) | Single Patch without Test (Acc) | 51.8
{
"harness": "Agentless (single patch without test)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Single Patch without Test (Acc)",
"judge": null
}Three separate SWE-bench Verified protocols printed as distinct table rows; do not merge. Competitor cells: V3-0324 36.6, Qwen3-235B 39.4, Sonnet 4 50.2, Opus 4 53.0, GPT-4.1 40.8, Gemini 2.5 Flash 32.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: SWE-bench Verified (Agentic Coding) / Single Attempt (Acc) · quote_snippet: SWE-bench Verified (Agentic Coding) | Single Attempt (Acc) | 65.8
{
"harness": "agentic (unspecified)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "Single Attempt (Acc)",
"judge": null
}Competitor cells carry asterisk (72.7* Sonnet 4, 72.5* Opus 4): the page prints no footnote explaining the asterisk semantics; treat those cells as vendor-reported-by-competitor values of unknown protocol. GPT-4.1 54.6, V3-0324 38.8, Qwen3-235B 34.4, Gemini 2.5 Flash not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: SWE-bench Verified (Agentic Coding) / Multiple Attempts (Acc) · quote_snippet: SWE-bench Verified (Agentic Coding) | Multiple Attempts (Acc) | 71.6
{
"harness": "agentic (unspecified)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Multiple Attempts (Acc)",
"judge": null
}Multiple-attempt aggregation is not pass@1 and is not comparable to single-attempt rows. Competitor cells (asterisked): Sonnet 4 80.2*, Opus 4 79.4*; all other columns not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: SWE-bench Multilingual(Agentic Coding) · quote_snippet: SWE-bench Multilingual(Agentic Coding) | Single Attempt (Acc) | 47.3
{
"harness": "agentic (unspecified)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "Single Attempt (Acc)",
"judge": null
}new-benchmark: swebench-multilingual not yet in data/benchmarks/ (id already used by prior batches' notes). Competitor cells: V3-0324 25.8, Qwen3-235B 20.9, Sonnet 4 51.0, GPT-4.1 31.5; Opus 4 and Gemini 2.5 Flash not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: TerminalBench / Inhouse Framework (Acc) · quote_snippet: TerminalBench | Inhouse Framework (Acc) | 30.0
{
"harness": "in-house framework (unspecified)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}Two TerminalBench harnesses printed as distinct rows; not comparable to each other. Competitor cells: Sonnet 4 35.5, Opus 4 43.2, GPT-4.1 8.3; V3-0324 / Qwen3-235B / Gemini 2.5 Flash not reported.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: TerminalBench / Terminus (Acc) · quote_snippet: TerminalBench | Terminus (Acc) | 25.0
{
"harness": "Terminus",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}Competitor cells: V3-0324 16.3, Qwen3-235B 6.6, GPT-4.1 30.3, Gemini 2.5 Flash 16.8; Claude columns not reported. Terminus is a third-party public harness, enabling cross-vendor comparison only when the harness version matches.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Coding Tasks section · row: Aider-Polyglot · quote_snippet: Aider-Polyglot | Acc | 60.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}Competitor cells: V3-0324 55.1, Qwen3-235B 61.8, Sonnet 4 56.4, Opus 4 70.7, GPT-4.1 52.4, Gemini 2.5 Flash 44.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Tool Use Tasks section · row: Tau2 retail · quote_snippet: Tau2 retail | Avg@4 | 70.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 4,
"aggregation": "Avg@4",
"judge": null
}Competitor cells: V3-0324 69.1, Qwen3-235B 57.0, Sonnet 4 75.0, Opus 4 81.8, GPT-4.1 74.8, Gemini 2.5 Flash 64.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Tool Use Tasks section · row: Tau2 airline · quote_snippet: Tau2 airline | Avg@4 | 56.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 4,
"aggregation": "Avg@4",
"judge": null
}Competitor cells: V3-0324 39.0, Qwen3-235B 26.5, Sonnet 4 55.5, Opus 4 60.0, GPT-4.1 54.5, Gemini 2.5 Flash 42.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Tool Use Tasks section · row: Tau2 telecom · quote_snippet: Tau2 telecom | Avg@4 | 65.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 4,
"aggregation": "Avg@4",
"judge": null
}Competitor cells: V3-0324 32.5, Qwen3-235B 22.1, Sonnet 4 45.2, Opus 4 57.0, GPT-4.1 38.6, Gemini 2.5 Flash 16.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Tool Use Tasks section · row: AceBench · quote_snippet: AceBench | Acc | 76.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}new-benchmark: acebench not yet in data/benchmarks/. Competitor cells: V3-0324 72.7, Qwen3-235B 70.5, Sonnet 4 76.2, Opus 4 75.6, GPT-4.1 80.1, Gemini 2.5 Flash 74.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: AIME 2024 · quote_snippet: AIME 2024 | Avg@64 | 69.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 64,
"aggregation": "Avg@64",
"judge": null
}Avg@64 aggregation differs from single-run AIME scores; not comparable. Competitor cells: V3-0324 59.4*, Qwen3-235B 40.1*, Sonnet 4 43.4, Opus 4 48.2, GPT-4.1 46.5, Gemini 2.5 Flash 61.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: AIME 2025 · quote_snippet: AIME 2025 | Avg@64 | 49.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 64,
"aggregation": "Avg@64",
"judge": null
}Competitor cells: V3-0324 46.7, Qwen3-235B 24.7*, Sonnet 4 33.1*, Opus 4 33.9*, GPT-4.1 37.0, Gemini 2.5 Flash 46.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: MATH-500 · quote_snippet: MATH-500 | Acc | 97.4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}Competitor cells: V3-0324 94.0*, Qwen3-235B 91.2*, Sonnet 4 94.0, Opus 4 94.4, GPT-4.1 92.4, Gemini 2.5 Flash 95.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: HMMT 2025 · quote_snippet: HMMT 2025 | Avg@32 | 38.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 32,
"aggregation": "Avg@32",
"judge": null
}new-benchmark: hmmt-25 already introduced by prior batches, still not in data/benchmarks/. Competitor cells: V3-0324 27.5, Qwen3-235B 11.9, Sonnet 4 15.9, Opus 4 15.9, GPT-4.1 19.4, Gemini 2.5 Flash 34.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: CNMO 2024 · quote_snippet: CNMO 2024 | Avg@16 | 74.3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 16,
"aggregation": "Avg@16",
"judge": null
}new-benchmark: cnmo-2024 not yet in data/benchmarks/. Competitor cells: V3-0324 74.7, Qwen3-235B 48.6, Sonnet 4 60.4, Opus 4 57.6, GPT-4.1 56.6, Gemini 2.5 Flash 75.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: PolyMath-en · quote_snippet: PolyMath-en | Avg@4 | 65.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 4,
"aggregation": "Avg@4",
"judge": null
}new-benchmark: polymath-en not yet in data/benchmarks/. Competitor cells: V3-0324 59.5, Qwen3-235B 51.9, Sonnet 4 52.8, Opus 4 49.8, GPT-4.1 54.0, Gemini 2.5 Flash 49.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: ZebraLogic · quote_snippet: ZebraLogic | Acc | 89.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}new-benchmark: zebralogic not yet in data/benchmarks/. Competitor cells: V3-0324 84.0, Qwen3-235B 37.7*, Sonnet 4 79.7, Opus 4 59.3, GPT-4.1 58.5, Gemini 2.5 Flash 57.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: AutoLogi · quote_snippet: AutoLogi | Acc | 89.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}new-benchmark: autologi not yet in data/benchmarks/. Competitor cells: V3-0324 88.9, Qwen3-235B 83.3*, Sonnet 4 89.8, Opus 4 86.1, GPT-4.1 88.2, Gemini 2.5 Flash 84.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: GPQA-Diamond · quote_snippet: GPQA-Diamond | Avg@8 | 75.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 8,
"aggregation": "Avg@8",
"judge": null
}Competitor cells: V3-0324 68.4*, Qwen3-235B 62.9*, Sonnet 4 70.0*, Opus 4 74.9*, GPT-4.1 66.3, Gemini 2.5 Flash 68.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: SuperGPQA · quote_snippet: SuperGPQA | Acc | 57.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}new-benchmark: supergpqa already introduced by prior batches, still not in data/benchmarks/. Competitor cells: V3-0324 53.7, Qwen3-235B 50.2, Sonnet 4 55.7, Opus 4 56.5, GPT-4.1 50.8, Gemini 2.5 Flash 49.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, Math & STEM Tasks section · row: Humanity's Last Exam (Text Only) · quote_snippet: Humanity's Last Exam (Text Only) | Acc | 4.7
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}Text-only subset, no tools; not comparable to HLE w/ tools rows in other releases (e.g. K2 Thinking 44.9). Competitor cells: V3-0324 5.2, Qwen3-235B 5.7, Sonnet 4 5.8, Opus 4 7.1, GPT-4.1 3.7, Gemini 2.5 Flash 5.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: MMLU · quote_snippet: MMLU | EM | 89.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "EM",
"judge": null
}Competitor cells: V3-0324 89.4, Qwen3-235B 87.0, Sonnet 4 91.5, Opus 4 92.9, GPT-4.1 90.4, Gemini 2.5 Flash 90.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: MMLU-Redux · quote_snippet: MMLU-Redux | EM | 92.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "EM",
"judge": null
}Competitor cells: V3-0324 90.5, Qwen3-235B 89.2*, Sonnet 4 93.6, Opus 4 94.2, GPT-4.1 92.4, Gemini 2.5 Flash 90.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: MMLU-Pro · quote_snippet: MMLU-Pro | EM | 81.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "EM",
"judge": null
}Competitor cells: V3-0324 81.2*, Qwen3-235B 77.3, Sonnet 4 83.7, Opus 4 86.6, GPT-4.1 81.8, Gemini 2.5 Flash 79.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: IFEval · quote_snippet: IFEval | Prompt Strict | 89.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Prompt Strict",
"judge": null
}Competitor cells: V3-0324 81.1, Qwen3-235B 83.2*, Sonnet 4 87.6, Opus 4 87.4, GPT-4.1 88.0, Gemini 2.5 Flash 84.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: Multi-Challenge · quote_snippet: Multi-Challenge | Acc | 54.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Acc",
"judge": null
}new-benchmark: multichallenge already introduced by prior batches, still not in data/benchmarks/. Competitor cells: V3-0324 31.4, Qwen3-235B 34.0, Sonnet 4 46.8, Opus 4 49.0, GPT-4.1 36.4, Gemini 2.5 Flash 39.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: SimpleQA · quote_snippet: SimpleQA | Correct | 31.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Correct",
"judge": null
}Competitor cells: V3-0324 27.7, Qwen3-235B 13.2, Sonnet 4 15.9, Opus 4 22.8, GPT-4.1 42.3, Gemini 2.5 Flash 23.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarking Kimi K2 · table: DOM table, General Tasks section · row: Livebench(2024/11/25) · quote_snippet: Livebench(2024/11/25) | Pass@1 | 76.4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Pass@1",
"judge": null
}LiveBench is continuously refreshed; the printed date-pinned snapshot (2024/11/25) is part of the protocol. Competitor cells: V3-0324 72.4, Qwen3-235B 67.6, Sonnet 4 74.8, Opus 4 74.6, GPT-4.1 69.8, Gemini 2.5 Flash 67.8.
Kimi-K2-Base
Kimi-K2-Base 为 Kimi K2 发布中的基座模型条目。发布页基准表均以 Kimi-K2-Instruct 为评测对象,Base 无独立评测行。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。