← 模型目录

Grok 4.1

xAI / Grok · 2025-11-17 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Grok 4.1

xAI 将 Grok 4.1 定位为一次可用性导向的风格升级:复用驱动 Grok 4 的大规模强化学习基础设施来优化风格、人格、有用性与对齐,使其更敏锐捕捉细微意图、更具对话魅力且人格连贯,同时完全保留前代的锋利智能与可靠性。已收录评测覆盖竞技场 Elo、情商、创意写作与事实性等领域,Text Arena Thinking 配置 1483 Elo(总榜第一)、EQ-Bench3 1586 Elo(归一化)。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

arena 1483 Elo (#1 overall) 模型 grok-4-1 · 版本 Text Arena, Thinking config · 指标 arena_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: LMArena Text Leaderboard · quote_snippet: Grok 4.1 Thinking (code name: quasarflux) holds the #1 overall position with 1483 Elo

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking mode",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard",
  "judge": null
}

Maps to existing benchmark arena (LMArena Text Arena). Page states a 31-point margin over the highest non-xAI model; codename quasarflux is the arena submission identity. Leaderboard Elo is a moving snapshot. Grok 4 by comparison ranked #33 overall per the same page.

打开官方来源

arena 1465 Elo (#2 overall) 模型 grok-4-1 · 版本 Text Arena, non-thinking config · 指标 arena_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: LMArena Text Leaderboard · quote_snippet: Grok 4.1 in its non-reasoning mode (code name: tensor)... ranks #2 at 1465 Elo

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-thinking (no thinking tokens, immediate response)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard",
  "judge": null
}

Page highlights that this NON-thinking configuration surpasses every other model's full-reasoning configuration on the public leaderboard — a cross-effort comparison claim worth flagging when reading the ranking.

打开官方来源

eq-bench3 1586 Elo (normalized) 模型 grok-4-1 · 版本 未说明 · 指标 normalized_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Emotional Intelligence / EQ-Bench · figure: EQ-Bench recharts SVG ("Emotional Intelligence Benchmark - Elo (Normalized)"); numeric labels present in the rendered SVG text of page.html · quote_snippet: EQ-Bench is a LLM-judged test... We report the rubric score and normalized Elo score by running the official benchmark repository

{
  "harness": "official EQ-Bench3 repository",
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking mode",
  "temperature": "default sampling parameters prescribed by the benchmark",
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "rubric score + normalized Elo",
  "judge": "Claude Sonnet 3.7 (prescribed judge)"
}

new-benchmark: eq-bench3 not yet in data/benchmarks/. Values recovered 2026-09-01 from the chart's own SVG text in the archived page.html (not an image OCR path): Grok 4.1 Thinking 1586, Grok 4.1 (non-thinking) 1585, Kimi K2 Instruct 1561, Horizon Alpha 1559, Gemini 2.5 Pro 1460, GPT-5 Chat 1364, Claude Opus 4 1304, Grok 4 1206. Row records the Thinking config; the 1-point non-thinking gap is within Elo noise. Page states scores were computed with default sampling parameters, prescribed judge (Claude Sonnet 3.7) and no system prompt, run from the official EQ-Bench3 repository; xAI notes it is working with the benchmark author to publish leaderboard numbers.

打开官方来源

creative-writing-v3 1721.9 Elo (normalized) 模型 grok-4-1 · 版本 未说明 · 指标 normalized_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Creative Writing / Creative Writing v3 · figure: Creative Writing v3 recharts SVG ("Judging creative writing reliably - Elo (Normalized)"); numeric labels present in the rendered SVG text of page.html · quote_snippet: models generate responses to 32 distinct writing prompts across 3 iterations

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking mode",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "rubrics + model battle normalized Elo",
  "judge": null
}

new-benchmark: creative-writing-v3 not yet in data/benchmarks/. Values recovered 2026-09-01 from the chart's own SVG text in the archived page.html: Polaris Alpha (early GPT 5.1) 1756.2 / Grok 4.1 Thinking 1721.9 / Grok 4.1 (non-thinking) 1708.6 / o3 1696.4 / Claude Sonnet 4.5 1648.7 / Kimi K2 Instruct 1627.5 / Grok 3 1126. Row records the Thinking config. Scores computed via rubrics + model-battle normalized Elo (same machinery as EQ-Bench3).

打开官方来源

factscore 2.97% 模型 grok-4-1 · 版本 未说明 · 指标 hallucination_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Reduced Hallucinations / FActScore · figure: FActScore recharts SVG (two-bar chart with % labels) plus Grok 4 Fast (Non-Reasoning) / Grok 4.1 (Non-Reasoning) legend; labels present in the rendered SVG text of page.html · quote_snippet: We also evaluate FActScore, which is a public benchmark consisting of 500 biography questions on individuals

{
  "harness": null,
  "tools": [
    "web search"
  ],
  "shots": null,
  "reasoning_effort": "non-reasoning (Fast)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: factscore not yet in data/benchmarks/. Direction flag: LOWER is better per the page's own "Lower score is better" subtitle. Values recovered 2026-09-01 from the chart's SVG text in the archived page.html: Grok 4.1 (Non-Reasoning) 2.97% vs baseline Grok 4 Fast (Non-Reasoning) 9.89%; the paired Hallucination Rate chart on the same grid shows 4.22% vs 12.09% on internal production-traffic queries (internal metric, kept in notes only). Methodology per prose: non-reasoning model evaluated with web search tools; hallucination rate = macro-average of atomic claims with major/minor errors.

打开官方来源