Grok 4.1
xAI / Grok · 2025-11-17 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Grok 4.1
xAI 将 Grok 4.1 定位为一次可用性导向的风格升级:复用驱动 Grok 4 的大规模强化学习基础设施来优化风格、人格、有用性与对齐,使其更敏锐捕捉细微意图、更具对话魅力且人格连贯,同时完全保留前代的锋利智能与可靠性。已收录评测覆盖竞技场 Elo、情商、创意写作与事实性等领域,Text Arena Thinking 配置 1483 Elo(总榜第一)、EQ-Bench3 1586 Elo(归一化)。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: LMArena Text Leaderboard · quote_snippet: Grok 4.1 Thinking (code name: quasarflux) holds the #1 overall position with 1483 Elo
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking mode",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard",
"judge": null
}Maps to existing benchmark arena (LMArena Text Arena). Page states a 31-point margin over the highest non-xAI model; codename quasarflux is the arena submission identity. Leaderboard Elo is a moving snapshot. Grok 4 by comparison ranked #33 overall per the same page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: LMArena Text Leaderboard · quote_snippet: Grok 4.1 in its non-reasoning mode (code name: tensor)... ranks #2 at 1465 Elo
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-thinking (no thinking tokens, immediate response)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard",
"judge": null
}Page highlights that this NON-thinking configuration surpasses every other model's full-reasoning configuration on the public leaderboard — a cross-effort comparison claim worth flagging when reading the ranking.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Emotional Intelligence / EQ-Bench · figure: EQ-Bench recharts SVG ("Emotional Intelligence Benchmark - Elo (Normalized)"); numeric labels present in the rendered SVG text of page.html · quote_snippet: EQ-Bench is a LLM-judged test... We report the rubric score and normalized Elo score by running the official benchmark repository
{
"harness": "official EQ-Bench3 repository",
"tools": null,
"shots": null,
"reasoning_effort": "Thinking mode",
"temperature": "default sampling parameters prescribed by the benchmark",
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "rubric score + normalized Elo",
"judge": "Claude Sonnet 3.7 (prescribed judge)"
}new-benchmark: eq-bench3 not yet in data/benchmarks/. Values recovered 2026-09-01 from the chart's own SVG text in the archived page.html (not an image OCR path): Grok 4.1 Thinking 1586, Grok 4.1 (non-thinking) 1585, Kimi K2 Instruct 1561, Horizon Alpha 1559, Gemini 2.5 Pro 1460, GPT-5 Chat 1364, Claude Opus 4 1304, Grok 4 1206. Row records the Thinking config; the 1-point non-thinking gap is within Elo noise. Page states scores were computed with default sampling parameters, prescribed judge (Claude Sonnet 3.7) and no system prompt, run from the official EQ-Bench3 repository; xAI notes it is working with the benchmark author to publish leaderboard numbers.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Creative Writing / Creative Writing v3 · figure: Creative Writing v3 recharts SVG ("Judging creative writing reliably - Elo (Normalized)"); numeric labels present in the rendered SVG text of page.html · quote_snippet: models generate responses to 32 distinct writing prompts across 3 iterations
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking mode",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "rubrics + model battle normalized Elo",
"judge": null
}new-benchmark: creative-writing-v3 not yet in data/benchmarks/. Values recovered 2026-09-01 from the chart's own SVG text in the archived page.html: Polaris Alpha (early GPT 5.1) 1756.2 / Grok 4.1 Thinking 1721.9 / Grok 4.1 (non-thinking) 1708.6 / o3 1696.4 / Claude Sonnet 4.5 1648.7 / Kimi K2 Instruct 1627.5 / Grok 3 1126. Row records the Thinking config. Scores computed via rubrics + model-battle normalized Elo (same machinery as EQ-Bench3).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Reduced Hallucinations / FActScore · figure: FActScore recharts SVG (two-bar chart with % labels) plus Grok 4 Fast (Non-Reasoning) / Grok 4.1 (Non-Reasoning) legend; labels present in the rendered SVG text of page.html · quote_snippet: We also evaluate FActScore, which is a public benchmark consisting of 500 biography questions on individuals
{
"harness": null,
"tools": [
"web search"
],
"shots": null,
"reasoning_effort": "non-reasoning (Fast)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: factscore not yet in data/benchmarks/. Direction flag: LOWER is better per the page's own "Lower score is better" subtitle. Values recovered 2026-09-01 from the chart's SVG text in the archived page.html: Grok 4.1 (Non-Reasoning) 2.97% vs baseline Grok 4 Fast (Non-Reasoning) 9.89%; the paired Hallucination Rate chart on the same grid shows 4.22% vs 12.09% on internal production-traffic queries (internal metric, kept in notes only). Methodology per prose: non-reasoning model evaluated with web search tools; hallucination rate = macro-average of atomic claims with major/minor errors.