Gemini 3.1 Flash-Lite
Google DeepMind / Gemini · 2026-05-08 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Gemini 3.1 Flash-Lite
Google 将 Gemini 3.1 Flash-Lite 定位为 Gemini 3 系列中最快、最具成本效率的型号,2026-03-03 预览、2026-05-07/08 转正(GA)。已收录 3 项评测(均出自预览公告)覆盖对话竞技场、科学与多模态推理:亮点 Arena.ai 排行 1432 Elo、GPQA 86.9%。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 0.25 / 输出 1.5 · 定价出自 2026-03-03 预览公告,GA 页未重列
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Cost-efficiency without compromise · quote_snippet: 3.1 Flash-Lite achieves an impressive Elo score of 1432 on the Arena.ai Leaderboard
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard",
"judge": null
}Maps to existing benchmark arena ('Arena.ai Leaderboard' — the arena.ai identity of LMArena). Elo snapshot at preview time.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Cost-efficiency without compromise · quote_snippet: outperforms other models of similar tier across reasoning and multimodal understanding benchmarks, including 86.9% on GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark gpqa. This is the number that had circulated only in secondary coverage during batch 5 — now grounded on the official preview post.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Cost-efficiency without compromise · quote_snippet: and 76.8% on MMMU Pro–even surpassing larger Gemini models from prior generations like 2.5 Flash
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark mmmu variant Pro. Context: 3.1 Pro's MMMU-Pro is 80.5 — the Flash-Lite tier sits 3.7pp under the Pro tier on this benchmark.