跳到主要内容
Google · 发布于 2026-05-18

Gemini 3.1 Flash Lite LLM

上下文 1M 最大输出 66K 类型 LLM 渠道 Google Vertex
google/gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s high-efficiency multimodal model, optimized for low latency, high throughput, and cost-effective AI workloads. It supports text, images, video, audio, and PDF inputs, making it ideal for lightweight agentic workflows, data extraction, content processing, and other high-volume applications where speed and cost are critical. The model supports configurable thinking levels (minimal, low, medium, and high), allowing developers to balance intelligence, latency, and cost for different workloads. At half the price of Gemini 3 Flash, Gemini 3.1 Flash Lite delivers excellent value for scalable production deployments.

定价

USD · 实时汇率
定价 USD / M
输入 $0.25
输出 $1.50
命中缓存 $0.03
图片输入

供应商

同一模型,多渠道实时比价 · 最优值绿色高亮
供应商 上下文 最大输出 输入 /M 输出 /M 缓存 /M 延迟 吞吐
Google Vertex 1M 66K $0.25 $1.50 $0.03 2120ms 197 t/s

接入示例

Model Center 会在不同服务提供商之间规范化处理请求和响应,为你统一接口。

提供兼容 OpenAI 的 Completion API:既可直接调用,也可通过 OpenAI SDK 调用,一把 key 调用全部模型。

获取 API Key 获取 API Key
from openai import OpenAI

client = OpenAI(
    base_url="https://router-integration.test.cogfoundry.ai/api/v1",
    api_key="$MODEL_CENTER_API_KEY",
)

completion = client.chat.completions.create(
    model="google/gemini-3.1-flash-lite",
    messages=[{"role": "user", "content": "9.11 和 9.8 哪个大?"}],
    stream=True,
)

for chunk in completion:
    if chunk.choices and chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="", flush=True)

Google 的更多模型