跳到主要内容
Google · 发布于 2025-06-24

Gemini 2.5 Flash Lite LLM

上下文 1M 最大输出 64K 类型 LLM 渠道 OpenRouter
google/gemini-2.5-flash-lite

Gemini 2.5 Flash-Lite is the lightweight member of the Gemini 2.5 family, optimized for ultra-fast responses, high throughput, and cost-efficient AI workloads. It delivers improved performance, faster token generation, and greater efficiency compared with previous Flash-class models. By default, thinking capabilities (multi-step reasoning) are disabled to maximize speed and minimize cost. Developers can enable reasoning through the Reasoning API parameter when deeper intelligence is required, providing flexible control over the balance between latency, capability, and resource usage.

定价

USD · 实时汇率
定价 USD / M
输入 $0.10
输出 $0.40
命中缓存 $0.01
图片输入

供应商

同一模型,多渠道实时比价 · 最优值绿色高亮
供应商 上下文 最大输出 输入 /M 输出 /M 缓存 /M 延迟 吞吐
OpenRouter 1M 64K $0.10 $0.40 $0.01 3653ms 2140 t/s
Google Vertex 1M 64K $0.10 $0.40 $0.03 956ms 306 t/s

接入示例

Model Center 会在不同服务提供商之间规范化处理请求和响应,为你统一接口。

提供兼容 OpenAI 的 Completion API:既可直接调用,也可通过 OpenAI SDK 调用,一把 key 调用全部模型。

获取 API Key 获取 API Key
from openai import OpenAI

client = OpenAI(
    base_url="https://router-integration.test.cogfoundry.ai/api/v1",
    api_key="$MODEL_CENTER_API_KEY",
)

completion = client.chat.completions.create(
    model="google/gemini-2.5-flash-lite",
    messages=[{"role": "user", "content": "9.11 和 9.8 哪个大?"}],
    stream=True,
)

for chunk in completion:
    if chunk.choices and chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="", flush=True)

Google 的更多模型