跳到主要内容

模型市场

11 个模型 · 一把 key 实时调用 · 按 token 计费,无月费

$0.10/M
最低输入价
1M
最大上下文
2
收录厂商

11 / 11 个模型

  • Gemini 3.5 Flash is Google’s fast and efficient multimodal model, combining advanced coding, reasoning, and agentic capabilities with the speed and affordability of the Flash family. Designed for scalable AI workflows, it supports text, image, video, audio, and PDF inputs, making it ideal for coding assistants, automation, and multi-agent applications. With medium thinking effort enabled by default, Gemini 3.5 Flash delivers an optimized balance between intelligence, speed, and cost. Developers can adjust thinking levels from minimal to high to fine-tune performance for different use cases.

    输入
    $1.50/M
    输出
    $9.00/M
    上下文
    1M
    1.3M tok
    google/gemini-3.5-flash
  • Gemini 3.1 Flash Lite is Google’s high-efficiency multimodal model, optimized for low latency, high throughput, and cost-effective AI workloads. It supports text, images, video, audio, and PDF inputs, making it ideal for lightweight agentic workflows, data extraction, content processing, and other high-volume applications where speed and cost are critical. The model supports configurable thinking levels (minimal, low, medium, and high), allowing developers to balance intelligence, latency, and cost for different workloads. At half the price of Gemini 3 Flash, Gemini 3.1 Flash Lite delivers excellent value for scalable production deployments.

    输入
    $0.25/M
    输出
    $1.50/M
    上下文
    1M
    641.4K tok
    google/gemini-3.1-flash-lite
  • Gemini 3.1 Flash Lite Preview is Google’s high-efficiency multimodal model, optimized for high-throughput, cost-sensitive AI workloads. It delivers higher overall quality than Gemini 2.5 Flash Lite and approaches Gemini 2.5 Flash across many core capabilities, while maintaining low latency and excellent cost efficiency. The model provides notable improvements in audio understanding and speech recognition (ASR), RAG retrieval and snippet ranking, translation, data extraction, and code completion. It supports configurable thinking levels (minimal, low, medium, and high) for flexible control over intelligence, latency, and cost. At half the price of Gemini 3 Flash, Gemini 3.1 Flash Lite Preview is an excellent choice for scalable production and high-volume AI applications.

    输入
    $0.25/M
    输出
    $1.50/M
    上下文
    1M
    对话与客服内容创作文档与办公
    google/gemini-3.1-flash-lite-preview
  • Gemini 3.1 Pro Preview is Google’s advanced reasoning model, designed for complex coding, agentic workflows, and high-context problem solving. It delivers improved software engineering performance, stronger agent reliability, and better token efficiency across demanding tasks. Built on the Gemini 3 multimodal foundation, it supports advanced reasoning across text, images, video, audio, and code, with a 1 million token context window for processing large-scale information and long-running workflows. It delivers improved performance on software engineering benchmarks and real-world coding environments, while enabling more reliable autonomous execution in structured domains such as finance and spreadsheet automation. Gemini 3.1 Pro Preview introduces a medium thinking level to provide better control over the trade-off between intelligence, latency, and cost. With stronger long-horizon reasoning, tool orchestration, and workflow automation capabilities, it is optimized for AI agents, coding assistants, multimodal analysis, enterprise automation, and complex data-driven applications.

    输入
    $2.00/M
    输出
    $12.00/M
    上下文
    1M
    编程智能代理与数据多模态理解
    google/gemini-3.1-pro-preview
  • Gemini 3 Flash Preview is Google’s fast, efficient thinking model for agentic workflows, coding, and interactive AI applications. It delivers near-Pro-level reasoning and tool-use performance with lower latency and cost, making it ideal for multi-turn conversations, autonomous agents, and collaborative development. With a 1 million token context window, multimodal input support (text, images, audio, video, and PDFs), configurable thinking levels, structured outputs, tool calling, and context caching, Gemini 3 Flash Preview provides a strong balance of intelligence, speed, and efficiency.

    输入
    $0.50/M
    输出
    $3.00/M
    上下文
    1M
    智能代理与数据编程对话与客服 1.1M tok
    google/gemini-3-flash
  • Gemini 2.5 Flash is Google’s efficient, high-performance AI model built for demanding reasoning, coding, mathematics, and scientific workloads. Its integrated thinking capabilities enable stronger problem solving, improved accuracy, and better handling of complex multi-step tasks. Developers can fine-tune reasoning depth using the “max tokens for reasoning” setting, providing flexible control over the trade-off between intelligence, speed, and cost for different AI applications.

    输入
    $0.30/M
    输出
    $2.50/M
    上下文
    1M
    多模态理解对话与客服编程 241K tok
    google/gemini-2.5-flash
  • Gemini 2.5 Pro is Google’s advanced AI model built for complex reasoning, software development, mathematics, and scientific applications. With integrated thinking capabilities, it performs deeper multi-step reasoning to deliver more accurate answers, stronger problem solving, and better understanding of complex contexts. Gemini 2.5 Pro achieves leading performance across a range of industry benchmarks, demonstrating strong capabilities in coding, reasoning, and complex task execution. Its advanced intelligence and reliability make it well suited for demanding applications such as AI agents, research workflows, and enterprise-grade automation.

    输入
    $1.25/M
    输出
    $10.00/M
    上下文
    1M
    编程多模态理解编程 81.3K tok
    google/gemini-2.5-pro
  • Gemini 2.5 Flash-Lite is the lightweight member of the Gemini 2.5 family, optimized for ultra-fast responses, high throughput, and cost-efficient AI workloads. It delivers improved performance, faster token generation, and greater efficiency compared with previous Flash-class models. By default, thinking capabilities (multi-step reasoning) are disabled to maximize speed and minimize cost. Developers can enable reasoning through the Reasoning API parameter when deeper intelligence is required, providing flexible control over the balance between latency, capability, and resource usage.

    输入
    $0.10/M
    输出
    $0.40/M
    上下文
    1M
    多模态理解编程数学与科学 1.3M tok
    google/gemini-2.5-flash-lite