跳到主要内容

模型市场

11 个模型 · 一把 key 实时调用 · 按 token 计费,无月费

$0.10/M
最低输入价
1M
最大上下文
2
收录厂商

11 / 11 个模型

  • DeepSeek

    DeepSeek V4 Pro (75% off) 75% OFF

    DeepSeek V4 Pro is DeepSeek's flagship Mixture-of-Experts (MoE) model, featuring 1.6 trillion total parameters, 49 billion activated parameters, and a 1 million-token context window. It delivers exceptional performance in reasoning, coding, mathematics, and software engineering, making it ideal for demanding AI workloads. Built on the same architecture as DeepSeek V4 Flash, it adds a hybrid attention system for more efficient long-context processing. It supports High and XHigh reasoning modes (with XHigh corresponding to maximum reasoning), making it well suited for complex tasks such as full codebase analysis, multi-step agent workflows, large-scale information synthesis, and enterprise-grade automation.

    输入
    $0.43/M $1.71
    输出
    $0.86/M $3.43
    上下文
    1M
    编程智能代理与数据数学与科学 322.7K tok
    deepseek/deepseek-v4-pro
  • Gemini 3.5 Flash is Google’s fast and efficient multimodal model, combining advanced coding, reasoning, and agentic capabilities with the speed and affordability of the Flash family. Designed for scalable AI workflows, it supports text, image, video, audio, and PDF inputs, making it ideal for coding assistants, automation, and multi-agent applications. With medium thinking effort enabled by default, Gemini 3.5 Flash delivers an optimized balance between intelligence, speed, and cost. Developers can adjust thinking levels from minimal to high to fine-tune performance for different use cases.

    输入
    $1.50/M
    输出
    $9.00/M
    上下文
    1M
    1.3M tok
    google/gemini-3.5-flash
  • Gemini 3.1 Flash Lite is Google’s high-efficiency multimodal model, optimized for low latency, high throughput, and cost-effective AI workloads. It supports text, images, video, audio, and PDF inputs, making it ideal for lightweight agentic workflows, data extraction, content processing, and other high-volume applications where speed and cost are critical. The model supports configurable thinking levels (minimal, low, medium, and high), allowing developers to balance intelligence, latency, and cost for different workloads. At half the price of Gemini 3 Flash, Gemini 3.1 Flash Lite delivers excellent value for scalable production deployments.

    输入
    $0.25/M
    输出
    $1.50/M
    上下文
    1M
    641.4K tok
    google/gemini-3.1-flash-lite
  • DeepSeek V4 Flash is DeepSeek's high-speed, cost-efficient Mixture-of-Experts (MoE) model, featuring 284 billion total parameters, 13 billion activated parameters, and a 1 million-token context window. It delivers fast inference, high throughput, and strong performance in reasoning, coding, and general-purpose AI tasks. Built with a hybrid attention architecture for efficient long-context processing, it supports High and XHigh reasoning modes (with XHigh corresponding to maximum reasoning). It is ideal for coding assistants, conversational AI, real-time applications, and agent workflows where speed, scalability, and cost efficiency are essential.

    输入
    $0.14/M
    输出
    $0.29/M $0.43
    上下文
    1M
    数学与科学智能代理与数据文档与办公 481.9K tok
    deepseek/deepseek-v4-flash
  • Gemini 3.1 Flash Lite Preview is Google’s high-efficiency multimodal model, optimized for high-throughput, cost-sensitive AI workloads. It delivers higher overall quality than Gemini 2.5 Flash Lite and approaches Gemini 2.5 Flash across many core capabilities, while maintaining low latency and excellent cost efficiency. The model provides notable improvements in audio understanding and speech recognition (ASR), RAG retrieval and snippet ranking, translation, data extraction, and code completion. It supports configurable thinking levels (minimal, low, medium, and high) for flexible control over intelligence, latency, and cost. At half the price of Gemini 3 Flash, Gemini 3.1 Flash Lite Preview is an excellent choice for scalable production and high-volume AI applications.

    输入
    $0.25/M
    输出
    $1.50/M
    上下文
    1M
    对话与客服内容创作文档与办公
    google/gemini-3.1-flash-lite-preview
  • Gemini 3.1 Pro Preview is Google’s advanced reasoning model, designed for complex coding, agentic workflows, and high-context problem solving. It delivers improved software engineering performance, stronger agent reliability, and better token efficiency across demanding tasks. Built on the Gemini 3 multimodal foundation, it supports advanced reasoning across text, images, video, audio, and code, with a 1 million token context window for processing large-scale information and long-running workflows. It delivers improved performance on software engineering benchmarks and real-world coding environments, while enabling more reliable autonomous execution in structured domains such as finance and spreadsheet automation. Gemini 3.1 Pro Preview introduces a medium thinking level to provide better control over the trade-off between intelligence, latency, and cost. With stronger long-horizon reasoning, tool orchestration, and workflow automation capabilities, it is optimized for AI agents, coding assistants, multimodal analysis, enterprise automation, and complex data-driven applications.

    输入
    $2.00/M
    输出
    $12.00/M
    上下文
    1M
    编程智能代理与数据多模态理解
    google/gemini-3.1-pro-preview
  • Gemini 3 Flash Preview is Google’s fast, efficient thinking model for agentic workflows, coding, and interactive AI applications. It delivers near-Pro-level reasoning and tool-use performance with lower latency and cost, making it ideal for multi-turn conversations, autonomous agents, and collaborative development. With a 1 million token context window, multimodal input support (text, images, audio, video, and PDFs), configurable thinking levels, structured outputs, tool calling, and context caching, Gemini 3 Flash Preview provides a strong balance of intelligence, speed, and efficiency.

    输入
    $0.50/M
    输出
    $3.00/M
    上下文
    1M
    智能代理与数据编程对话与客服 1.1M tok
    google/gemini-3-flash
  • Gemini 2.5 Flash is Google’s efficient, high-performance AI model built for demanding reasoning, coding, mathematics, and scientific workloads. Its integrated thinking capabilities enable stronger problem solving, improved accuracy, and better handling of complex multi-step tasks. Developers can fine-tune reasoning depth using the “max tokens for reasoning” setting, providing flexible control over the trade-off between intelligence, speed, and cost for different AI applications.

    输入
    $0.30/M
    输出
    $2.50/M
    上下文
    1M
    多模态理解对话与客服编程 241K tok
    google/gemini-2.5-flash
  • Gemini 2.5 Pro is Google’s advanced AI model built for complex reasoning, software development, mathematics, and scientific applications. With integrated thinking capabilities, it performs deeper multi-step reasoning to deliver more accurate answers, stronger problem solving, and better understanding of complex contexts. Gemini 2.5 Pro achieves leading performance across a range of industry benchmarks, demonstrating strong capabilities in coding, reasoning, and complex task execution. Its advanced intelligence and reliability make it well suited for demanding applications such as AI agents, research workflows, and enterprise-grade automation.

    输入
    $1.25/M
    输出
    $10.00/M
    上下文
    1M
    编程多模态理解编程 81.3K tok
    google/gemini-2.5-pro
  • DeepSeek

    DeepSeek V3.1

    DeepSeek V3.1 is DeepSeek's advanced hybrid reasoning model, featuring 671 billion total parameters, 37 billion activated parameters, and a 128K-token context window. It supports both reasoning and non-reasoning modes, allowing developers to optimize for either response quality or speed depending on the task. The model delivers major improvements in reasoning, coding, tool use, and agent workflows, with performance approaching DeepSeek R1 while providing faster responses and more efficient inference. It supports structured tool calling, code agents, search agents, and complex multi-step workflows, making it an excellent choice for research, software development, enterprise automation, and general-purpose AI applications.

    输入
    $0.57/M
    输出
    $1.71/M
    上下文
    128K
    智能代理与数据多模态理解数学与科学 22 tok
    deepseek/deepseek-v3.1
  • Gemini 2.5 Flash-Lite is the lightweight member of the Gemini 2.5 family, optimized for ultra-fast responses, high throughput, and cost-efficient AI workloads. It delivers improved performance, faster token generation, and greater efficiency compared with previous Flash-class models. By default, thinking capabilities (multi-step reasoning) are disabled to maximize speed and minimize cost. Developers can enable reasoning through the Reasoning API parameter when deeper intelligence is required, providing flexible control over the balance between latency, capability, and resource usage.

    输入
    $0.10/M
    输出
    $0.40/M
    上下文
    1M
    多模态理解编程数学与科学 1.3M tok
    google/gemini-2.5-flash-lite