Google Web Search
~$14
Charged only when web search is enabled; the precharge is reconciled against actual usage.
Precharged queries: 1
Google's Gemini model family — from the ultrafast Flash, including the Gemini 3.8 Flash API, to the powerful Pro, featuring massive context windows and native multimodal understanding.
What's on your mind?
Ask APIMart
gemini-3.8-flash
Transparent pricing with no hidden fees. Pay only for what you use.
Per 1M tokens
| Rate | GoldCurrent price | Platinum | Diamond | Official | Your savings |
|---|---|---|---|---|---|
| Input | 6~$0.6 | 5.7~$0.57 | 5.4~$0.54 | 7.5~$0.75 | 20% |
| Cached input | 0.6~$0.06 | 0.57~$0.057 | 0.54~$0.054 | 0.75~$0.075 | 20% |
| Output | 30~$3 | 28.5~$2.85 | 27~$2.7 | 37.5~$3.75 | 20% |
~$14
Charged only when web search is enabled; the precharge is reconciled against actual usage.
Precharged queries: 1
~$0.0833
~$1
Charged once based on token count and the selected TTL when an explicit cache is created.
Access the full Gemini lineup, from the latest Gemini 3.8 Flash and Flash Lite to Gemini 3.1 Pro. Industry-leading 1M+ context windows, native multimodal capabilities, and advanced thinking modes.
50K+
Active Users
99.9%
Uptime
2x
Faster
70%
Cost Savings
What makes Gemini a leading choice for multimodal AI
How teams leverage Gemini models in production
Start using Google's models in minutes
Create your free APIMart account and top up your balance.
Generate an API key to access Gemini models programmatically.
Use the playground above or integrate via the OpenAI-compatible API.
What developers say about our Gemini API
“The 1M context window is a game changer. We process entire codebases in a single call for our code review tool.”
Alex Chen
Engineering Lead
“Gemini Flash is incredibly fast and cheap. Perfect for our real-time translation service.”
Sarah Li
Product Manager
“The multimodal capabilities let us build a single pipeline for text, images, and documents.”
Mike Wang
ML Engineer
“Gemini 2.5 Pro Thinking is our go-to for complex analysis tasks. The reasoning quality is excellent.”
Emily Zhou
Data Scientist
“Gemini Thinking is excellent for our research product. It works through complex information step by step.”
David Liu
Senior Developer
“We use Flash Lite for high-volume summarization. The cost savings over GPT-4 are significant.”
Lisa Zhang
Platform Engineer
Common questions about using the Gemini API
We offer the latest Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash, 3.5 Flash, and 3.5 Flash Lite, plus Gemini 3.1 Pro Preview and Gemini 3 Pro Preview, with thinking variants. Earlier Gemini 2.5 Pro/Flash and 2.0 Flash models remain available.
Thinking variants use extended chain-of-thought reasoning. You can control the thinking budget (128 to 3000 tokens).
Token-based pricing (input + output). Flash Lite is the most affordable, Flash is mid-range, and Pro is premium.
Yes. Gemini models use the standard chat completions format and work with the OpenAI SDK.
Yes. Your API data is not stored or used for training. All requests are encrypted.
Gemini 3.5 Flash Lite for high-volume, low-cost tasks. Gemini 3.8 Flash for balanced performance. Gemini 3.1 Pro for maximum quality. Add thinking mode for complex reasoning.
Yes. The Google Gemini API is billed per million tokens, with input and output tokens priced separately. The pricing table on this page lists every model and specification, and the prices shown are Gold member prices (20% off). Platinum and Diamond members save even more, up to 28%, and you only pay for successful requests.
Gemini 3 API calls, such as Gemini 3 Pro Preview or Gemini 3.8 Flash, are billed per million input and output tokens, and each model has its own rate: Flash Lite is the most affordable tier and Pro is the premium tier. See this page's live pricing table for current Gemini API pricing.
Sign up for an APIMart account, open the API Keys page and create a key. The same key works for the Gemini API and every other model on APIMart, with pay-as-you-go billing and no subscription.
Gemini models on APIMart use the OpenAI-compatible Chat Completions format. Use your APIMart key, pick a Gemini model such as Gemini 3 Pro or Gemini 3.8 Flash, and send your messages with the OpenAI SDK or any HTTP client. Model IDs, parameters, and code examples are in the APIMart docs (docs.apimart.ai).
No. APIMart does not provide a free quota for the Gemini API. Gemini models are pay-as-you-go by token usage, you only pay for successful requests, and members save 20%–28% on every price.
You can reach us via the live chat in the bottom-right corner, email us at [email protected], or join our Discord community. Our team will get back to you as soon as possible.
Explore more models in the same category.
Claude Haiku 5.5
Claude Haiku 5.5 API on APIMart delivers fast, cost-efficient AI, offering responsive conversations, capable coding assistance, and reliable performance for high-volume workflows.
Kimi K3
Kimi-K3 is a next-generation large language model launched by Moonshot AI. It features ultra-long context, multimodal understanding, and strong coding capabilities, making it suitable for complex reasoning, software development, knowledge analysis, and agent automation tasks.
Claude Opus 5
Claude Opus 5 is Anthropic’s next-generation flagship large language model, featuring enhanced capabilities in code development, complex reasoning, knowledge analysis, and agent execution, making it ideal for large-scale project development and professional work scenarios.
Gemini 3.5 Flash Lite
A lightweight multimodal model launched by Google that emphasizes low cost, low latency, and high throughput. It is suitable for document parsing, data extraction, structured output, and large-scale agent workflows, delivering faster response times while maintaining core reasoning capabilities.