LLM API for Every Leading Model
Call GPT, Claude, Gemini, DeepSeek, Qwen, Grok, Kimi, GLM and MiniMax models through one OpenAI-compatible API. Switch models by changing one parameter, and every model is 20% off the list price.
LLM API Model Families
Each family page lists every available version with its context window, token pricing and a playground to test prompts before you integrate.
GPT API
The complete OpenAI model lineup — GPT-6.1, GPT-5.6 and everything back to GPT-3.5, plus the o-series reasoning models, all accessible through a single API.
Text GenerationReasoningCodeView models & pricing →Claude API
Anthropic's Claude model family — from the lightweight Haiku to the Claude Opus 5.5 API for the most complex tasks — known for safety, long context windows, and nuanced reasoning.
ReasoningLong ContextSafe AIView models & pricing →Gemini API
Google's Gemini model family — from the ultrafast Flash, including the Gemini 3.8 Flash API, to the powerful Pro, featuring massive context windows and native multimodal understanding.
MultimodalLong ContextThinkingView models & pricing →DeepSeek API
DeepSeek V4 API access to DeepSeek's open-source model family — V4 Pro, V4 Flash and V4.1 Flash chat models plus the R1 reasoning model, delivering frontier-level performance at a fraction of the cost.
ReasoningOpen SourceCost-EffectiveView models & pricing →Qwen API
The Qwen family covers general and multimodal workloads. This page highlights qwen3.8-max, with up to 983,616 input tokens, 131,072 output tokens, required reasoning, caching, and built-in tools through the Responses API.
Long ContextRequired ReasoningBuilt-in ToolsView models & pricing →Grok API
The Grok API gives you xAI's Grok model family — Grok 4.7, 4.6, 4.5 and 4.3 plus reasoning, coding, and agent-oriented variants — through a unified API.
ReasoningCodingAgentsView models & pricing →Doubao API
ByteDance's Doubao model family — offering strong Chinese-English capabilities, vision understanding, and advanced thinking modes for complex tasks.
Chinese NLPVisionThinkingView models & pricing →GLM API
Zhipu AI's GLM model family — Chinese-English bilingual models with rapid iteration from GLM-4.6 to GLM-5.3, delivering strong reasoning and generation quality.
BilingualReasoningChatView models & pricing →Kimi API
Kimi K3 API access to Moonshot AI's Kimi model family — K3 and K2.7 Code with strong reasoning, coding capabilities, and extended thinking modes.
ReasoningCodingThinkingView models & pricing →MiniMax API
MiniMax's language model family — from M2.1 to M3, offering competitive performance for general chat, content creation, and bilingual tasks.
ChatContent CreationBilingualView models & pricing →Codex API
Codex API for OpenAI's coding models — GPT-5.3 Codex, GPT-5.2 Codex, GPT-5.1 Codex Max, GPT-5.1 Codex Mini and more, all through one OpenAI-compatible API.
Code GenerationReasoningCode ReviewView models & pricing →
How to Choose an LLM API
Complex reasoning and coding
Claude Opus 5.5, GPT-6.1 and Qwen 3.8 Max handle multi-step reasoning, agents and large codebases.
Fast and low-cost chat
Gemini 3.8 Flash, DeepSeek V4.1 Flash and GLM-5.3 Flash keep latency and token cost low for high-volume chat and extraction.
Long context
Several models support very long context windows for document analysis and retrieval; check each model's context limit on its family page.
Drop-in OpenAI compatibility
Use the OpenAI SDK you already have: point the base URL to APIMart and change the model name to switch providers.
LLM API FAQ
Is the APIMart LLM API OpenAI-compatible?
Yes. Chat models use the OpenAI-compatible Chat Completions format, so existing OpenAI SDK code works after changing the base URL, API key and model name. Claude models are also available through the Messages format.
How is LLM API pricing calculated?
LLM models are billed per million input and output tokens, and some models also price cached input. Each family page shows the token pricing per model, and every model is 20% off the list price, with membership tiers saving up to 28%.
Which LLM APIs are available?
APIMart offers 140+ chat models from OpenAI, Anthropic, Google, DeepSeek, Alibaba Qwen, xAI, Moonshot Kimi, Zhipu GLM and MiniMax, with new versions added as they launch.
Can I use one API key for multiple LLM providers?
Yes. A single APIMart key works for every chat, image and video model, with one balance and one invoice.
Do you support streaming and function calling?
Yes. Streaming responses, function or tool calling and vision input are supported on models whose providers offer those features.