通过 Kimi K3 API 调用 Moonshot AI 的 Kimi 模型系列 — 以 K3 和 K2.7 Code 为核心,具备强大的推理、编码能力和扩展思考模式。
您在想什么?
询问 APIMart
kimi-k3
透明定价,无隐藏费用。按量付费,用多少付多少。
每 100 万 token
| 计费项 | 黄金价当前价格 | 铂金价 | 钻石价 | 官方价 | 为你节省 |
|---|---|---|---|---|---|
| 输入 | 24~$2.4 | 22.8~$2.28 | 21.6~$2.16 | 30~$3 | 20% |
| 缓存输入 | 2.4~$0.24 | 2.28~$0.228 | 2.16~$0.216 | 3~$0.3 | 20% |
| 输出 | 120~$12 | 114~$11.4 | 108~$10.8 | 150~$15 | 20% |
使用 Kimi K3、K2.7 Code、K2-Thinking 和 K2.6。Moonshot AI 的前沿模型,在推理、编码和扩展思考方面表现出色。
50K+
活跃用户
99.9%
在线率
2x
更快
70%
成本节省
Kimi 模型为何受到 AI 开发者的青睐
团队如何在生产环境中使用 Kimi 模型
几分钟内即可开始使用 Moonshot 的模型
创建免费的 APIMart 账户并充值余额。
生成 API Key,以编程方式访问 Kimi 模型。
使用上方的 Playground 或通过 OpenAI 兼容 API 进行集成。
开发者对我们 Kimi API 的评价
“Kimi K2 的编码能力真的令人印象深刻,轻松处理复杂的重构任务。”
Alex Chen
高级开发者
“K2-Thinking 解决了 GPT-4o 无法解决的数学问题,推理深度令人瞩目。”
Sarah Li
AI 研究员
“K2.5 是一个重大升级。我们已将其集成到中英文内容生产流水线中。”
Mike Wang
内容负责人
“价格有竞争力,性能扎实。Kimi 是大厂模型的优秀替代方案。”
Emily Zhou
产品经理
“思考模式非常适合我们的法律分析工具,能够系统地处理复杂案例。”
David Liu
法律科技工程师
“API 速度快、文档清晰、运行稳定。非常满意通过 APIMart 使用 Kimi。”
Lisa Zhang
后端开发者
关于 Kimi API 的常见问题
我们以最新的 Kimi K3 为主,另有 K2.7 Code、K2.7 Code Highspeed、K2.6 和 K2.5。K2 Instruct、K2 Thinking 等旧型号仍可继续使用。
K3 是 Moonshot AI 最新的 Kimi 模型,上下文窗口最高可达 100 万 token。K2.7 Code 面向编程场景,K2-Thinking 增加了扩展推理,适用于复杂问题。
Kimi K3 是 Moonshot AI 最新模型,上下文窗口最高可达 100 万 token,适合需要最新能力和长上下文的场景。K2.6、K2.5 仍可用于现有业务,K2.7 Code 面向编程,K2-Thinking 提供扩展推理。所有模型共用同一个 APIMart API Key,切换时只需更换模型 ID。
按 token 计费(输入 + 输出)。请查看上方的价格表了解具体费率。
是的。Kimi 模型使用标准的 chat completions 格式,可与 OpenAI SDK 配合使用。
是的。您的 API 数据不会被存储或用于训练,所有请求均已加密。
最新能力和长上下文选 K3,编程选 K2.7 Code,需要更快的编程响应选 K2.7 Code Highspeed,复杂推理选 K2-Thinking。
Moonshot AI Kimi K3 API 按每百万输入、输出 token 计费,本页价格表列出了全部规格。页面展示的是黄金会员价(8 折,立省 20%),铂金、钻石会员更省,最高省 28%。只为成功请求付费。
注册 APIMart 账号,在控制台的 API Keys 页面创建密钥即可。同一个密钥可调用 Moonshot AI Kimi API 及平台上的全部模型,按量付费,无需订阅。
不是。Kimi K3 API 没有免费额度,采用按量付费,只为成功请求付费,所有价格均享 8 折(立省 20%)。
Moonshot AI Kimi API 兼容 OpenAI:使用 OpenAI SDK 或任意 HTTP 客户端调用 APIMart 的 Chat Completions 接口,用 APIMart API key 鉴权,并将 model 设为本页中您需要的 Kimi 模型。完整参数和代码示例见 APIMart 文档(docs.apimart.ai)。
你可以通过页面右下角的在线客服、发送邮件至 [email protected],或在 Discord 社区中联系我们。我们的团队会尽快回复你的问题。
探索同类型的其他模型。
Claude Opus 5
Claude Opus 5 is Anthropic’s next-generation flagship large language model, featuring enhanced capabilities in code development, complex reasoning, knowledge analysis, and agent execution, making it ideal for large-scale project development and professional work scenarios.
Gemini 3.5 Flash Lite
A lightweight multimodal model launched by Google that emphasizes low cost, low latency, and high throughput. It is suitable for document parsing, data extraction, structured output, and large-scale agent workflows, delivering faster response times while maintaining core reasoning capabilities.
Gemini 3.6 Flash
Gemini 3.6 Flash is a faster, more token-efficient, and more code-capable version of the Gemini Flash model, suitable for everyday development, AI agents, and high-concurrency applications.
Qwen 3.7 Max
Launched by Alibaba Cloud Tongyi Qianwen, this flagship large language model features enhanced capabilities in reasoning, coding, and handling complex tasks, making it suitable for enterprise-level applications, chatbots, and high-difficulty task scenarios.