What Is Moonshot AI and How Do You Use Its Kimi API?
Moonshot AI is the company behind the Kimi model family, and its Kimi API Platform (platform.kimi.ai) is where developers access those models programmatically. The current flagship is Kimi K3, a 2.8-trillion-parameter model with a 1M-token context window and Tool Calling, positioned for software engineering, knowledge work, and deep reasoning. If you want to build an AI application on Kimi models, you sign up through the platform, pick a model, and call it through the API — the sections below cover what each model is for, what tools ship with the platform, and how to get started.
What Moonshot AI and Kimi K3 are
Moonshot AI builds the Kimi large language models. The platform describes K3 as its "most capable flagship model to date," designed for frontier intelligence scenarios such as software engineering, knowledge work, and deep reasoning. Its headline specifications:
- 1M-token context window — enough to hold very large codebases, long documents, or extended multi-step agent transcripts in a single request.
- Tool Calling — the model can invoke external tools rather than only generating text.
- Code generation, intelligent dialogue, and visual reasoning — the platform lists these as core capability areas.
The platform also states it is "trusted by millions of professional developers," and its tooling is described as production-ready: integrate once and start building.
Choosing between the three models
The platform lists three models. They differ mainly in context window, capability focus, and price, so the choice usually comes down to task type and how much context you need.
| Model | Best for | Context window | Input | Output | Cache hit |
|---|---|---|---|---|---|
| Kimi K3 | Frontier tasks: software engineering, knowledge work, deep reasoning, long-horizon agents | 1M tokens | $3.00 / MTok | $15.00 / MTok | $0.30 / MTok |
| Kimi K2.7 Code | Coding tasks; more reliable instruction-following in long contexts | 256k tokens | $0.95 / MTok | $4.00 / MTok | $0.19 / MTok |
| Kimi K2.6 | General-purpose work: vision + text input, thinking and non-thinking modes, conversations and coding agent tasks | 256k tokens | $0.95 / MTok | $4.00 / MTok | $0.16 / MTok |
A practical way to decide:
- Need the largest context or the hardest reasoning/coding tasks? Use K3 — it is roughly 3x the input price and nearly 4x the output price of the other two, so reserve it for work that needs the extra capability.
- Doing focused coding work within 256k tokens? K2.7 Code is tuned for that, with higher success rates on coding tasks and more reliable long-context instruction-following.
- Need multimodal input (vision + text) or a mix of chat and agent tasks? K2.6 supports both vision and text input plus thinking/non-thinking modes.
K3 also has a cache write price of $3.00 / MTok, which the other two models do not list. Cache hits are dramatically cheaper than fresh input on all three — $0.30 vs. $3.00 on K3 — so repeated prompts or shared prefixes are where caching pays off.
The official plug-and-play tools
The platform ships a set of built-in tools you can attach to a model instead of building them yourself. The stated design goal is "integrate once and start building."
- Web Search — gives the model access to current information and cites authoritative sources, so results are verifiable.
- Code-Runner — executes Python code.
- Quick Js — safely executes JavaScript using the QuickJS engine.
- Excel — analysis for Excel and CSV files.
- Memory — a storage and retrieval system that persists conversation history, user preferences, and similar data.
- Fetch — extracts URL content and formats it as markdown.
- Rethink — an intelligent idea-organization tool.
- Random-Choice — a random selection tool.
- Date — date and time processing.
- Convert — unit conversion, including physics units and currency.
For agent-style work, the combination matters more than any single tool: Web Search plus Fetch handles retrieval, Code-Runner plus Quick Js handles execution, and Memory carries state across turns. The platform frames this as a "complete tool ecosystem that integrates general intelligence with vertical models — built for real world complexity," with agent programming and deep research/reasoning as the two named scenarios.
How to get started
The platform's own navigation points to a Quick Start path. The concrete entry points are:
- Get Started — the primary call-to-action on the platform landing page.
- Developer Console — where you manage API access (
platform.kimi.ai/console/payis the payment/console link). - Documentation — the reference for model behavior and tool integration.
- Pricing — the chat pricing page (
platform.kimi.ai/docs/pricing/chat) for current rates.
A typical first integration looks like this: create an account and get API credentials in the Developer Console, read the Quick Start and tool documentation, then send a request to the model you chose. If you plan to use tools like Web Search or Code-Runner, check the documentation for how they are declared in a request — the platform lists them as plug-and-play, but you still need to enable the ones you want.
Two things to verify before you commit:
- Current pricing. The rates above come from the platform's model listing. Confirm them on the pricing page, since model pricing changes.
- Access terms. The platform shows a "Kimi Business" membership option alongside developer pricing. Whether a given plan or region has usage limits, rate limits, or login requirements is not specified in the material here — check the console and pricing pages directly rather than assuming.
If you are deciding whether to build on Kimi at all, the strongest signals are the 1M-token context on K3, the built-in tool set, and the price gap between K3 and the cheaper K2.x models — that gap lets you route easy traffic to K2.6 or K2.7 Code and reserve K3 for the hard cases.