Website Review
What is Kimi API Platform?
Kimi API Platform is the developer-facing service for Moonshot AI's Kimi large language models. It gives you API access to the Kimi K3 flagship model, the K2.7 Code coding model, and the K2.6 general-purpose model, plus a set of hosted tools such as web search, code execution, memory, and file analysis that you can call alongside the models.
What you actually get
- Model access. K3 is positioned for frontier work such as software engineering, knowledge work, and deep reasoning, with a 1M-token context window and tool calling. K2.7 Code targets coding tasks with a 256k-token context window. K2.6 handles vision and text input, with thinking and non-thinking modes, for chat and coding-agent tasks.
- Hosted tools. Rather than wiring up your own infrastructure, you can call built-in tools for web search, Python execution, JavaScript execution via QuickJS, Excel/CSV analysis, URL fetching and markdown conversion, memory persistence, date handling, unit conversion, and Base64 encoding.
- Agent-oriented workflows. The platform is pitched at autonomous programming agents and multi-step research, where the long context window matters for holding large codebases or document sets in view.
Who it suits
Developers building AI applications that need long-context reasoning, coding automation, or agentic pipelines with tool use. The hosted tool set is the main practical draw: if your app needs search, code execution, or persistent memory, you can integrate once instead of building each capability yourself. Teams that only need straightforward chat completion may find a lighter API sufficient.
Cost trade-off to weigh
The published rates show a clear tiering. K3 is the most expensive option at $3.00 per million input tokens and $15.00 per million output tokens, with cache hits at $0.30. K2.7 Code and K2.6 both sit at $0.95 input and $4.00 output, with cache hits at $0.19 and $0.16 respectively. Caching is the lever worth designing around: if your workload repeats large prompts or system context, the cache-hit rate is roughly a tenth of the standard input rate.
A practical next step
Pick your model by task rather than by defaulting to the flagship. Route coding and routine generation to K2.7 Code or K2.6, and reserve K3 for the long-context reasoning and complex multi-step work where its 1M-token window earns the price difference. Before committing, check the official documentation and current pricing at Kimi API Platform, and prototype one real workflow end to end — including a tool call and a cached prompt — so you can measure actual token spend against your budget.
How much does it cost to use Kimi K3 and other models via the API?
Kimi's API is priced per million tokens, with separate rates for input, output, and cached tokens. Kimi K3, the flagship model, is the most expensive tier; the coding and general-purpose models cost substantially less per token.
| Model | Input (per MTok) | Output (per MTok) | Cache hit (per MTok) |
|---|---|---|---|
| Kimi K3 | $3.00 | $15.00 | $0.30 |
| Kimi K2.7 Code | $0.95 | $4.00 | $0.19 |
| Kimi K2.6 | $0.95 | $4.00 | $0.16 |
K3 also shows a cache write rate of $3.00 per MTok. Cache hits are roughly a tenth of the input rate on K3, so reusing a stable prompt prefix is the main lever on cost for repeated workloads.
What the price difference means in practice
K3 pairs a 1M-token context window with frontier reasoning and long-horizon coding, so it suits software engineering, deep research, and multi-step agent workflows where quality matters more than volume. K2.7 Code targets coding tasks specifically with a 256k-token context and more reliable long-context instruction following. K2.6 is general purpose, handling vision and text input plus thinking and non-thinking modes, with the same 256k window.
A practical split: use K2.6 or K2.7 Code for high-volume, well-scoped requests, and reserve K3 for the hard cases—large-repo refactors, long document analysis, or agent runs that need to hold a lot of context at once.
Next step
Check the official pricing page for current rates and any model or region variations before you commit to a budget: Kimi API Platform. If you are estimating spend, start from your expected input/output ratio—output tokens dominate the bill on K3 at five times the input rate.
How do I get started with the Kimi API and integrate it into my application?
To start with the Kimi API, create a developer account on the platform, generate an API key in the developer console, and call the chat completions endpoint from your application using that key. The platform lists a Quick Start section and separate Documentation and Pricing pages, so the practical path is: sign up → get a key → follow the Quick Start example → pick a model → test tool calling and context limits.
H3 Which model to pick
The page lists three current models, and the choice depends on your task:
| Model | Best for | Context window | Trade-off |
|---|---|---|---|
| Kimi K3 | Software engineering, knowledge work, deep reasoning | 1M tokens | Highest listed cost per token |
| Kimi K2.7 Code | Coding tasks, long instructions | 256k tokens | Narrower than K3, cheaper |
| Kimi K2.6 | General chat, vision plus text, agent tasks | 256k tokens | Less specialised for frontier reasoning |
K3 is positioned as the flagship for long-horizon work; K2.7 Code is tuned for reliable instruction-following in code; K2.6 covers general and multimodal use. If you are unsure, prototype on K2.6 or K2.7 Code and move to K3 only where the extra context or reasoning pays for itself.
H3 A concrete first integration
Say you are building an internal support assistant that reads long ticket threads. Start by sending a single conversation to the chat endpoint, confirm the response format, then add the Web Search tool if answers need current information, and Memory if you want conversation history or user preferences to persist. The page also lists tools for code execution, Excel/CSV analysis, URL fetching, and unit conversion — these are described as plug-and-play, so you can enable them without building your own infrastructure.
H3 Practical checks before you commit
- Confirm current pricing on the official pricing page rather than relying on second-hand figures.
- Test your real prompt lengths against the 1M-token K3 window, since long contexts affect both cost and latency.
- Verify tool-calling behaviour with your own data, especially if you chain several tools.
- Check regional availability and any business terms if you are deploying commercially.
For background on the model family, see Kimi and Moonshot AI.
Which Kimi model should I choose for coding versus general dialogue tasks?
For coding, start with Kimi K2.7 Code; for general dialogue, vision, and mixed agent tasks, Kimi K2.6 is the more natural default. Choose Kimi K3 when the task is genuinely frontier-level: large refactors, deep multi-step reasoning, or work that needs the whole repository or a long document set in view at once.
How they differ in practice
| K3 | K2.7 Code | K2.6 | |
|---|---|---|---|
| Best fit | Frontier reasoning, software engineering, knowledge work | Coding tasks in long contexts | General dialogue, vision + text, agent tasks |
| Context window | 1M tokens | 256k tokens | 256k tokens |
| Modes | — | Instruction-following optimised for code | Thinking and non-thinking |
| Input / output price | $3.00 / $15.00 per MTok | $0.95 / $4.00 per MTok | $0.95 / $4.00 per MTok |
| Cache hit | $0.30 per MTok | $0.19 per MTok | $0.16 per MTok |
The practical trade-off is cost against scope. K3's output rate is roughly 3.75× K2.7 Code's, so it pays off when a task would otherwise need several retries or when you must feed in a very large codebase. For routine completion, test generation, or chat, the cheaper models are usually the better default.
A concrete scenario: a team building a coding agent would route everyday edits and test writing to K2.7 Code, send screenshots of UI bugs and ordinary chat to K2.6, and escalate only cross-file refactors or architectural analysis to K3. If your product mixes chat and code in one interface, K2.6's vision input and thinking/non-thinking switch may let you avoid running two models at all.
Next step: check the current per-model rates on the platform's pricing page before committing, since these figures can change: Kimi API Platform.
What can I build with Kimi's Tool Calling and the official built-in tools?
You can build agent-style applications that combine Kimi's Tool Calling with ready-made utilities, so the model can fetch live information, run code, read files and remember context instead of relying only on its training data.
What the official tools enable
The platform lists a set of plug-and-play tools you can wire into a model call. Practical uses:
- Web Search — answer questions about current events and cite sources, useful for research assistants or news summarizers.
- Code-Runner (Python) and Quick Js — execute generated code safely, so you can build a data-analysis or calculation assistant that returns verified results rather than guessed ones.
- Excel — parse spreadsheets and CSVs, good for finance, reporting or operations tools.
- Fetch — pull a URL and convert it to markdown, which suits documentation Q&A or content pipelines.
- Memory — persist conversation history and user preferences across sessions, the basis for a personalized assistant.
- Date, Convert, Base 64, Rethink, Random-Choice — smaller helpers that handle time handling, unit and currency conversion, encoding, idea organization and random selection.
Matching a model to the job
| Goal | Sensible starting point |
|---|---|
| Long-horizon coding agents, deep reasoning, large codebases | Kimi K3, with its 1M-token context window |
| Focused coding tasks | Kimi K2.7 Code, 256k context |
| Mixed vision and text, chat plus agent work | Kimi K2.6, 256k context |
The trade-off is cost versus capability: the flagship model is the most expensive per token, so reserve it for tasks where long context or harder reasoning actually matters, and route routine work to the smaller models. Cached input is charged at a much lower rate than fresh input, which favors designs that reuse a stable prompt or document prefix.
A concrete scenario
Say you want an internal assistant that reviews a quarterly spreadsheet and writes a summary with current market context. You would let the model call Excel to read the file, Web Search for recent figures, Code-Runner for any calculations, and Memory to remember the team's preferred report format. That combination is hard to assemble by hand and straightforward once the tools are registered.
A useful next step is to open the developer console, pick one tool such as Web Search, and build a single-turn example before adding memory or code execution. If you are weighing this against alternatives, the pricing page at Kimi API Platform and the developer console are the right places to confirm current rates, and you can compare tool ecosystems with OpenAI or Anthropic if you need a second reference point.
How does the 1M-token context window help with long documents or complex agent workflows?
A 1M-token context window changes what you can put in front of the model at once. Instead of retrieving a few fragments and hoping they contain the answer, you can supply an entire body of material — a long contract set, a full codebase snapshot, months of meeting notes — and let the model reason across all of it in a single pass. For agent workflows, the bigger benefit is continuity: intermediate results, tool outputs and earlier decisions can stay in context rather than being summarized away between steps.
Where it helps most
- Long-document analysis. Cross-referencing clauses, tracing a definition through a 300-page filing, or comparing several versions of a spec. The model can cite and reconcile details that would be split across many retrieval chunks.
- Repository-scale coding. Kimi's platform material describes K3 as built for software engineering and long-horizon coding, with the context window paired to autonomous agents handling debugging, refactoring and multi-step development. Keeping more of the repo in view reduces the "it edited the wrong file" failure mode.
- Multi-step agents. Tool Calling plus a large window means an agent can accumulate search results, file contents and its own prior reasoning without losing the thread. The platform lists plug-and-play tools including Web Search, Code-Runner, Excel, Memory and Fetch, which is the kind of loop where context accumulates fast.
- Deep research. Synthesizing many sources into one argument rather than summarizing each separately.
The trade-offs
A large window is not free or automatically better. Cost scales with tokens, and long inputs can dilute attention — models sometimes miss a detail buried in the middle. Latency rises, and stuffing everything in can be worse than retrieving the right 20 pages. Treat 1M as headroom for genuinely connected material, not as a replacement for good retrieval.
A practical next step
Pick one task you currently solve with chunked retrieval — say, answering questions across a 200-page policy manual. Run it twice: once with your existing retrieval pipeline, once with the full document in context. Compare answer accuracy, latency and token cost. That comparison, not the headline number, tells you whether the long window earns its place in your workflow.
For model tiers and current rates, see Kimi API Platform; competing long-context options include OpenAI and Anthropic.
User reviews (0)