What Is the K3 Large Language Model and How Do You Use It?

Kimi K3 is Moonshot AI's flagship large language model, available through the Kimi API Platform at platform.kimi.ai. It pairs a 1M-token context window with Tool Calling, and the platform positions it for software engineering, knowledge work, and deep reasoning. If you need a single model that can hold very long inputs and drive multi-step agent workflows, K3 is the default choice on this platform. If your task is narrower — pure coding or mixed vision-and-text at lower cost — the sibling models below are usually the better fit.

What K3 actually is

K3 is described as Kimi's most capable flagship model to date. The two specs that matter most for planning:

  • 1M-token context window — large enough to hold whole codebases, long document sets, or extended agent transcripts in one request.
  • Tool Calling — the model can invoke tools rather than only producing text, which is what makes agent-style workflows possible.

The platform frames K3 around "frontier intelligence scenarios": software engineering, knowledge work, and deep reasoning. Its long-horizon coding capability is called out specifically as the basis for autonomous programming agents that handle debugging, refactoring, and multi-step development.

Choosing between K3 and its siblings

The platform lists three current models. They differ mainly in context size, modality, and price, so the decision is usually about task fit rather than raw capability.

Model Context Best for Input / MTok Output / MTok Cache hit / MTok
K3 1M tokens Software engineering, knowledge work, deep reasoning, long-horizon agents $3.00 $15.00 $0.30
K2.7 Code 256k tokens Coding tasks, reliable instruction-following in long contexts $0.95 $4.00 $0.19
K2.6 256k tokens General tasks, vision + text input, thinking and non-thinking modes, conversations and coding agents $0.95 $4.00 $0.16

A practical rule: start with K2.6 or K2.7 Code when the task fits in 256k tokens and doesn't need frontier reasoning. Move to K3 when you need the 1M window, when the task is genuinely hard reasoning, or when an agent has to sustain a long multi-step workflow. K3's cache write is also $3.00 / MTok, which matters if you repeatedly send the same long prefix.

Extending K3 with official tools

K3's Tool Calling connects to a set of production-ready tools the platform describes as "integrate once and start building." These are the ones most likely to change what you can build:

  • Web Search — lets the model pull current information and cite sources, so results are verifiable.
  • Memory — persistent storage and retrieval for conversation history and user preferences.
  • Code-Runner — executes Python.
  • Quick Js — safely executes JavaScript via the QuickJS engine.
  • Excel — analysis for Excel and CSV files.
  • Fetch — extracts URL content and formats it as markdown.
  • Rethink, Random-Choice, Date, Convert, Base 64 — idea organization, random selection, date/time handling, unit and currency conversion, and encoding/decoding.

For an agent that reads a repository, runs tests, and searches for a fix, the combination of K3's long context, Code-Runner, and Web Search covers most of the loop without custom infrastructure.

Getting started

The platform's quick-start path runs through the developer console:

  1. Go to platform.kimi.ai and open the developer console.
  2. Obtain API access (the console is the entry point for getting started).
  3. Make a first request against the K3 endpoint.
  4. Add tools as needed by following the documentation.

The documentation and pricing pages are linked from the platform's main navigation, and the console is where you manage access and payment.

What to watch for

  • Cost scales with context. K3's 1M window is a capability, not a default — sending a million tokens on every call is expensive at $3.00 / MTok input. Use cache hits ($0.30 / MTok) for repeated long prefixes.
  • Pick the model per task, not per project. Routing simple requests to K2.6 and hard ones to K3 is the cheapest way to use the platform.
  • Tool Calling is the differentiator. If you only need text generation, the smaller models are sufficient; K3's value shows up in agent and long-context work.
platform.kimi.ai
Kimi API Platform, providing the 2.8-trillion-parameter Kimi K3 large language model API with a 1M-token context window and Tool Calling. Professiona…