Reasoning / 1M context / multimodal / coding

Kimi (Moonshot AI)

Kimi K3 - 1M ctx, multimodal, open MoE (self-host available)

Kimi K3 (released 07/16/2026) is Moonshot AI's flagship model: approx. 2.8T-parameter MoE (16 of 896 experts active), a 1M-token context window and native multimodality (text+image; video in API).

Moonshot calls it the first open ~3T-class model; full weights were published around 07/26-07/27/2026 on Hugging Face under the Kimi K3 License (not OSI open source - separate MaaS terms above revenue thresholds). Available via Kimi.com, Kimi Work, Kimi Code and the API (`kimi-k3`).

Older K2.7 Code and K2.6 variants remain in the API for narrow legacy/specialized cases.

  • Verified2026-08-01
  • Models covered3
  • Sources4
  • Deployment channels3

Purchase decision (when to choose / when to avoid)

Choose if...

  • You analyze very long documents or repositories and need a 1M context window at a good price.
  • You run long coding/reasoning/agent tasks and want flagship Kimi K3 instead of more expensive premium models.
  • You want multimodality (image/video) and a 1M context window in the API or Kimi Code.

Avoid if...

  • Your priority is Polish language and local nuances (often weaker than top-3 + Bielik).
  • You need enterprise governance in the EU - check terms, regions and retention.
  • You need self-host without supernode-class hardware - K3 requires a 64+ accelerator cluster; consider API or a lighter open-weight model.
  • You need an OSI license (MIT/Apache) - Kimi K3 License has separate commercial terms (incl. MaaS).

Cost in practice (scenarios)

Long-doc analysis

Often cheaper than premium top-3 for long contexts (check pricing).

  • large documents
  • frequent summaries
Reasoning/coding

Good cost/quality in API, but K3 always emits reasoning_content - reasoning tokens count toward the limit alongside output.

  • reasoning tasks
  • multi-turn sessions
Self-host K3

Public weights from ~07/27/2026 (Kimi K3 License); needs supernode infrastructure (64+ accelerators) - budget for cluster and MLOps, not a single GPU.

  • on-prem
  • data control
  • long context
  • MaaS terms above revenue thresholds
These are estimates/scenarios (not an invoice). Actual cost depends on context length, number of users, limits and retention policies.

Deployment / data / enterprise

Deployment channels

  • Self-host K3 (public weights from ~07/27/2026; supernode 64+ accelerators; Kimi K3 License)
  • Kimi OpenPlatform (API)
  • Kimi.com / Kimi Work / Kimi Code (SaaS)

Data policy

Training on data
Customer content may be used to maintain and improve services - exclusion requires an enterprise agreement (check OpenPlatform terms).
Retention
Varies by data type and purpose; account, input and payment data may be retained for the life of the account.
Data residency
Servers in Singapore (Kimi OpenPlatform).
Cross-border transfer and use of data for service improvement require legal/security sign-off before processing store customer data.

Enterprise readiness

Admin
API + billing; enterprise depends on offering.
SSO/SCIM
Depends on enterprise offering.
Audit
Depends on enterprise offering.
DPA
Depends on agreement.
Certifications
Depends on agreement.
1M context at a good price; integrations and compliance need checking.

Best use cases

  • very long document and repository analysis (reports, contracts, specifications) - 1M context
  • long coding, reasoning and agentic tasks in pipelines - especially Kimi K3
  • multimodal pipelines: image and video analysis in a single query through K3.

Strengths

  • 1M context and strong reasoning at a competitive price versus premium models.
  • K3 as one flagship model instead of split K2.x lines; public weights for self-host.
  • API available globally; good quality in Chinese and English; Kimi Code for terminal agents.

Weaknesses / risks

  • Smaller integration ecosystem than OpenAI/Claude; documentation partly in Chinese.
  • Limited Polish language quality; EU compliance requires verification.
  • Kimi K3 License is not OSI open source; self-host needs supernode-class hardware (2.8T MoE, 64+ accelerators).

Current models (examples)

  • Kimi K3 (`kimi-k3`) - flagship model (1M ctx); open weights from ~07/27/2026 (self-host: supernode).
  • Kimi K2.7 Code - coding variant (256K ctx); legacy/specialized.
  • Kimi K2.6 / K2.5 / K2 - previous generations, currently legacy.

Alternatives (if this model doesn't fit)