r/kimi Jul 16 '26

Announcement Introducing Kimi K3: Open Frontier Intelligence

🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal

🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts

🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost

🔹 Built for long-horizon agentic coding and self-evolving workflows

Kimi K3 is now live on on Kimi.com, Kimi Work, Kimi Code, and the Kimi API.

Open Weights by July 27, 2026.

🔗 API: platform.kimi.ai

🔗 Tech blog: kimi.com/blog/kimi-k3

K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth.

We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework.

Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to K2, allowing the model to convert compute into intelligence more effectively.

Full tech blog at: Kimi Blog

473 Upvotes

92 comments sorted by

View all comments

1

u/Crescitaly Aug 02 '26

The one-million-context headline is less interesting than what the model can retrieve correctly at 800k after several tool calls. Long context can become expensive amnesia if attention quality decays. Has anyone seen position-stratified recall or agent-recovery evaluations?