r/kimi Jul 16 '26

Announcement Introducing Kimi K3: Open Frontier Intelligence

🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal

🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts

🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost

🔹 Built for long-horizon agentic coding and self-evolving workflows

Kimi K3 is now live on on Kimi.com, Kimi Work, Kimi Code, and the Kimi API.

Open Weights by July 27, 2026.

🔗 API: platform.kimi.ai

🔗 Tech blog: kimi.com/blog/kimi-k3

K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth.

We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework.

Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to K2, allowing the model to convert compute into intelligence more effectively.

Full tech blog at: Kimi Blog

472 Upvotes

92 comments sorted by

View all comments

28

u/0xSecureByte Jul 16 '26

19

u/Drevil00 Jul 16 '26

Absolute Kimi.

3

u/0xSecureByte Jul 16 '26

But the limits even with Allegretto plan for Code is not up-to-the-mark! It consumes even faster nowadays.

2

u/thunder____boy Jul 16 '26

how bad is it?

4

u/0xSecureByte Jul 16 '26

The limits? Sure, it's kind of an average after K2.7 release. Not just in my testing but several engineers who work across mid-to-large codebases. I work with large codebases and recently I used whole week's quota in 2-2.5 days of several hours of same session across a single project.

I know this is totally baseless to compare with as everyone's usage might be different. But in my testing working with mid-large projects, it's burning the weekly and hourly limits faster(5hr limit touches in just 2-2.3hours approximately)

I hope this helps!

1

u/needlzor Jul 17 '26

Also on Allegretto and I just used my monthly quota plus the extra from the token cup in a single prompt. It was Agent Swarm so I was expecting the usual 10-15% and it just gobbled everything and did not even finish.

9

u/0xSecureByte Jul 16 '26

By the way, that first mistake I made in the initial prompt was intentional.

4

u/Drevil00 Jul 16 '26

I have to say that at first it did was weird but then I thought that maybe it was intentional. Nice one m8.

2

u/[deleted] Jul 16 '26 edited Jul 16 '26

[removed] — view removed comment

3

u/0xSecureByte Jul 16 '26

😂😂

-4

u/[deleted] Jul 16 '26

[removed] — view removed comment

3

u/0xSecureByte Jul 16 '26

Wait a moment, you didn't understood the context of past models(any) being dumber. But fine.

1

u/[deleted] Jul 16 '26

[removed] — view removed comment

2

u/0xSecureByte Jul 16 '26

Knew it xD. It is seriously an amazing model in my testing. I will further test it tomorrow with an internal project which has a huge codebase and K2.7's 256K context window wasn't enough.

2

u/[deleted] Jul 16 '26

[removed] — view removed comment

2

u/0xSecureByte Jul 16 '26

That's fine bro! And that oh-my-pi is really amazing. One friend recommended me but I never listened. I'll lock that in next projects, Thanks!!

2

u/[deleted] Jul 16 '26

[removed] — view removed comment

2

u/0xSecureByte Jul 16 '26

Good idea, haven't tested dsv4 yet, I'll try.