MOONSHOT AI / 2026
Kimi K3
A native multimodal 2.8T-class model for long-horizon coding, knowledge work, and million-token contexts.
DECISION BRIEF
Long-horizon agents
Stable LatentMoE routes each token through 16 of 896 experts. Its native artifact is still 1.56 TB, making exact residency and expert communication decisive deployment constraints.
Publicly available checkpoints; verify terms before commercial deployment.
Unknown remains explicit until the exact checkpoint license is verified and sourced.
Architecture values are linked to the publishing organization.
ARCHITECTURE
Compute is sparse. Residency is not.
Stable LatentMoE routes each token through 16 of 896 experts. Its native artifact is still 1.56 TB, making exact residency and expert communication decisive deployment constraints.
Expert topology: 896 / top-16. The active-parameter count approximates token-level compute; it does not determine checkpoint memory, KV-cache demand, expert placement, or interconnect pressure.
Moonshot AI — Kimi K3
moonshotai/Kimi-K3@9f62e4e9fffb · 1.56 TB
NEXT ACTION
Test the workload, not the headline.
Start with checkpoint residency, then layer in context, concurrency, runtime overhead, topology, and resilience.
Build an assurance plan