Moonshot AI · Launched July 16, 2026

Kimi K3.
Open Frontier Intelligence.

The world's first open 3T-class model — 2.8 trillion parameters, a 1-million-token context window, and native vision. #1 on Frontend Code Arena, ahead of Claude Opus 4.8 and GPT-5.5, at roughly half their cost.

Via Moonshot's referral program — you get up to 1-year membership credits, and so do we.

Official Kimi K3 launch visual by Moonshot AI
OFFICIAL KIMI K3 LAUNCH VISUAL — MOONSHOT AI
PARAMETERS
2.8T MoE
ACTIVE / EXPERTS
16 of 896
CONTEXT
1M tokens
OPEN WEIGHTS
July 27

What Is Kimi K3?

Kimi K3 is the newest flagship AI model from Moonshot AI, officially launched on July 16, 2026. It is the largest open-weight model ever released — a 2.8-trillion-parameter mixture-of-experts system that routes each token through just 16 of its 896 specialized experts, delivering frontier intelligence at practical cost. (Many people also search for it as “Kimi 3.0” — the official name is Kimi K3. Full specifications below.)

K3 introduces a 1-million-token context window, native image and video understanding, and two architectural breakthroughs — Kimi Delta Attention and Attention Residuals — that make it up to 6.3× faster at long contexts and ~2.5× more scaling-efficient than Kimi K2. On independent testing it ranks #4 out of 189 AI models, ahead of Claude Opus 4.8 and GPT-5.5, while costing roughly half as much per task.

Best of all: Moonshot's brand-new referral program means you can start using Kimi K3 with free bonus membership credits — up to a full year — just by signing up through an invite link like ours.

Watch · Official demo

Kimi K3 designed a chip — for itself.

In a single 48-hour autonomous run, K3 built, optimized and verified a working chip using open-source EDA tools. This video was also written, shot and rendered by K3.

48 hrs

fully autonomous run

3.981 mm²

Nangate 45nm · 100 MHz

8,721 tok/s

decode throughput in sim

1.46M

standard cells placed

SOURCE: MOONSHOT AI — KIMI K3 LAUNCH BLOG

Benchmarks · Official + independent

The numbers, straight from Moonshot's launch suite.

Kimi K3 ranks #1 on 4 of 8 real-world agentic benchmarks (Program Bench, SWE Marathon, SpreadsheetBench 2, Automation Bench — plus BrowseComp), and is the only open model within striking distance of GPT-5.6 Sol and Claude Fable 5. Moonshot itself notes K3 still trails those two overall — and independent testers at Artificial Analysis place it #4 of 189 models with a score of 57.

Official Kimi K3 coding benchmarks: DeepSWE, Terminal-Bench 2.1, FrontierSWE, Program Bench, Kimi Code Bench 2.0, SWE Marathon
OFFICIAL CODING BENCHMARKS — MOONSHOT AI, JULY 2026 · ALL MODELS AT MAX THINKING EFFORT
Official Kimi K3 general and visual agent benchmarks: GDPval-AA, JobBench, AA-Briefcase, SpreadsheetBench 2, Automation Bench, BrowseComp, CharXiv, Zerobench
OFFICIAL GENERAL & VISUAL AGENT BENCHMARKS — MOONSHOT AI, JULY 2026

Where Kimi K3 finishes first (or within 0.5 of first)

RECREATED 1:1 FROM THE OFFICIAL CHARTS ABOVE

Terminal-Bench 2.1

GPT-5.6 Sol
88.8
Kimi K3
88.3
Claude Opus 4.8
84.6
Claude Fable 5
84.6
GPT-5.5
83.4

Program Bench — K3 #1

Kimi K3
77.8
GPT-5.6 Sol
77.6
Claude Fable 5
76.8
Claude Opus 4.8
71.9
GPT-5.5
70.8

SWE Marathon — K3 #1

Kimi K3
42
Claude Opus 4.8
40
GPT-5.6 Sol
39
Claude Fable 5
35
GPT-5.5
14

BrowseComp — K3 #1

Kimi K3
91.2
GPT-5.6 Sol
90.4
Claude Fable 5
88
GPT-5.5
84.4
Claude Opus 4.8
84.3

SpreadsheetBench 2 — K3 #1

Kimi K3
34.8
Claude Fable 5
34.7
GPT-5.6 Sol
32.4
Claude Opus 4.8
31.6
GPT-5.5
29.1

Automation Bench — K3 #1

Kimi K3
30.8
GPT-5.6 Sol
29.7
Claude Fable 5
29.1
Claude Opus 4.8
27.2
GPT-5.5
22.7

Note (from Moonshot's own footnotes): Fable 5 results include potential fallback behavior; GPT-5.6 Sol results include potential cyberguards. BrowseComp uses context compaction at 300K; with a full 1M-token window K3 scores 90.4–91.2. Harnesses: KimiCode, Claude Code, or Codex per benchmark.

LMArena · Community-voted

#1 on Frontend Code Arena — above every closed model.

Real users, blind votes. Kimi K3 jumped 17 places from K2.6 (#18 → #1) with a 76% pairwise win rate — vs 63% for Claude Fable 5 and 58% for GPT-5.6 Sol.

FRONTEND CODE ARENA — ELO

Kimi K3
1679
Claude Fable 5
1631
GPT-5.6 Sol
1590

#1

Frontend Arena

76%

pairwise win rate

#9

Text Arena (was #38)

#1 IN 6 OF 7 FRONTEND DOMAINS

  • Brand & Marketing #1
  • Reference-Based Design #1
  • Data & Analytics #1
  • Consumer Product #1
  • Simulations #1
  • Content Creation Tools #1
  • Gaming #2 · behind Fable 5

SOURCE: @ARENA ON X · JULY 16–17, 2026

Features

Why Kimi K3 Is a Big Deal

Six breakthroughs that make K3 the most exciting AI launch of 2026.

🧠

2.8T MoE Architecture

896 expert subnetworks with only 16 activated per token — massive capacity, efficient compute. The largest open-weight model ever released.

📚

1M Token Context Window

Feed entire codebases, full books, or dozens of reports in a single prompt. K3 remembers everything from start to finish.

👁️

Native Vision & Video

Natively understands images, screenshots, diagrams and video — optimize frontend design, CAD, and UI from a single screenshot.

Kimi Delta Attention

New KDA architecture delivers up to 6.3× faster decoding at million-token contexts and ~2.5× better scaling efficiency than K2.

🤖

Long-Horizon Agentic Coding

Sustains multi-hour engineering sessions with minimal supervision — in testing it ran a 48-hour chip-design pipeline autonomously.

🔓

Open Weights (July 27)

Full model weights released under a Modified MIT license — download, fine-tune, and run K3 locally with full data privacy.

Under the hood

The architecture behind K3.

Nine of the past twelve months, Kimi models have set the upper bound of open-model sizes. K3 pushes it again — not with brute force, but with new attention and routing math.

Kimi K3 architecture diagram: the Stable LatentMoE and Kimi Delta Attention modules (left), the AttnRes operation (top right), and the Block Attention Residuals backbone (right)
OFFICIAL KIMI K3 ARCHITECTURE — STABLE LATENTMOE + KDA MODULES, ATTNRES, AND BLOCK ATTENTION RESIDUALS BACKBONE · SOURCE: MOONSHOT AI

Kimi Delta Attention (KDA)

A linear-attention variant that forms the efficient backbone for scaling — up to 6.3× faster decoding at million-token contexts, with a prefill-cache implementation contributed to vLLM.

Attention Residuals (AttnRes)

Selectively retrieves representations across model depth instead of accumulating them uniformly — roughly +25% training efficiency at under 2% extra compute.

Stable LatentMoE — 16 of 896 experts

Extreme sparsity: each token activates just 16 of 896 experts (~50B of 2.8T parameters), keeping inference practical at 3T-class scale.

Quantile Balancing

Derives expert allocation directly from router-score quantiles — loss-free load balancing with no auxiliary loss, no learning rate, no sensitive hyperparameter.

Per-Head Muon

Extends the Muon optimizer by optimizing attention heads independently for more adaptive learning at scale.

SiTU + Gated MLA

Sigmoid Tanh Unit improves activation control; Gated MLA improves attention selectivity — together enabling stable 2.8T-scale training.

MXFP4 quantization-aware training

QAT from the SFT stage onward — MXFP4 weights with MXFP8 activations for broad hardware compatibility and cheap serving.

~2.5× scaling efficiency vs K2

The combined structural changes convert compute into intelligence ~2.5× more effectively than the K2 generation.

FULL DETAILS LAND WITH THE K3 TECHNICAL REPORT · WEIGHTS BY JULY 27, 2026 (MODIFIED MIT)

Vision-in-the-loop · Game dev

One prompt. Fully playable games.

K3 combines 3D reasoning, coding and native vision — it writes code, takes a screenshot of the result, sees what looks wrong, and iterates. Every frame below was generated by Kimi K3 from a single prompt, taken from Moonshot's launch blog.

Open-world exploration game — generated by Kimi K3
Open-world exploration game
"Swordrealm of the Nine Heavens" — wuxia RPG — generated by Kimi K3
"Swordrealm of the Nine Heavens" — wuxia RPG
Cyberpunk web-swinging game — generated by Kimi K3
Cyberpunk web-swinging game
Voxel colosseum — generated by Kimi K3
Voxel colosseum
First-person balloon-shooter arena — generated by Kimi K3
First-person balloon-shooter arena
Gargantua black-hole simulation — generated by Kimi K3
Gargantua black-hole simulation

SCREENSHOTS: MOONSHOT AI — KIMI K3 GAME CASES

Kimi Work · Widgets & Dashboard

A "living" dashboard that builds itself.

In Kimi Work, K3 generates interactive widgets — stock watchlists, Pomodoro timers, habit trackers, music players — that connect to local data and keep updating. Dashboard pins them into one persistent view built around your project.

Case studies · Long-horizon autonomy

What a 2.8T agent actually does all day.

01

MiniTriton — a GPU compiler from scratch

K3 built a compact Triton-like compiler with its own tile-level IR over MLIR, optimization passes and a PTX codegen pipeline. It matches or beats Triton on supported roofline benchmarks and sustains end-to-end nanoGPT training.

DSL → IR → PTX → runtime

02

Kernel optimization vs the frontier

Given 24 hours in an identical sandbox to profile and rewrite GPU kernels (AttnRes, KDA, 512-head-dim MLA), K3 performed competitively with Claude Fable 5 and substantially outperformed Opus 4.8, GPT-5.6 Sol and GPT-5.5. An early K3 build already handled most of Moonshot's own kernel work.

NVIDIA H200 + alt-vendor GPGPU

03

2 hours of K3 ≈ 1–2 weeks of a researcher

To reproduce the I–Love–Q universal relations in computational astrophysics, K3 cross-validated 20+ papers, implemented the full numerical pipeline, evaluated 300+ equations of state, caught inconsistencies in published formulas, and wrote 3,000+ lines of Python plus an interactive dashboard.

Computational astrophysics

04

42 years of the AI ASIC industry, one report

Through 120+ rounds of recursive self-improvement, K3 produced an interactive research site: 2.8k+ web searches, 1.1k+ terminal data pulls, 11k+ pages across 87 quarterly reports and 99 source PDFs — rendered as bespoke charts and animated diagrams.

Kimi Work · deep research

SOURCE: MOONSHOT AI — KIMI K3 LAUNCH BLOG

Kimi K3 vs Kimi K2

Not an update. A new generation.

Spec Kimi K2 Kimi K3
Total parameters 1.0T 2.8T
Architecture MLA MoE KDA + AttnRes + Stable LatentMoE
Experts 384 896 (16 active / token)
Context window 256K tokens 1,048,576 tokens
Vision / video input Native
Long-context decode baseline up to 6.3× faster
Scaling efficiency baseline ~2.5× better
Frontend Code Arena #18 (K2.6) #1
Weights license Modified MIT Modified MIT · July 27

Read before use

K3's limitations — documented by Moonshot itself.

Rare honesty for an AI launch, and exactly why this model is worth trusting. These are from the official launch post and independent testing.

01

Don't switch models mid-session

K3 is trained with preserved thinking history. If the harness drops it — or you switch from another model mid-session — output quality can become unstable. Start a fresh K3 session instead.

02

It can be overly proactive

Tuned for long-horizon tasks, K3 may make decisions on your behalf when intent is ambiguous. Set explicit boundaries in the system prompt or AGENTS.md.

03

Not above every closed model

Moonshot's own words: K3 "still trails the most powerful proprietary models, Claude Fable 5 and GPT-5.6 Sol" in overall user experience. It's the best open model — not the best model, period.

04

Speed is the trade-off

Independent tests measure ~62 output tokens/sec with ~2s TTFT — below the median for its price tier. Brilliant for long runs and batch work; less ideal for snappy chat UX.

Referral program · Live now

Get free Kimi K3 membership credits.

Moonshot launched its official referral program alongside K3. Sign up through our invite link and both of us receive a guaranteed benefit — up to 1-year membership credits.

01

Open the invite link

It leads to Kimi's official page with our invitation code pre-applied.

02

Create your free account

Under a minute. No credit card required to start using Kimi K3.

03

Both sides get rewarded

Moonshot's referral program gives you — and us — up to 1-year membership credits.

Claim Your Free Credits →

Invitation code KRPVU2 · applied automatically

Availability

Four ways to use Kimi K3 today.

Kimi app & kimi.com

Chat with K3 free on iOS, Android, HarmonyOS, or the web. The fastest way to try it.

Sign up with bonus credits →

Kimi Work

Download Kimi K3 for PC — the desktop app (Windows / Apple silicon, v3.1.0+) for deep research, widgets and dashboards.

Get Kimi Work →

Kimi Code

Terminal + IDE agent. Run it and switch to K3 with the /model command.

Docs →

Kimi API

OpenAI-compatible endpoint. Create a Kimi API key, select kimi-k3 — reasoning effort is max by default at launch. Full setup guides in the official Kimi docs.

API platform →
kimi_k3_quickstart.py
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_MOONSHOT_API_KEY",
    base_url="https://api.moonshot.ai/v1",
)

resp = client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Refactor this repo"}],
    reasoning_effort="max",  # default at launch
)

Pricing

Frontier-class intelligence at Sonnet-tier prices.

Kimi K3 API Per 1M tokens
Cache-hit input $0.30
Cache-miss input $3.00
Output $15.00
Avg. cost per benchmark task $0.94

Mooncake disaggregated inference · >90% cache-hit rate in coding workloads

  • ~50% cheaper than Opus 4.8 / GPT-5.5 per completed task, and ~3× cheaper than Claude Fable 5 ($2.75/task).
  • Free to try in the Kimi app — no credit card. Heavy launch traffic means free tiers can queue; membership skips the line.
  • Referral bonus: sign up through our link and you get up to 1-year membership credits — free.
Try Kimi K3 Free →

FAQ

Kimi K3, answered.

What is Kimi K3? +

Kimi K3 is Moonshot AI's flagship model launched July 16, 2026 — the world's first open 3T-class model. It has 2.8 trillion parameters in a mixture-of-experts design (16 of 896 experts active per token), a 1,048,576-token context window, native image/video understanding, and is built on the new Kimi Delta Attention and Attention Residuals architecture.

When was Kimi K3 released? (Kimi 3.0 release date) +

Kimi K3 officially launched on July 16, 2026, and is available today on Kimi.com, Kimi Work, Kimi Code, and the Kimi API. Some users search for it as "Kimi 3.0" — Moonshot's official name is Kimi K3. The full open-weight release (model weights + technical report) follows by July 27, 2026.

What are Kimi K3's full specifications? +

Kimi K3 specs: 2.8T total parameters (Stable LatentMoE, 16 of 896 experts active per token), 1M-token context window, native image/video understanding, Kimi Delta Attention + Attention Residuals architecture, Gated MLA, MXFP4 weights with MXFP8 activations, max thinking effort by default, and a Modified MIT open-weight license.

Is this the official Kimi AI website? +

No — k3-kimi.com is an independent fan-made guide. The official Kimi AI website is kimi.com, where you can log in (Kimi login) and use K3 free. Official docs live on Moonshot's Kimi API Platform. We just track benchmarks, pricing, and news — and share an invite link that gets you free membership credits.

How do I download Kimi K3 for PC? +

On desktop, download the Kimi Work app (version 3.1.0 or later) for Windows or Apple silicon Mac from the official Kimi site — K3 is built in. Alternatively use kimi.com in any browser with no download, or run Kimi Code in your terminal and pick Kimi K3 with the /model command. Mobile apps are on iOS, Android, and HarmonyOS.

Where can I download Kimi K3 weights? +

Moonshot will publish the full Kimi K3 weights by July 27, 2026 under a Modified MIT license, alongside the technical report. Weights ship in MXFP4 for broad hardware compatibility — but note a 2.8T model needs supernode-class infrastructure (64+ accelerators recommended), so most users should stick to the app or API.

How much does Kimi K3 cost? (pricing) +

Kimi K3 is free in the Kimi app and on kimi.com. API pricing: $0.30 per million tokens for cache-hit input, $3.00/M for cache-miss input, and $15.00/M for output — roughly half the cost of comparable frontier models, with a 90%+ cache hit rate in coding workloads.

How do I get a Kimi API key for Kimi Code? +

Sign up on Moonshot's Kimi API Platform, create an API key, and select the kimi-k3 model. The official docs on the platform cover authentication, endpoints, and the KDA prefill-cache behavior. In the Kimi Code CLI, run /model and choose Kimi K3 — no manual key needed if you're logged into your Kimi account.

Kimi K3 review — is it worth using? +

Our take: yes for long-horizon coding, research, and agentic work. It's #1 on Frontend Code Arena, beats Claude Opus 4.8 and GPT-5.5 on most coding benchmarks, handles 1M-token contexts, and costs about half as much. Weaknesses: still trails Claude Fable 5 and GPT-5.6 Sol in overall polish, and output speed (~62 tok/s) favors long runs over snappy chat.

Is Kimi K3 free to use? +

Yes. You can use K3 free in the Kimi app (iOS, Android, HarmonyOS) and on kimi.com. Through Moonshot's new referral program, signing up via an invite link earns both sides up to 1-year membership credits. API access is paid: $0.30/M cached input, $3/M input, $15/M output tokens.

How do I get free Kimi K3 membership credits? +

Click the invite link on this page (invitation code KRPVU2 is applied automatically), create your free Kimi account, and Moonshot's referral program credits both you and the inviter with a guaranteed benefit — up to 1-year membership credits.

How good is Kimi K3 really? (benchmarks) +

On Moonshot's launch suite, K3 ranks #1 on Program Bench (77.8), SWE Marathon (42.0), SpreadsheetBench 2 (34.8), Automation Bench (30.8) and BrowseComp (91.2), and #2 on Terminal-Bench 2.1 (88.3 vs Sol's 88.8). On the community-voted LMArena it is #1 in Frontend Code with a 76% win rate. Independent Artificial Analysis testing scores it 57 — #4 of 189 models. Moonshot itself says it still trails Fable 5 and GPT-5.6 Sol overall.

Kimi K3 vs Kimi K2 — what changed? +

Nearly everything: 2.8T vs 1.0T parameters, 896 vs 384 experts, 1M vs 256K context, native vision for the first time, the new KDA + Attention Residuals architecture (~2.5× scaling efficiency, up to 6.3× faster long-context decoding), and a jump from #18 to #1 on Frontend Code Arena.

Is Kimi K3 open source? Can I run it locally? +

The full weights release by July 27, 2026 under a Modified MIT license, with MXFP4 quantization for broad hardware support. Realistically, a 2.8T-parameter model needs serious infrastructure — Moonshot recommends supernode deployments with 64+ accelerators. Most people should use the app or API.

What are Kimi K3's limitations? +

Per Moonshot: don't switch models mid-session (thinking-history sensitivity), it can be overly proactive on ambiguous tasks, and it still trails Fable 5 / GPT-5.6 Sol in overall UX. Independent tests also note slower output speed (~62 tok/s) — better for long runs than snappy chat.

What can Kimi K3 actually do? +

Long-horizon agentic work: it built playable games from single prompts, designed a working chip autonomously in 48 hours, wrote a GPU compiler (MiniTriton) from scratch that rivals Triton, reproduced astrophysics research in 2 hours that typically takes 1–2 weeks, and produces interactive research reports and video edits in Kimi Work.

Still curious? Try Kimi K3 free →

Open weights land July 27

Be early to Kimi K3.

The referral program rewards early adopters most. Sign up today — you get up to 1-year membership credits, and so do we.

Sign Up & Claim Free Credits

Free account · No credit card · ~60 seconds