WorldbyFlow•Structured Research
Generated July 26, 2026· technology· 24 sources
Profile: Kimi K3 — Open-Weight Frontier Language Model
Product Profile
Bottom Line
Kimi K3 is the largest open-weight model ever released — a 2.8-trillion-parameter MoE that benchmarks within 3–5 points of the closed frontier at one-third the price, with open weights dropping July 27, 2026; the simultaneous distillation dispute with Anthropic and the White House means the model's global distribution and the US government's response to it will both occur in the same week, making the next seven days the highest-stakes period in its short life.
Open-weight frontier mixture-of-experts language model · Moonshot AI (Beijing)
Maker: Moonshot AI (Beijing) · vKimi K3 — API-available July 16, 2026; open weights committed by July 27, 2026 · Since 2026 · GENERALLY AVAILABLE
Also known as: Kimi 3, kimi-k3, moonshotai/kimi-k3
Brief
Kimi K3 is Moonshot AI's flagship large language model, released on July 16, 2026, as the world's largest open-weight model at 2.8 trillion parameters — surpassing the prior record held by DeepSeek's 1.6-trillion-parameter model. Built on a Mixture-of-Experts architecture with Kimi Delta Attention and Attention Residuals, it activates only 16 of 896 experts per token, enabling frontier-class performance at inference costs comparable to mid-tier closed models. It profiles for long-horizon coding, agentic knowledge work, and multimodal reasoning, with a 1-million-token context window and native vision. The model warrants profiling now because it debuted at the top of Arena.ai's Frontend Code leaderboard, triggered US government distillation allegations and sanctions threats against Moonshot AI, and is scheduled to release open weights on July 27, 2026 — a distribution event that, once it occurs, cannot be reversed by any export control.
Product Profile
Kimi K3 is a 2.8-trillion-parameter MoE model with always-on reasoning, native vision and video input, and a flat 1-million-token context window. It is positioned as a frontier-capable alternative to closed proprietary models at substantially lower API pricing, available both through Moonshot's hosted surfaces and — from July 27, 2026 — as open weights for self-hosting. The architecture prioritizes long-horizon agentic coding and complex knowledge workflows over simple conversational queries.
Total parameters · verified
2.8 trillion (MoE)
Active parameters per token · verified
16 of 896 experts activated per forward pass
Context window · verified
1,048,576 tokens (1M), flat pricing across full window
Attention architecture · verified
Kimi Delta Attention (KDA) — hybrid linear attention — with Attention Residuals (AttnRes)
Quantization · reported
MXFP4 weights, MXFP8 activations; quantization-aware training from SFT stage onward
Modalities · verified
Text, images, video (native); video is K3-only in the Kimi model family
Reasoning mode · verified
Always-on; reasoning_effort currently locked to max on the API — no lite or fast mode yet available
KEY CAPABILITIES
- Long-horizon agentic coding: #1 on Arena.ai Frontend Code Arena (1,679 Elo, July 16, 2026) across 6 of 7 tracked domains, per Moonshot-reported and community-observed results
- 1M-token flat context with no per-tier pricing penalty for long prompts — differentiates against Claude Fable 5 and Gemini, which charge more above a threshold
- Native multimodal input: text, image, and video supported in a single model (video is K3-exclusive in the Kimi lineup)
- Agent-first surface integration: Kimi Code CLI (KFC), Agent Swarm (parallel task execution), and Kimi Slides accessible across the product suite
- Open-weight release under Modified MIT license (scheduled July 27, 2026), enabling self-hosting, fine-tuning, and local deployment on sufficient hardware
KEY LIMITATIONS
- Self-hosting hardware barrier is severe: minimum ~8× H100 80GB GPUs (~1.4TB in MXFP4 quantization, ~18 H100 80GB at full precision) — beyond consumer or small-team reach
- Reasoning verbosity inflates output token counts: with reasoning_effort locked to max and no lighter mode yet available, effective API cost per task can exceed headline $15/M output rate
- Independent benchmark reproduction unavailable prior to July 27 open-weight release — all pre-release scores are vendor-reported or derived from API-only access, not independently verified on weights
- Distillation allegation unresolved: White House and Anthropic have alleged covert extraction of Claude model outputs; Moonshot AI denies it; statistical evidence published by a Redwood Research researcher (July 2026) is suggestive but not conclusive — provenance risk is live for regulated enterprise deployments
- Independent hallucination testing (TechTimes, July 24, 2026) found a 51% hallucination rate in one evaluation — not independently corroborated and methodology not fully disclosed, but the figure was omitted from Moonshot's own benchmark disclosures
Adoption Picture
Kimi K3's consumer launch in China generated enough demand that Moonshot briefly stopped accepting new users — a supply-side constraint, not an opt-out. Developer reception in the West has been strongly positive on coding benchmarks. The open-weight release scheduled for July 27, 2026 is expected to sharply expand adoption among self-hosting enterprises and researchers globally, but also makes usage effectively unmonitorable after distribution.
USER BASE
Consumer users on Kimi.com and mobile (China-heavy, some international); developer API users building coding agents and agentic workflows; enterprise teams evaluating for Copilot-alternative workloads (Microsoft reportedly evaluating K3 for Copilot document summarization and coding tasks, per Yourstory and Memeburn reporting, July 2026 — unconfirmed by Microsoft)
BenchLM public leaderboard rank
#5 of 215 models, score 79.98/100 with 61 published benchmark scores (BenchLM, July 25, 2026)
Artificial Analysis Intelligence Index
Score of 57 (third globally, behind Claude Fable 5 at 60 and GPT-5.6 Sol; per Artificial Analysis, July 2026)
Arena.ai Frontend Code Arena Elo
1,679 (#1 on launch day, July 16, 2026 — ahead of Claude Fable 5; community-observed, not vendor-claimed)
GROWTH SIGNALS
Service demand exceeded capacity at China consumer launch (July 16, 2026), forcing temporary suspension of new user registrations
Demand-side signal comparable to DeepSeek's January 2025 launch surge; suggests rapid organic adoption among Chinese knowledge workers
Frontend Code Arena: 17-place jump from Kimi K2.6 (#18) to K3 (#1) on launch day, July 16, 2026
Arena.ai leaderboard is judged by developer preference in blind pairwise matchups — not vendor-controlled, making this a more credible adoption signal than a vendor-run benchmark
Cloudflare Workers AI integrated Kimi K3 as a named model (developer docs live as of July 25, 2026)
Platform-level integration accelerates developer access without requiring direct Moonshot API relationships; expands the distribution surface before open weights drop
Micron analyst commentary citing Kimi K3 demand surge as a tailwind for HBM and DRAM consumption (Seeking Alpha, Yahoo Finance, July 24, 2026)
Supply-chain-level signal of infrastructure demand being pulled by K3 usage in China — historically correlated with sustained adoption rather than one-time spikes
Release History
Moonshot AI's Kimi model line began as a consumer chatbot and pivoted sharply toward open-weight, frontier-class models starting in July 2025 with Kimi K2. K3 represents the third generation of this strategy, roughly doubling the parameter count over K2 while introducing a novel attention architecture. The open-weight release cadence — API first, weights weeks later — is a deliberate pattern that builds hosted revenue before relinquishing the model to the self-hosting community.
Current: Kimi K3 — API generally available July 16, 2026; open weights scheduled July 27, 2026Last major update: 2026-07 — K3 launch with new KDA+AttnRes architecture, 2.8T parameters, and 1M-token context
RELEASE TIMELINE
2026-07-16 · vKimi K3
API launch of 2.8T-parameter MoE flagship; #1 on Arena.ai Frontend Code Arena on day one; open weights committed by July 27, 2026
2026-04 · vKimi K2.6
Intermediate release in the K2 family; ranked #18 on Frontend Code Arena before K3's arrival
2026-01 · vKimi K2.5
1-trillion-parameter MoE with multimodal and agentic capabilities; first K2-family model with vision and thinking modes
2025-10 · vKimi Linear
48B-parameter MoE introducing Kimi Delta Attention (KDA) — the attention architecture later scaled into K3
2025-07 · vKimi K2
1-trillion-parameter open-source MoE; Moonshot's strategic pivot to open weights; released under modified MIT license
RECENT UPDATES
2026-07Cloudflare Workers AI adds Kimi K3 as a supported model (model ID: moonshotai/kimi-k3)
↳ Extends K3 API access to Cloudflare's global edge network without a direct Moonshot relationship
2026-07vLLM publishes production-scale K3 support preview with Triton kernel implementation for KDA+AttnRes fusion
↳ Enables enterprise self-hosting on the dominant open-source inference framework; prerequisite for large-scale deployment
2026-07Kimi.com pricing page notes upcoming split of Kimi consumer and Kimi Code into separate subscription products
↳ Signals monetization restructuring as Kimi Code becomes a standalone revenue line
Competitive Landscape
Kimi K3 occupies a distinct position: it is the only model in the 2.8-trillion-parameter class with open weights, and it prices at roughly one-third the per-token rate of its nearest closed-weight competitor at comparable capability. Its primary benchmark rivals are Claude Fable 5 and GPT-5.6 Sol — both closed-weight, both more expensive. DeepSeek's V4 Pro remains the primary open-weight alternative, with 1.6 trillion parameters and an even lower price point but narrower capability on agentic tasks per current benchmarks.
ECOSYSTEM POSITION
Kimi K3 is currently the apex of the open-weight model tier — larger than any prior open release and benchmarking within 3–5 points of the closed frontier. Its strategic position is that it forces closed-model vendors to compete on dimensions other than capability exclusivity, since K3 at open weights makes near-frontier performance freely available to anyone with sufficient GPU infrastructure. The distillation dispute also positions Anthropic and the US government as active adversaries, adding regulatory and reputational complexity to the competitive picture.
COMPETING PRODUCTS
Claude Fable 5 / Fable 5 Max
Anthropic
Scores 60 on Artificial Analysis Intelligence Index vs K3's 57; beats K3 on most breadth and professional-knowledge benchmarks but costs approximately 3.3× more per token and is closed-weight
GPT-5.6 Sol
OpenAI
Scored 1,747.8 on GDPval-AA v2 (vs K3's 1,687); closed-weight with multi-effort-tier family (Sol/Terra/Luna) offering routing flexibility K3 currently lacks
DeepSeek V4 Pro
DeepSeek
1.6T-parameter open-weight MoE — largest prior open-weight record holder; lower price point but placed behind K3 on agentic coding benchmarks per current evaluations
Llama 4 family
Meta
Open-weight with broad deployment ecosystem; smaller parameter counts than K3 but benefits from Meta's distribution infrastructure and established fine-tuning community
INTEGRATIONS
- Cloudflare Workers AI (model ID: moonshotai/kimi-k3)
- OpenRouter (API aggregator access)
- vLLM (open-source inference framework — preview support July 22, 2026)
- Kimi Code (KFC) — Moonshot's own coding CLI and agent harness
- OpenCode Go — third-party agent harness
DEPENDENCIES
- NVIDIA H100/H200 GPU infrastructure for self-hosted deployment (minimum ~8× H100 80GB for full weights)
- Hugging Face — planned distribution channel for open weights (July 27, 2026)
- MXFP4/MXFP8 quantization tooling for practical deployment at Q4 precision
Commercial Model
Kimi K3 monetizes through two parallel tracks: a consumer subscription (Kimi.com, five tiers from free to $199/month) and a pay-per-token developer API ($3/$15 per million input/output tokens). The open-weight release on July 27, 2026 creates a third track — zero-marginal-cost self-hosting — which Moonshot is accepting as a strategic trade-off to build ecosystem adoption and developer lock-in at the infrastructure layer. Kimi and Kimi Code subscriptions are slated to separate into distinct products, suggesting a forthcoming monetization restructuring.
Revenue: hybrid — consumer subscription (freemium with four paid tiers) plus usage-based API billing; open-weight release is a deliberate platform play, not a revenue pathDistribution: Managed cloud (Moonshot-hosted at kimi.com and platform.kimi.ai) primary; third-party managed cloud (Cloudflare, OpenRouter) secondary; open-weight self-hosted (July 27, 2026 scheduled)
PRICING TIERS
Adagio (Free)$0/month
Unlimited basic chat, file uploads, web search, ~6 agent credits; no Agent Swarm, no Kimi Code CLI
Moderato$19/month (monthly billing); lower effective rate on annual billing
K3 access, expanded agent credits, Kimi Slides, Swarm access
Allegretto$39/month (monthly billing)
Higher agent credit multiplier and concurrency vs Moderato
Allegro$99/month (monthly billing); unlocks K3's full 1M-token context
Full 1M context window, highest consumer-tier agent concurrency
Vivace$199/month (monthly billing); unlocks K3's full 1M-token context
Maximum agent concurrency and credit multiplier; 1M context
API — Kimi K3 (pay-as-you-go)$3.00/M input tokens (uncached), $0.30/M input tokens (cached), $15.00/M output tokens — as of July 2026 per Moonshot official pricing
Full K3 capabilities including vision, 1M context, always-on reasoning; separate from subscription tiers
- 90% prompt-cache discount ($0.30 vs $3.00 per million cached input tokens) is structured to reward high-volume API customers with stable system prompts — typical of enterprise/agent workflow targeting
- Imminent Kimi / Kimi Code subscription split signals intent to create a dedicated developer revenue line at higher price points than consumer tiers
- Microsoft evaluation for Copilot workloads (reported July 2026, unconfirmed by Microsoft) would represent a major OEM-style licensing or API revenue event if confirmed
- Launch credit promotion (10–30% bonus on prepaid API balances, running July 15 – August 11, 2026) is a standard demand-seeding mechanism to build usage data and developer habit before open weights commoditize the API
Outlook
The July 27, 2026 open-weight release is the single most consequential near-term event for Kimi K3: once weights are globally distributed, the model becomes effectively irreversible from a policy perspective regardless of any US sanctions or export controls that follow. The distillation dispute with Anthropic and the White House is unresolved and could produce sanctions, Entity List restrictions, or downstream liability — none of which would retrieve already-distributed weights. Independently, the commercial trajectory depends on whether Microsoft's reported Copilot evaluation converts, whether the Kimi/Kimi Code subscription split attracts enterprise coding teams, and whether Moonshot can demonstrate independent architectural quality once open weights enable third-party inspection of the MoE structure.
ROADMAP
Open-weight release on Hugging Face under Modified MIT license· 2026-07-27
Enables global self-hosting and fine-tuning; makes any subsequent sanctions or export controls largely ineffective against already-distributed weights
Kimi / Kimi Code subscription separation into distinct products· Qualitative — near-term per in-app banner notice (as of July 2026)
Creates a dedicated developer monetization track; pricing and feature split not yet disclosed
Lighter reasoning effort modes (non-max reasoning_effort API parameter)· No date committed — Moonshot has signaled intent without a timeline
Unlocks lower-cost, lower-latency workloads currently inaccessible due to mandatory max reasoning; critical for high-throughput classification and summarization use cases
Third-party independent evaluation on open weights· 2026-08 (shortly after July 27 weight release)
First opportunity to independently verify or challenge vendor benchmark claims; hallucination rate, distillation evidence, and coding benchmark reproducibility all hinge on this
WATCH SIGNALS
Confirmation of open-weight availability on Hugging Face by July 27, 2026, including license terms and full MoE checkpoint
A delay or partial release would suggest regulatory intervention or Moonshot capacity constraints; a clean release validates the open-weight strategy and resets the competitive calculus for every lab with a closed flagship
White House or Treasury Department formal sanctions or Entity List action against Moonshot AI
Would signal US willingness to use export-control infrastructure against AI labs on distillation grounds — a precedent affecting the entire open-weight AI sector, not just Moonshot
Independent third-party reproduction of K3 coding benchmarks on open weights using a non-KimiCode harness
All pre-weight benchmarks use Moonshot's own KimiCode harness; independent results on SWE-bench or FrontierSWE will either validate or substantially revise K3's coding-leadership claim
Microsoft official confirmation or denial of Kimi K3 integration into Copilot
A confirmed OEM-scale deployment would be the largest enterprise adoption signal in K3's short life and would materially weaken OpenAI's default position inside Microsoft's product stack
Anthropic or Redwood Research publication of formal distillation evidence tied specifically to Kimi K3 (vs the February 2026 findings about earlier Moonshot activity)
Current public evidence does not directly connect the documented 3.4-million-exchange extraction campaign to K3's training data; direct evidence would escalate the legal and regulatory exposure substantially
PRODUCT / COMPETITIVE RISKS
- Distillation allegations — if the White House and Anthropic produce direct evidence that K3's training data included covertly harvested Claude Fable 5 outputs, Moonshot faces potential Entity List designation, sanctions, and reputational damage that could unwind enterprise adoption built on the K3 launch
- Open-weight irreversibility — once July 27 weights distribute globally, Moonshot loses control of the model's deployment and Anthropic/US government cannot recover IP embedded in the weights; this creates asymmetric risk for Moonshot's future legal exposure
- Hardware concentration for self-hosting — the 8× H100 80GB minimum for full K3 deployment limits self-hosting to well-capitalized organizations; Q4 MXFP4 quantized deployment (~18 H100 80GB equivalent) remains beyond most teams, keeping actual adoption concentrated in the hosted API and Cloudflare paths
- Benchmark reproducibility risk — all pre-weight scores use Moonshot's KimiCode harness; independent evaluations frequently show 10–20% performance gaps versus vendor-reported scores on open-source coding benchmarks; K3's coding leadership claim has not yet been independently confirmed
- Regulatory ban risk — Tom's Hardware and PCMag coverage notes the White House is reportedly reviving efforts to ban Chinese AI models; a formal prohibition on K3 API usage in US federal contexts or critical infrastructure would sharply curtail enterprise adoption
medium uncertainty· model's epistemic confidence in this analysis
BOTTOM LINE
Kimi K3 is the largest open-weight model ever released — a 2.8-trillion-parameter MoE that benchmarks within 3–5 points of the closed frontier at one-third the price, with open weights dropping July 27, 2026; the simultaneous distillation dispute with Anthropic and the White House means the model's global distribution and the US government's response to it will both occur in the same week, making the next seven days the highest-stakes period in its short life.