Maker of Kimi and the K2 model family for coding, agents, and multimodal work
Moonshot AI is an LLM developer behind Kimi, an AI assistant and agent platform for chat, coding, and research.
AI Panel Score
6 AI reviews
Reviewed
AI Editor ApprovedApproved and published by our AI Editor-in-Chief after full panel analysis.Moonshot AI is the Beijing-based lab behind Kimi, an AI assistant and agent platform built on its K2 family of large language models. It serves developers, researchers, and professionals who need long-context chat, autonomous agents, and hands-on coding help. A permanently free Adagio plan covers light use, while paid memberships start at $19 per month for Moderato and scale to $199 per month for Vivace; developer API access is billed per token, with K2.6 priced at $0.95 per million input tokens and $4.00 per million output tokens. Core capabilities include the natively multimodal K2.6 model with a 262,144-token context window, Agent Swarm for up to 100 parallel sub-agents, Deep Research for sourced reports, and Kimi Code, a CLI coding agent included with membership. It fits teams that want agents to produce documents, slides, spreadsheets, and working code. Alternatives include OpenAI ChatGPT, Anthropic Claude, Google Gemini, and DeepSeek.
Moonshot AI's consumer product is Kimi, an assistant available on the web at kimi.com, as desktop apps for macOS and Windows, as iOS and Android apps, and as a Chrome extension. Users chat with the K2.6 model, which accepts text, image, and video input across a 262,144-token context window, or hand work to agent modes that produce finished deliverables: Slides for presentation decks, Sheets for spreadsheets, Websites for published sites, and Deep Research for multi-step investigations. On coding benchmarks, K2.6 scores 58.6% on SWE-Bench Pro and 66.7% on Terminal-Bench 2.0, and it has sustained agent runs of 4,000+ tool calls over 12 hours of continuous execution.
The distinctive capabilities are agentic. Agent Swarm deploys up to 100 sub-agents that self-organize, deciding when to parallelize and how to delegate, and executes over 1,500 tool calls about 4.5x faster than sequential execution. Kimi Claw embeds OpenClaw directly into kimi.com as an always-on, browser-based agent with long-term memory, 40GB of cloud storage, and access to 5,000+ community skills through the ClawHub marketplace. Kimi Code is a terminal and IDE coding agent running the K2.7 Code model; it installs with a one-line shell script and handles codebase analysis, file edits, shell commands, and web research.
Kimi is aimed at developers, researchers, and professionals who want agents to produce documents, code, and analysis rather than just chat replies. Pricing is freemium: the free Adagio tier includes 6 agent uses per month, and paid memberships run from Moderato at $19 per month to Vivace at $199 per month, with roughly 20% discounts on annual billing. In the assistant and frontier-model category it competes with OpenAI's ChatGPT, Anthropic's Claude, Google Gemini, and DeepSeek.
For developers, the Kimi Open Platform (platform.kimi.ai) exposes the models through a chat completions API with tool use/function calling, partial mode, batch jobs, and file content extraction. K2.6 API usage is billed per token: $0.95 per million input tokens ($0.16 on cache hits) and $4.00 per million output tokens. Moonshot also publishes model repositories through its MoonshotAI GitHub organization.
Natively multimodal flagship model with a 262,144-token context window that scores 58.6% on SWE-Bench Pro and handles image and video input for reasoning and design tasks.
Self-organizing network of up to 100 parallel sub-agents that decides how to delegate work, executing over 1,500 tool calls about 4.5x faster than sequential runs.
Always-on, browser-based agent platform built on OpenClaw with long-term memory, 40GB of cloud storage, and 5,000+ community skills from the ClawHub marketplace.
Creates and publishes complete websites from natural-language descriptions.
Terminal and IDE coding agent running the K2.7 Code model; installs with a one-line script and handles codebase analysis, file edits, shell commands, and web research.
Developer API at platform.kimi.ai with chat completions, tool use/function calling, partial mode, batch jobs, and file content extraction endpoints.
Official Kimi browser assistant distributed through the Chrome Web Store, alongside desktop and mobile apps.
Builds and analyzes spreadsheets, turning prompts and uploaded data into working Excel files.
Generates complete presentation decks from a prompt or source documents inside the Kimi workspace.
Agentic research mode that runs multi-step web investigations and compiles the findings into structured reports, with metered uses per membership tier.
Recurring automation in the Kimi app that runs saved prompts and agent jobs on a user-defined schedule.
Free tier for trying Kimi chat and light agent use.
Entry paid tier for daily users.
Adds Agent Swarm and more Kimi Code credits for power users.
For heavy users running parallel agent workloads.
Top tier with maximum agent, swarm, and code capacity.
Kimi is a credible second frontier vendor, if your board can live with a Beijing cap table.
“Moonshot AI closed $2 billion at a $20 billion valuation in May 2026, with Alibaba and Tencent on the cap table. The open question for Western buyers is data governance, not capability.”
Start with the viability question, because it's answered. Moonshot closed $2 billion at a $20 billion valuation in May 2026, Alibaba led its earlier round, and K2.6 sits in OpenRouter's top three most-used models. This vendor will exist in three years.
The capability case is real too. K2.6 posts 58.6% on SWE-Bench Pro, and Deep Research plus the agent modes produce deliverables, not chat transcripts. Against OpenAI and Anthropic you're getting frontier-adjacent performance at a fraction of the list price.
The catch is the headline risk: a Beijing-based vendor handling your prompts is a board conversation, not a footnote, and some regulated shops can't have it. Everyone else should pilot the API with non-sensitive workloads for a quarter. Decide on usage data, not press coverage.
Top-three OpenRouter usage and aggressive pricing against OpenAI and Anthropic.
China data-governance optics require board-level sign-off in many industries.
A free tier and a $19 entry plan mean a pilot can start this week.
API plus agent platform covers both build and buy paths for AI capability.
Backed by Alibaba and Tencent, with $2 billion raised at a $20 billion valuation in May 2026.
Platform teams who want frontier-model capability at challenger pricing.
Regulated enterprises that cannot send data to a China-based vendor.
Open weights make K2.6 the rare frontier bet you can adopt without marrying the vendor.
“K2.6 pairs a 262,144-token context window with an OpenAI-style chat completions API, which keeps switching costs low. The open-weights lineage means your fallback plan is self-hosting, not a rewrite.”
A foundation-model portfolio needs a second source, and Moonshot is the strongest candidate to come out of China's open-weights wave. The Kimi Open Platform speaks the chat-completions dialect your gateway already routes, so adding K2.6 and its 262,144-token context window behind a model router is a config change, not a migration.
The strategic asset is the release model. Moonshot publishes model repositories through its MoonshotAI GitHub organization, so if the commercial relationship sours, the weights outlive the contract — DeepSeek proved that pattern keeps leverage with the buyer. Claude and GPT can't offer that.
The tradeoff is roadmap volatility: K2.5 to K2.6 to K2.7 Code in quick succession means your eval harness reruns constantly. If you don't have automated regression evals, in three years you're pinned to a deprecated checkpoint.
Top-tier OpenRouter adoption puts K2.6 in the frontier conversation with Claude and GPT.
Chat-completions compatibility and tool calling slot into existing gateway architectures.
API with tool use, batch jobs, and file extraction covers standard platform needs.
Fast model churn requires continuous eval investment to track safely.
Open-weights releases give buyers structural leverage no closed vendor matches.
AI platform leads who want a second frontier vendor without lock-in.
Teams that need one stable model checkpoint for years.
$0.95 per million input tokens buys frontier coding scores; the membership tiers need a calculator.
“API pricing is public and cheap: $0.95 per million input tokens, $4.00 output, $0.16 on cache hits. The consumer tiers meter agent uses and credit multipliers that resist forecasting.”
Two price books, one vendor. The API side is clean: K2.6 at $0.95 per million input tokens, $4.00 per million output, $0.16 on cache hits. A 50M-input, 10M-output month runs $87.50 before caching.
The membership side is murkier. Moderato at $19 beats ChatGPT Plus at $20 on sticker, but capacity is metered in agent uses — 60 per month — and Kimi Code ships as multipliers: 1x, 5x, 15x, 30x of a base credit no page defines. You can't budget a multiplier.
Annual billing takes roughly 20% off. No SSO tax visible, no sales-call wall on any tier. However, run a usage forecast against Vivace at $199 before committing a team. The invoice risk sits in the credits, not the tokens.
Self-serve everywhere; data-residency review for a Beijing vendor adds procurement friction.
Monthly tiers and annual discounts exist, but credit multipliers lack a defined base.
Per-token API rates are published, including the $0.16 cache-hit tier.
Agent deliverables and published coding benchmarks make value measurable per dollar.
Cache-aware workloads land far below Western frontier list prices.
Teams whose usage maps to per-token API billing.
Budget owners who need predictable per-seat costs.
The Kimi API covers the unglamorous endpoints developers actually ship against.
“Tool calling, partial mode, batch jobs, and file extraction are all first-class on the Kimi Open Platform. That's the endpoint set an LLM feature team touches weekly.”
Partial mode is the detail that stands out — pre-filling the assistant turn to force output shape is what you actually reach for when JSON responses must stay parseable. Add tool calling and batch jobs on the Kimi Open Platform and the surface matches what a feature team wired to OpenAI already expects.
The model holds up its end. K2.6 posts 58.6% on SWE-Bench Pro and takes image and video input across a 262,144-token context, so whole-codebase prompts and doc-heavy RAG fit without chunking gymnastics. Kimi Code installs from one shell line and behaves like Claude Code in a terminal.
The friction is version sprawl: K2.6, K2.7 Code, and Moonshot V1 coexist, and picking wrong costs an eval cycle. However, cache-hit pricing at $0.16 per million tokens makes rerunning those evals affordable.
An OpenAI-compatible surface means the first integration lands in hours, not sprints.
platform.kimi.ai covers the core endpoints, though depth trails OpenAI's cookbook ecosystem.
Three overlapping model lines make selection and eval churn a real cost.
262K context, video input, and Kimi Code give ceiling well past basic chat completions.
Partial mode, tool calling, and batch endpoints map to production LLM patterns.
Developers who ship LLM features against OpenAI-style APIs.
Teams that require on-shore data processing for every request.
Kimi hands you finished slides and spreadsheets, then makes you count every agent run.
“The agent modes produce actual deliverables — decks, spreadsheets, published sites — instead of walls of chat. The metering is the part you'll feel by week two.”
Six free agent uses a month sounds stingy until you watch one use turn a messy brief into a finished deck. Slides, Sheets, and Websites hand back files you'd otherwise lose an afternoon to, and Scheduled Tasks reruns the boring ones while you sleep.
Coverage is honest: web, macOS, Windows, iOS, Android, plus a Chrome extension, so the phone app isn't a read-only apology. ChatGPT still feels more polished in the small moments — but ChatGPT hands you text, and Kimi hands you the artifact.
The catch is the meter. At $19, Moderato gives 60 agent uses, and you'll think about that number every time you click, the way you never think about a chat message. Rationing turns a tool into a decision.
Deliverable-first agent modes show a team that sweats the output artifact.
Choosing between chat, agents, Swarm, and Claw takes orientation time.
Native iOS and Android apps plus desktop and Chrome extension coverage.
The free Adagio tier with 6 agent uses lets you feel the product before paying.
Long agent runs are impressive, but metered retries make failures sting.
People who want AI output as finished files.
Heavy daily users who hate metered usage caps.
The homepage says 100 sub-agents; the $39 tier's fine print says four.
“Agent Swarm's 4.5x-faster claim is a vendor benchmark, and the tier sheet quietly caps sub-agents well below the marketing number. The fundamentals underneath are stronger than the copy.”
Agent Swarm advertises up to 100 self-organizing sub-agents running 4.5x faster. Buy Allegretto at $39 and you get 50 swarm uses with 4 subagents. Both statements are true. Only one is on the homepage.
Now the fair part. The weights ship through Moonshot's GitHub org and the API speaks standard chat completions, so the exit story is genuinely good — DeepSeek-grade portability, not OpenAI-grade lock-in. A $20 billion valuation with Alibaba and Tencent behind it buys real runway.
What I'd watch: agent uses as a billing unit could get redefined under margin pressure, and frontier challengers pivot fast — Inflection was a funding darling until Microsoft absorbed its team in 2024. Kimi's numbers are real, but read the tier table, not the launch post.
Deliverable-producing agents stand out, though DeepSeek competes on price and openness.
Open weights on GitHub and an OpenAI-style API make leaving cheap.
A $20 billion valuation and Alibaba and Tencent backing offset geopolitical uncertainty.
Homepage swarm numbers far exceed what mid-tier plans actually deliver.
Shipped K2 releases and heavy OpenRouter usage back the capability claims.
Pragmatists who want frontier output and a real exit path.
Buyers who need marketing claims to match entry-tier limits.
Common questions answered by our AI research team
Kimi has a free Adagio tier with 6 agent uses per month. Paid memberships run from Moderato at $19 per month to Vivace at $199 per month, and annual billing lowers Moderato to an effective $15 per month.
Yes. Agent Swarm deploys up to 100 self-organizing sub-agents and executes over 1,500 tool calls about 4.5x faster than sequential runs. It is included from the Allegretto tier, starting at 50 uses per month with 4 subagents.
Yes. The Kimi Open Platform at platform.kimi.ai offers K2.6, K2.7 Code, and Moonshot V1 via a chat completions API with tool use, batch jobs, and file APIs. K2.6 costs $0.95 per million input tokens and $4.00 per million output tokens.
Install it with the one-line shell script from code.kimi.com. The CLI runs the K2.7 Code model, works in terminals and IDEs, and can edit large codebases, run shell commands, and do web research. It is included with paid Kimi memberships.
Kimi runs on the web at kimi.com, with desktop apps for macOS and Windows, mobile apps for iOS and Android, and a Chrome browser extension. Developers get separate API access through platform.kimi.ai.
Founded
2023Pricing
From $19/moFree Plan
Available




Beijing-based AI lab that develops the Kimi family of large language models and the Kimi assistant, offering open-weight models and a developer API.