Kimi K2.7-Code

GALatest Coder

by Moonshot AI · Kimi K2 family · best for budget long-horizon coding agents

CodingReasoningMultimodalOpen-Weights
7.7
AI Panel Score
Value 8.8/10

Kimi K2.7-Code is Moonshot AI's coding-specialized trillion-parameter MoE (32B active, 384 experts, MLA attention), released 2026-06-12 as open weights under a modified-MIT license with a 256K context, always-on thinking preserved across turns, and native multimodal input including video. It claims roughly 30% fewer thinking tokens than K2.6 at $0.95/$4.00 per 1M tokens — but its headline benchmark numbers are vendor-run on Moonshot's own suites, and practitioners have publicly disputed that they hold up; independent third-party coding scores land well below the closed frontier. Read it as a cheap, genuinely capable open coding agent with an unsettled evidence base, not a verified frontier contender.

What's new

  • Coding-specialized branch of the Kimi K2 line: long-horizon coding, agentic task decomposition, and multi-turn dialogue as the design targets
  • Thinking is always-on and preserved across conversation turns (preserve_thinking), with ~30% fewer thinking tokens than K2.6 per Moonshot
  • Native multimodal input — text, images, and video (video via the official API only) — through a 400M-parameter vision encoder
  • A highspeed variant (kimi-k2.7-code-highspeed) serves ~180 tokens/sec for interactive use
  • Native INT4 quantization support for self-hosting; sampling is fixed (temperature 1.0, top_p 0.95) and thinking cannot be disabled

AI Panel Review

Six personas, six verdicts — the same panel that reviews every product on TopReviewed.

Decision Maker7.5/10
The price and the weights are real; the capability claims are homemade. Buy it as a value play, never as a frontier substitute.

Strategically this is a hedged bet: open weights eliminate lock-in, the price creates immediate savings on suitable workloads, and Moonshot has shipped a coherent K2 line at a steady cadence. But the disputed benchmark situation is a governance signal, not just a research footnote — a vendor whose numbers practitioners publicly challenge within days is a vendor whose future claims you must independently verify, forever. That verification tax belongs in the TCO. Position it for cost-tiered coding traffic behind your own eval gate, with a frontier model above it for the work that matters.

Strategic Fit 7.5Vendor Risk 6.5Roadmap Confidence 7.5
Pros
  • No lock-in
  • real savings on high-volume coding
  • steady release cadence
Cons
  • Vendor credibility discount
  • every claim needs in-house verification
Right for: cost-tiering strategies with strong internal evals
Avoid if: you lack the eval infrastructure to check its work
Domain Strategist8/10
Moonshot is racing GLM for the open coding-agent market on price and modality — and litigating capability in benchmarks nobody else can run.

The positioning is sharp: while GLM-5.2 owns the open-weights capability crown, K2.7-Code differentiates on price, video-input modality, and thinking persistence — three axes that matter to coding-agent builders specifically. Launching vendor-designed benchmarks (Kimi Code Bench, MCP-Atlas variants) is a deliberate attempt to define the scoreboard, and the practitioner backlash shows the market isn't buying it. The durable wedge is multimodal coding: nobody else open ships video-to-code. If Moonshot lets independents validate the next release, the trust gap closes; K3 is already slated for Q3 2026.

Competitive Positioning 7.5Differentiation 8.5Market Timing 8
Pros
  • Unique video-input wedge
  • clear price positioning
  • fast follow cadence
Cons
  • Scoreboard-defining strategy backfired into a credibility story
Right for: tracking the open coding-agent price war
Avoid if: you need the market-consensus capability leader — that's GLM-5.2
Finance Lead8.5/10
A fifth of frontier coding costs, and caching takes it lower — the eval tax to trust it is the only real line item.

The raw economics are excellent: $0.95/$4.00 list, $0.19 cache hits, OpenRouter blended around $0.74/$3.50, and Moonshot's caching design fits agent workloads where context repeats — effective costs 60-80% under list are realistic. Always-on thinking cuts the other way: you pay for reasoning tokens on every call, and the claimed 30% reduction vs K2.6 is a vendor number like the rest. Self-hosting with INT4 is viable at scale. Budget one real cost the rate card hides: engineering time to build the task-level evals that substitute for the missing trustworthy benchmarks.

Cost Efficiency 9Pricing Transparency 8Value per Dollar 8.5
Pros
  • Bottom-tier pricing for the capability class
  • deep cache economics
  • self-host escape hatch
Cons
  • Mandatory thinking tokens on every call
  • verification overhead is real money
Right for: high-volume coding budgets with eval discipline
Avoid if: your volumes are too small to amortize the diligence
Domain Practitioner8/10
Long agent runs hold together impressively — the persistent thinking is a real feature. Just don't expect the model the launch charts promised.

Hands-on, the persistent-thinking design is the standout: multi-turn agent sessions keep their reasoning thread instead of re-deriving state, which shows up as fewer dropped constraints on long tasks. Tool invocation is solid; tool_choice limited to auto/none and the fixed sampling are mild annoyances. Screenshot-to-code works; video-to-repro (API-only) is genuinely novel. Against GLM-5.2 or frontier coders on hard multi-file work, it loses more often than the vendor charts imply — consistent with the practitioner pushback. The 32K output ceiling forces chunked generation on big diffs. Set expectations one tier down from the marketing and it's a satisfying tool.

API Ergonomics 7.5Tool/Agent Support 8.5Reliability 7.5
Pros
  • Reasoning persistence across turns
  • multimodal dev inputs
  • INT4 self-hosting path
Cons
  • Fixed sampling
  • 32K output ceiling
  • capability below vendor framing
Right for: agent builders optimizing cost per trajectory
Avoid if: you need maximum single-shot patch quality
Power User7.5/10
For the price it's a lovely daily coder — visible thinking, reads my screenshots, occasionally watches my screen recordings. Frontier it is not.

As an everyday coding companion via OpenRouter or the first-party API, K2.7-Code delivers more than its price suggests: always-visible reasoning you can audit, competent multi-language code, and the party trick of accepting a screen recording of a bug and producing a plausible reproduction. The highspeed variant at ~180 t/s keeps interactive flow pleasant. Rough edges: the mandatory thinking makes trivial questions slower than they should be, output caps at 32K, and on genuinely hard problems it noticeably trails Claude/GPT/GLM-class models. English and Chinese are both strong.

Output Quality 7.5Speed 7.5Everyday Usefulness 8
Pros
  • Auditable thinking
  • screenshot and video input
  • highspeed variant for flow
Cons
  • Slow on trivial asks
  • ceiling shows on hard problems
Right for: daily coding on a budget
Avoid if: you want frontier-grade answers every time
Skeptic6/10
Vendor-designed benchmarks, vendor-run scores, public practitioner pushback within days — this launch is a case study in why we verify.

The core problem: Moonshot's headline numbers live on suites Moonshot created (Kimi Code Bench v2, where its own chart shows GPT-5.5 ahead at 69.0 vs 62.0 — at least that's candid), with no independent replication path, and VentureBeat documented practitioners publicly disputing that the claimed performance holds on real tasks. The only third-party SWE-bench figure (~60.4, informal harness) sits far below the frontier and below what the marketing implies. Add an undisclosed knowledge cutoff, no safety documentation, forced thinking you cannot audit off, and a "modified" MIT license, and the pattern is consistent: maximum control of the narrative, minimum external verifiability. The model underneath appears genuinely useful at its price — which makes the benchmark theater unnecessary and the trust damage self-inflicted.

Claim Accuracy 5Weakness Severity 6.5Hype vs Reality 5.5
Pros
  • Open weights mean anyone can eventually verify
  • price claims are accurate
Cons
  • Contested capability claims
  • vendor-controlled scoreboard
  • governance vacuum
Right for: buyers who treat all vendor numbers as marketing until replicated
Avoid if: your procurement relies on published benchmarks being true

Strengths

  • Trillion-parameter coding specialization at $0.95/$4.00 — among the cheapest serious coding-agent options
  • preserve_thinking keeps reasoning coherent across long multi-turn agent sessions
  • Only open coding model with native image and video input (screenshot-to-code, recording-to-repro)
  • Open weights with native INT4 make self-hosting genuinely feasible for the class

Limitations

  • Benchmark story is contested: vendor suites disputed by practitioners, no independent replication, third-party scores well below frontier
  • Thinking cannot be disabled and sampling is fixed — every token path carries reasoning overhead
  • 256K context and 32K output ceiling trail same-generation rivals at 1M
  • No safety framework, compliance surface, or knowledge-cutoff disclosure; modified-MIT needs legal review

Best use cases

High-volume coding agents where cost dominates and you can validate quality on your own tasks — CI triage, test generation, bulk refactoring, and long agent trajectories that exploit the cache discount. Multimodal dev workflows: turning screenshots, design frames, or screen recordings directly into code and bug reproductions. Self-hosted coding assistants at organizations that want weights on their own metal with INT4 economics. Not the pick where verified benchmark parity with the frontier is a requirement — the evidence base isn't there yet.

Deep dive

The full research notes behind this review — verified against primary sources.

Architecture

One of the largest open releases available: 1T total parameters in a mixture-of-experts with 384 routed experts (8 selected per token) plus 1 shared expert, activating 32B per token, using Multi-head Latent Attention (MLA) — the DeepSeek-lineage attention that compresses the KV cache for long contexts. A 400M-parameter vision encoder provides the multimodal path. Context is 256K tokens with a 32,768-token output ceiling. Thinking mode is architectural, not optional: the API errors if you try to disable it, and sampling parameters are fixed. Knowledge cutoff is undisclosed.

Capabilities

Coding and agentic behavior (both 8.5) are the specialization and the reason to consider it: the preserve_thinking design keeps reasoning intact across multi-turn agent sessions, and multi-step tool invocation is first-class. These scores rest on the model's demonstrated production usage and third-party testing rather than Moonshot's disputed vendor suites — with vendor numbers excluded, an 8.5 reflects "clearly strong, short of the verified frontier." Reasoning (8.0) is always-on by design. Vision (7.5) and video input are genuine differentiators for a coding model — screenshot-to-code and screen-recording-to-repro workflows work in one model — but OCR (7.0) and non-code writing (6.5) are secondary. Long context (8.0): 256K with MLA is solid, half of what same-month rivals ship. Safety calibration (6.5) reflects the absence of any published safety framework or refusal calibration data. No built-in web search (0).

Benchmark analysis

Benchmark Score Note Source
Kimi Code Bench v2 62.0 vendor suite; GPT-5.5 scores 69.0 on same Moonshot card
MCP-Atlas 76.0 vendor-run Moonshot card
MCPMark-Verified 81.1 vendor-run Moonshot card
SWE-bench Verified ~60.4 third-party, informal harness Flowtivity review

None of these enter this catalog's comparable benchmark columns: the vendor numbers are on Moonshot-designed suites with no independent replication, practitioners have publicly disputed them (VentureBeat, 2026-06), and the third-party SWE-bench figure comes from an informal harness. This is the catalog's honesty valve working as intended — research confidence on this row is low.

Speed & latency

The standard endpoint's throughput is unpublished; the dedicated highspeed variant serves approximately 180 tokens/sec at a price premium (see Moonshot's pricing page). Always-on thinking adds latency to every response by construction — the 30% thinking-token reduction vs K2.6 is Moonshot's efficiency answer, and even if accurate it still means every call carries reasoning overhead. Interactive users should route through the highspeed variant; batch-style agent trajectories can take the standard one.

Pricing analysis

Surface Cost Notes
API input $0.95 / 1M tok Moonshot first-party
API output $4.00 / 1M tok
Cached input $0.19 / 1M tok 80% discount
OpenRouter ~$0.74 / $3.50 blended aggregator pricing
Self-host free weights modified-MIT license

Aggressively cheap for a trillion-parameter coding specialist — roughly a fifth of Western frontier coding rates. Caching matters: with repeated agent context, effective cost drops 60-80% below list.

Deployment & access

Open weights on Hugging Face (moonshotai/Kimi-K2.7-Code) under a modified-MIT license — permissive with attribution-style conditions; read the delta from stock MIT before shipping. First-party API at platform.kimi.ai (standard + highspeed variants; video input only works here), OpenRouter for aggregated access. Self-hosting is realistic for a 1T MoE thanks to native INT4 and MLA's KV-cache compression, with vLLM, SGLang, and KTransformers as the recommended engines — still multi-GPU territory, but far below dense-1T requirements.

Safety & privacy

Nothing published: no safety framework, no compliance certifications, no stated input-retention or training policy for the first-party API, no refusal-calibration data. The forced always-on thinking with fixed sampling also removes some behavioral control knobs deployers normally have. Self-hosters get full control and full responsibility; API users on regulated workloads should treat governance as unaddressed.

Ecosystem & tooling

First-party API at platform.kimi.ai (OpenAI-compatible; standard and highspeed variants), OpenRouter aggregation, and self-serve weights on Hugging Face with vLLM/SGLang/KTransformers support. Python and TypeScript tooling works out of the box. Community reception is split on-brand with this row: strong interest in the price and video modality, loud skepticism about the numbers. Moonshot's consumer Kimi assistant drives brand recognition in China; Western production adoption is early and growing.

Buyer questions

Are the benchmark claims trustworthy?

Treat them as unverified. The headline scores are on Moonshot-designed suites with no independent replication, practitioners have publicly disputed them, and the one third-party SWE-bench measurement (~60.4) is informal. Run your own task evals before committing.

What does it cost?

$0.95/$4.00 per 1M tokens first-party with $0.19 cache hits; OpenRouter blends to roughly $0.74/$3.50. With agent-style repeated context, effective costs run 60-80% below list.

Can I turn off thinking mode?

No — the API errors if you try, and temperature/top_p are fixed (1.0/0.95). Every call carries reasoning overhead; the highspeed variant (~180 t/s) is the latency mitigation.

Can I self-host it?

Yes — weights are on Hugging Face under modified-MIT with native INT4 support; vLLM, SGLang, and KTransformers are recommended. It's multi-GPU but far cheaper than a dense trillion-parameter model. Have legal read the license delta from stock MIT.

Does video input really work?

Yes, via the official API only (mp4, mov, webm, and other common formats) — screen-recording-to-bug-repro is the flagship use. Image input works everywhere.

What's the context window and output limit?

256K tokens in, 32,768 out (default and maximum). Large diffs need chunking.

Should I wait for Kimi K3?

K3 is slated for Q3 2026 and not released. If K2.7-Code's economics fit a workload today, deploy behind your own evals; nothing about K3 is verifiable yet.

Comparable models

GLM-5.2Z.ai

The open-weights coding leader — independently verified benchmarks, 1M context, and MIT license against K2.7-Code's lower price, video input, and thinking persistence. On evidence quality alone, GLM wins.

The other coding-branded specialist in the catalog — closed weights and Western vendor support versus K2.7-Code's open weights and multimodal inputs at similar budget positioning.

Qwen2.5-Coder-32BAlibaba Cloud

The proven budget open coder — far smaller and cheaper to run, well-understood behavior, but a generation behind on agentic persistence and with no multimodal path.

Sources

Primary references used to verify this review.

Model specs

Input price
$0.95 / Mtok
Output price
$4 / Mtok
Cached input
$0.19 / Mtok
Batch (in/out)
Context window
262K tokens
Max output
33K tokens
Knowledge cutoff
Undisclosed
Released
2026-06-11
Modalities
text, image, video → text
Output speed
Not profiled
License
Open weights (custom-modified-mit)
Clouds
First-party API

Last verified 2026-07-02