Moonshot's Kimi K2.7-Code cuts thinking tokens 30% in open-weights coding model
Moonshot AI has released Kimi K2.7-Code on Hugging Face, an open-weights model tuned for agentic coding work. Built on Kimi K2.6 with the same architecture as the K2.5/K2.6 line, its headline improvement is token efficiency: thinking-token usage drops roughly 30% versus K2.6 while the company claims stronger end-to-end completion on long-horizon software engineering tasks. The model ships with native INT4 quantization, the same approach used in Kimi-K2-Thinking, which matters for anyone trying to serve a model of this class on constrained hardware.
Moonshot benchmarks it against GPT-5.5 (Codex, xhigh mode) and Claude Opus 4.8 (Claude Code, xhigh mode) across a mix of in-house and external suites. These include Kimi Code Bench V2 (realistic engineering tasks spanning 10+ languages), Program Bench (reimplementing programs like FFmpeg and SQLite from only a compiled binary and docs, judged by 248,000 fuzz-generated tests), MLS-Bench for ML research capability, and agentic tool-use benchmarks like MCP-Atlas and MCPMark-Verified. As usual with vendor-run evaluations, the in-house benchmarks should be read with some skepticism until independent results land.
Deployment targets vLLM, SGLang, and KTransformers, with OpenAI- and Anthropic-compatible APIs available through Moonshot’s platform. Notable constraints: the model forces thinking mode on (instant mode is unsupported), recommends temperature 1.0 and top-p 0.95, and supports a 262K-token context. The release continues the pattern of Chinese labs shipping competitive open-weights coding models that undercut closed frontier offerings on cost and self-hosting flexibility.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.