RC RANDOM CHAOS

Microsoft debuts MAI-Code-1-Flash, a Copilot-tuned model that undercuts Haiku 4.5

· via Hacker News

Original source

MAI-Code-1-Flash

Hacker News →

Microsoft has released MAI-Code-1-Flash, an in-house coding model now rolling out to GitHub Copilot individual users in VS Code through both the explicit model picker and the default auto router. Unlike models tuned primarily for leaderboard scores, this one was trained directly inside the GitHub Copilot production harness, with checkpoints evaluated on real Copilot tasks like repo Q&A, refactoring, and telemetry-grounded workflows.

The headline claim is efficiency: adaptive solution length control lets the model truncate reasoning on easy prompts and expand it on hard ones, reportedly cutting token use by up to 60% on SWE-Bench Verified. Microsoft pitches it against Anthropic’s Claude Haiku 4.5 and reports wins across SWE-Bench Verified, Pro, and Multilingual plus Terminal Bench 2, with a 16-point lead on SWE-Bench Pro (51.2% vs 35.2%) and a 28.9-point lead on IF Bench instruction following.

Microsoft also describes a custom 186-question adversarial benchmark designed to defeat memorization — inverted Monty Hall variants, impossible tasks, underdetermined scenarios — on which MAI-Code-1-Flash hit 85.8% adjusted accuracy. The company concedes weaknesses, noting Einstellung-trap categories still score below 50%, hinting that pattern-matching habits persist even in a model marketed for reasoning.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.