RC RANDOM CHAOS

VibeThinker-3B: tiny model claims frontier-level math and code reasoning

· via Hacker News

Original source

VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO

Hacker News →

A new technical report introduces VibeThinker-3B, a 3-billion-parameter dense model built to test how far verifiable reasoning can be pushed at small scale. The authors report results that would normally require models orders of magnitude larger: 94.3 on AIME26 (rising to 97.1 with claim-level test-time scaling), 80.2 Pass@1 on LiveCodeBench v6, and a 96.1% acceptance rate on recent, unseen LeetCode contests. They position these numbers alongside flagship systems like DeepSeek V3.2, GLM-5, and Gemini 3 Pro, while a 93.4 on IFEval suggests the aggressive reasoning tuning did not erode instruction-following.

The gains come from a post-training pipeline rather than raw scale — curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation, layered on the team’s earlier 1.5B work and a ‘Spectrum-to-Signal’ paradigm. From this the authors propose a Parametric Compression-Coverage Hypothesis: verifiable reasoning compresses into a compact ‘reasoning core,’ while broad open-domain knowledge and general competence still demand wide parameter coverage over facts and long-tail cases.

The significance, if the results hold up to independent replication, is that small models may be a complementary path to frontier reasoning rather than just cheaper deployment substitutes — relevant for anyone weighing on-device or cost-constrained inference. The usual caveat applies: these are self-reported benchmarks from a single technical report, and verifiable-task scores say little about general-purpose knowledge, where the authors themselves concede small models fall short.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.