Sakana's Fugu-Ultra: orchestrating frontier models beats any single one at agentic tasks
Sakana AI’s Fugu project makes a single argument across six wildly different benchmarks: an agent that coordinates several strong models can outdo any individual frontier model on open-ended, code-driven work. The flagship system, Fugu-Ultra, is pitted against three anonymized frontier baselines (Models A, B, and C) in head-to-head trials, and it edges or routs them nearly everywhere.
The demonstrations span autonomous ML research and hard reasoning tasks. Using Karpathy et al.’s AutoResearch loop, which iteratively rewrites training code and keeps only edits that lower validation bits-per-byte, Fugu-Ultra ran 123 experiments in about 14 hours on one H100 and tuned its way to the best mean BPB of the field. It also reconstructed the reading order of a 1610 scattered-script Japanese kana letter far more accurately than rivals (0.80 vs. 0.24 NED, with one model failing to produce working code at all), wrote a from-scratch pure-Python Rubik’s Cube solver that cleared all 300 scrambled cubes near the move-optimal frontier while two baselines crashed outright, generated a functioning CAD mechanical iris where others left gaps or couldn’t close the aperture, won four consecutive blindfold chess games against frontier models and a 2100-Elo Stockfish, and grew a simulated 50-week stock portfolio by a mean 19.4% versus under 15% for competitors.
The consistent failure mode worth noting is that competing models often shipped plausible-looking code that simply didn’t run — a reminder that these are execution-graded benchmarks, not paper evaluations. The trading and reading-order results carry obvious caveats (a single equity, one historical window, one letter), so the broader claim is less ‘Fugu is the best model’ than ‘orchestration is a real axis of capability.’ Anthropic’s own Claude models, including Opus 4.8 and Fable 5, are the current frontier systems an orchestration approach like this would typically draw on, though Sakana keeps the specific baselines anonymized.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.