Sakana AI says Fugu Ultra beat Anthropic’s Fable 5, but benchmark results vary

The Japan-based startup’s claim on science reasoning and coding tests faces scrutiny as critics say model scores can swing by 10 to 20 points depending on scaffolding.

Summary

verifying reliability

Terms & Concepts

No specialized terms available for this topic.