Models & Research

Together AI: GLM-5.3 ties Claude Fable 5 at a fifth of the price

Together AI published two more DeepSWE comparisons this week, and the headline result is a statistical tie with a 5.4x price gap inside it. Z.ai’s GLM-5.3 solved 69.0% of tasks on the first attempt against 69.7% for Claude Fable 5, a 0.7 point difference that sits inside the error bars on both sides. GLM-5.3 did it at $3.99 a rollout. Fable cost $21.63.

Measured as solved work per dollar, that’s 17 solves per $100 for GLM-5.3 against Fable’s 3. Retries widen the gap instead of closing it. GLM-5.3 leads pass@2 at 81.1% to 77.1% and pass@4 at 87.6% to 84.1%, which means the cheaper model also holds the higher ceiling. Together’s write-up puts it flatly: there’s no attempt count at which paying 5.4x for Fable buys more coverage.

The GPT-5.6 Sol matchup ended differently, though not by much. Sol kept the single-shot crown at 72.7%, but GLM-5.3 tied it at two attempts and passed it at four. The best row on the board was the cascade Together already demonstrated with its DeepSeek runs: GLM-5.3 goes first, Sol takes over when the test suite rejects the answer, and the pair solves 85.9% of tasks at $6.61 each. That beats Sol alone at $8.37, and it beats a perfect one-shot router at 83.8%.

Modelpass@1pass@4Cost per rollout
GLM-5.3 (max)69.0%87.6%$3.99
GPT-5.6 Sol (max)72.7%85.8%$8.37
Claude Fable 5 (max)69.7%84.1%$21.63
Together AI, 452 rollouts per model across 113 DeepSWE tasks, four trials each.

The routing detail cuts finer than the totals. GLM-5.3 posted the best JavaScript score on the board at 90% to Fable’s 75%, while Fable answered with Rust at 85% to 70%, the widest single-language gap in the matchup. Protocol conformance is GLM-5.3’s one clear hole at 44%. On failures, Sol broke tests that already passed in 20% of its misses, against 11% for both GLM-5.3 and Fable.

Against Fable, though, the routing argument mostly collapses. The two models correlate at 0.65 per task, the highest agreement Together measured, so they succeed and fail on largely the same problems.

They are near-substitutes, and when two models substitute, you keep the lower-cost one.

Zain Hasan and Shobhit Dixit, Together AI

GLM-5.3 itself is barely a week old. Z.ai launched it on 14 August as “the most capable open-weights model for coding”, though the weights were held back for safety hardening, due on Hugging Face two weeks after launch. Nathan Lambert at Interconnects puts the model at roughly 750 billion parameters, a third the size of Kimi K3, and credits the jump to scaled post-training. That is why the public DeepSWE leaderboard reads so differently generation to generation: 44% for GLM-5.2 against 69% for GLM-5.3.

Fable 5 sits at the other end of the market. Anthropic launched it in June as a Mythos-class model, and Together calls it the most expensive model measured in these runs. But what the premium buys, on this evidence, is a genuine Rust and serialization specialty rather than a general lead.

The caveats from Together’s earlier comparisons still apply. Together sells hosted inference for open-weight models like GLM-5.3, so it isn’t a neutral referee, and every figure comes from its own analysis of the published per-trial records. The public leaderboard is close but not identical, listing claude-fable-5 at 70% and gpt-5.6-sol at 73%, with overlapping error bars all round.

What’s worth watching is the last week of August. If the GLM-5.3 weights land on Hugging Face on schedule, the cheapest row on this table becomes a model you can serve yourself, and the cascade math gets rewritten by whoever hosts it cheapest. The catch hasn’t changed since the DeepSeek version, though: the cascade only works when your test suite can actually reject a bad answer.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *