Models & Research

ORCA-bench: GLM-5 invents a root cause in 66% of hard incidents

ORCA-bench, a new oncall benchmark from Cornell Tech, Traversal and Columbia University, reports that the weakest of five frontier agents named an implausible root cause in roughly 40% of its incident reports. The paper’s own appendix puts that figure at 66.1% once you look only at the Hard tasks, the vague-report setting closest to a real oncall page. That is the number the summary page doesn’t print.

The benchmark pairs 1,079 root cause analysis tasks with six days of metrics, logs and traces from an OpenTelemetry-instrumented microservice system under continuous simulated user load. Agents query that recorded history through Prometheus, Jaeger and OpenSearch via Grafana, with full access to the application source code. Faults come from toggling 11 of the OpenTelemetry demo’s own feature flags, such as productCatalogFailure, rather than from a synthetic fault library.

Scores fall as the user’s complaint gets vaguer, which means the realistic settings are the weak ones. The best RCA accuracy, which requires naming every plausible root cause, is 25.3% on Medium tasks and 10.0% on Hard. Pooled across all 884 incident tasks, Claude Sonnet 4.6 leads on accuracy at 30.9% and GPT-5.5 leads on RCA depth at 48.8%.

ModelAccuracy, all 884Invented a cause, all 884Accuracy, HardInvented a cause, Hard
Claude Sonnet 4.630.9%14.8%8.6%13.6%
Claude Opus 4.728.6%12.6%10.0%12.1%
GPT-5.524.3%25.0%8.6%22.5%
GLM-517.6%40.2%1.1%66.1%
DeepSeek-V4-Pro15.0%7.2%5.4%12.1%
RCA accuracy and hallucination rate, from ORCA-bench Tables J.1 and J.4. Hard tasks give the agent the least specific user report.

But the hallucination column is where the difficulty ladder bites hardest. GLM-5’s rate climbs from 21.5% on Easy tasks to 34.2% on Medium and 66.1% on Hard, while its accuracy collapses to 1.1%. GPT-5.5 moves the other way, inventing causes most often on Easy tasks at 28.8%, which is why a single pooled number flatters some models and punishes others. That pattern is the familiar problem with headline benchmark figures.

DeepSeek-V4-Pro looks like the cautious one at 7.2%, and it isn’t. It also posts the lowest RCA depth at 19.7%, and the authors say it produces more empty reports on harder tasks. On the 195 control tasks, where no incident is present at all, the scoring rules state that “an empty report is a trivial pass”. So silence scores well here, even though it resolves nothing.

The catch is that source code matters more than the agents act like it does. Removing it drops RCA accuracy by 9 to 16 percentage points and raises the hallucination rate for every model tested. Yet Claude Opus 4.7 and GLM-5 spend only 16% and 20% of their commands reading source code, against 72% and 70% on telemetry queries. Between 26% and 40% of those telemetry calls either error out or come back empty.

Since real production systems are orders of magnitude larger, more dynamic, and more idiosyncratic, the gap we report underscores the engineering work still needed before agents can be entrusted with production reliability.

Gong et al., ORCA-bench

The release is partial, and that matters if you plan to score your own stack. The public split carries 755 tasks across 77 incidents, not the 1,079 the paper scores, so leaderboard entries won’t line up with the table above. The evaluation code and the 50 GB, six-day testbed are both public.

There’s a reason to care beyond the leaderboard. When Boston Consulting Group polled 1,261 managers in January, 22 percent said their organisations had added AI agents to the org chart, WIRED reported, and the same researchers found managers caught 18 percent fewer errors when told the work came from an AI employee rather than an AI tool. That matters because a confident wrong diagnosis is the failure mode that survives that kind of review.

Claude Fable 5 does better, on a much smaller sample. It reaches 40.6% RCA accuracy on the 32 incident tasks of the hand-checked Verified subset, against 25.0% for Claude Opus 4.7, but the standard error is 8.8 points and the gains come mostly from Medium tasks rather than Easy ones. Whether that holds across all 884 tasks is the thing to watch, because the gap between a controlled testbed and a live system has caught other benchmarks out before.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

Rundowns AI Desk covers artificial intelligence: model releases, research, funding and policy. Every story is written from primary sources, with each claim linked to the announcement, filing or paper it came from, and checked against those sources before publication.

Leave a Reply

Your email address will not be published. Required fields are marked *