Models & Research

ServiceNow’s synthetic data lifts Gemma to 27.18% on ITSM agent tasks

ServiceNow CoreAI published a synthetic data pipeline on Friday that it says lifted a 26B Gemma model from 18.77% to 27.18% on enterprise IT service tasks. Read next to ServiceNow’s own benchmark paper, that number lands about 1.3 points below the 28.5% the best of 14 frontier models scored on the same domain. The gain is real, and what it buys is a small open model that still fails nearly three tasks in four.

The pipeline is called AutoSynthData, and the write-up on Hugging Face describes it as a loop rather than a generator. It runs a target model and a stronger teacher on evaluation tasks, then records where the target fails and how the teacher succeeds. Those findings get distilled into what the authors call sanitized capability specification cards.

The generator never sees the original evaluation prompts, entities, trajectories or verifier details. It only gets the cards, which is the design choice that keeps the training set from collapsing into a copy of the test set. That matters because the usual failure mode of synthetic data is quietly memorising the benchmark you meant to beat.

Even so, a card is only a prompt, which means each candidate task then has to survive three checks. A difficulty filter keeps tasks the target solves on no more than one of three trials while the stronger solver manages at least two of three. A positive gate replays the reference solution to confirm it passes the verifier, and a negative gate can mutate the expected outcome to confirm those states no longer pass.

Two runs, and a 3.7x difference in generation time

ServiceNow ran that whole loop twice, in the Hybrid and ITSM domains of its EnterpriseOps-Gym environment. The target was Google’s Gemma-4-26B-A4B-it both times, and the teacher changed.

RunTeacherSamplesGeneration timeReported result
HybridQwen3.8-27B2,000about 18 hoursmean Pass@1 up 7.2 points, 35% relative; verifier success 63.01% to 68.55%
ITSMDeepSeek-V4.1-Flash1,99466 hoursmean Pass@1 18.77% to 27.18%
Source: ServiceNow CoreAI, AutoSynthData, 2 October 2026.

Those two rows carry a detail worth pulling out. Six fewer samples took 3.7 times as long, which is our calculation from the 66 hours and 18 hours the post reports. ServiceNow attributes the gap to a larger teacher model and to the ITSM run predating throughput optimisations.

What 27% means on this benchmark

EnterpriseOps-Gym is ServiceNow’s own work, released in March and described by Malay and eight co-authors. It holds 1,150 expert-curated tasks across eight domains, with 512 tools and 164 database tables in a containerised sandbox. The code repository notes that tasks run against live MCP servers and are scored by SQL verifiers checking final state, not action sequences.

That paper evaluated 14 frontier models, and the scores are the context for 27.18%. Claude Opus 4.5 took the best average at 37.4%, Gemini-3-Flash came second overall at 31.9% and topped ITSM at 28.5%. On 30 deliberately infeasible tasks, the best model refused correctly only 53.9% of the time.

Our findings underscore that current agents are not yet ready for autonomous enterprise deployment.

Malay et al., EnterpriseOps-Gym, arXiv

So the tuned 26B model sits within about 1.3 points of the strongest frontier ITSM score, on a benchmark its own authors say nothing is ready for. One caveat belongs here: the blog reports mean Pass@1 and the paper reports average task completion, and neither document states that the two runs share an identical evaluation setup. Treat the comparison as a scale check, not a leaderboard entry.

There’s a second gap in the reporting. The Hybrid result is given as a relative improvement and a verifier success rate, with no absolute Pass@1 figure in the text, and the post says the checkpoint closes 59% of the original Pass@1 gap between Gemma and the reference model without naming that reference model. The best Hybrid checkpoint was epoch 5.

The result is narrower than the number suggests, and ServiceNow is explicit that the experiments cover supervised fine-tuning only, and that extending the same loop to reinforcement learning is planned rather than done. The released dataset is Apache-2.0 and was downloaded 21,446 times last month, so anyone can check whether a second lab reproduces the gain. Until someone does, this is one environment, one target model and an agent that is still wrong most of the time.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

Rundowns AI Desk covers artificial intelligence: model releases, research, funding and policy. Every story is written from primary sources, with each claim linked to the announcement, filing or paper it came from, and checked against those sources before publication.

Leave a Reply

Your email address will not be published. Required fields are marked *