ServiceNow AI tests failure-guided training data for enterprise agents
AutoSynthData generates and checks enterprise-agent tasks around weaknesses found in a target model, with reported gains in two controlled environments.
Training on the tasks an agent still misses
ServiceNow AI described AutoSynthData on October 2 as a way to create training tasks from weaknesses observed in an enterprise agent. The team evaluates a target model in a stateful environment, compares its failures with a stronger teacher’s successful runs, and turns the gaps into new tasks. Each task includes a system specification, a user request and a verifier. The pipeline checks whether the task can be completed, whether a reference solution works and whether incorrect outcomes fail verification. It then uses accepted tasks for supervised fine-tuning and can shift the next batch toward remaining weaknesses.
ServiceNow reports tests in two EnterpriseOps Gym domains. In its Hybrid experiment, it says 2,000 generated training samples raised mean Pass@1 by 7.2 percentage points, while verifier success rose from 63.01% to 68.55%. In a separate ITSM experiment, it reports mean Pass@1 rising from 18.77% to 27.18% after training on 1,994 generated samples. These are the team’s results in controlled benchmark environments, not independently established improvements for deployed customer agents. The report describes the method and links the Gym dataset; it does not establish that the AutoSynthData pipeline itself is publicly available as a product.