AI tools are rapidly advancing in monitoring IT systems. Yet a new study from Datadog and Carnegie Mellon University shows that engineers maintain a significant lead in handling complex incidents. Based on real failures observed in production, this test compares several advanced models to human specialists. The results above all reveal the current limits of the models when faced with critical and unforeseen situations.

In brief
- A Datadog study shows that AI models remain less efficient than engineers in handling complex IT failures.
- The tests are based on 63 real incidents and more than 5 million data points from production emergencies.
- GPT-5 dominates general models with 62.7% accuracy, but human experts still reach 72.7%.
- Researchers believe that collaboration between humans and AI could greatly improve incident response in the future.
AI is progressing, but remains limited in the face of complex incidents
Tech companies are now introducing AI agents that can automatically analyze production incidents, despite recent advances made by these models. These systems must help teams detect anomalies and identify the origin of failures. However, the ARFBench benchmark shows that this automation still remains imperfect. The project is based on real incidents observed during emergency situations, with data manually validated to avoid artificial scenarios.
L'study is based in particular on several key figures:
- 63 real incidents analyzed from Slack exchanges in emergency situations;
- 750 questions created around the incidents studied;
- 142 monitoring indicators used in the benchmark;
- Over 5 million data points examined.
The tests evaluate both anomaly detection and the models' ability to understand complex relationships between multiple metrics. GPT-5 achieves an F1 score of 47.5% on the most difficult questions, while maintaining an overall accuracy of 62.7%. The researchers also point out that trillions of dollars are lost each year due to system failures, which reinforces the strategic importance of AI tools in modern digital infrastructures.
Engineers maintain a clear lead over current models
Faced with model results, human engineers maintain better overall accuracy. Experts in the field obtained a score of 72.7%, well above the best models tested. Even Datadog's non-experts achieved 69.7%, higher than automated systems.
These results show that engineers interpret the overall context of an incident even better. They more easily understand the interactions between several technical signals and the unusual behavior of infrastructures.
No AI model has managed to exceed human performance benchmarks. However, some specialized systems are gradually narrowing the gap. The hybrid Toto-1.0-QA-Experimental model, developed by Datadog, achieves an accuracy of 63.9%. This system combines an internal forecasting model with Qwen3-VL 32B.
In anomaly detection, Toto even obtains an F1 score at least 8.8 points higher than other competing models. This result confirms that a specialized model on observability data can better respond to a specific technical task than a general system.
Despite these advances, engineers remain essential during critical incidents. Models sometimes lose business context, ignore certain metadata, or misinterpret multiple indicators simultaneously.
A collaboration between AI and humans becomes the most credible scenario
Above all, the study highlights that the errors of humans and those of models are different. AI systems detect certain anomalies quickly, while humans better understand ambiguous situations and operational constraints.
Researchers explain that these differences create complementary skills. Models sometimes miss details of context, while humans make more errors on precise timestamps or complex instructions.
To measure this potential, the researchers imagined an “expert oracle” capable of systematically choosing the best response between a human and an AI. In this theoretical scenario, accuracy increases to 87.2%, with an F1 score of 82.8%.
This result does not yet represent a concrete product. However, it shows that collaboration between artificial intelligence and engineers could greatly improve the management of computer incidents in the coming years. Automated systems therefore seem intended to assist human teams rather than completely replace them in the short term.
Maximize your Tremplin.io experience with our 'Read to Earn' program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.
