Skip to content

August 05, 2026

Generated: 2026-08-05 18:01 UTC
Total Duration: 1m 49s
Iterations: 1
Judge (classifier) model: gpt-4.1

HolmesGPT is continuously evaluated against real-world Kubernetes and cloud troubleshooting scenarios.

If you find scenarios that HolmesGPT does not perform well on, please consider adding them as evals to the benchmark.

Model Accuracy Comparison

Model Pass Fail Skip/Error Total Success Rate

Results are automatically generated and updated weekly. For detailed traces and analysis, see our Braintrust dashboard.