August 05, 2026¶
Generated: 2026-08-05 18:01 UTC
Total Duration: 1m 49s
Iterations: 1
Judge (classifier) model: gpt-4.1
HolmesGPT is continuously evaluated against real-world Kubernetes and cloud troubleshooting scenarios.
If you find scenarios that HolmesGPT does not perform well on, please consider adding them as evals to the benchmark.
Model Accuracy Comparison¶
| Model | Pass | Fail | Skip/Error | Total | Success Rate |
|---|---|---|---|---|---|
Results are automatically generated and updated weekly. For detailed traces and analysis, see our Braintrust dashboard.