Switching Bluebox to Sonnet 5.5: the eval results
Bluebox runs on Anthropic's Sonnet 5.5 model since October 1. In our evals with SREGym, Sonnet 5.5 diagnosed more problems correctly than Sonnet 5 (86.7% vs 73.3%), used 3.3x fewer tokens, finished 3.7x faster, and cost 4.3x less per investigation. Why SREGym We evaluate Bluebox continuously, against our own benchmarks and against public ones. SREGym is a public, open benchmark of fault scenarios. Each scenario has a target application on Kubernetes and an injected fault, with a known ground