Sierra said on September 8, 2026, that it is open-sourcing hyper-τ-bench, a long-horizon benchmark scoring whether AI coding agents can construct a working customer-service agent. The strongest automated configuration passed 23.9% of held-out evaluation tasks, Sierra reported, against 82.2% for a reference pairing an engineer with a frontier model. From Acting as an Agent to Building One Sierra built the original τ-bench in 2024 to answer a question it said felt novel at the time: whether a…