# Hyper-τ-Bench

> Updated 2026-09-11 · type: tool · category: mcp-agents · status: active · rev 1

Hyper-τ-Bench is Sierra's open-source benchmark for whether an AI agent can research, design and build a working customer-service agent end to end.

- Open source: yes (MIT)
- Self-hostable: yes
- Pricing model: free
- Best for: Teams deciding whether 'have an agent build our support/service agent' is real yet — the benchmark quantifies the gap between automated construction and expert-built systems.
- Last verified: 2026-09-11

- **Canonical:** https://gtmstacker.com/registry/tool/hyper-tau-bench/
- **Source:** [github · sierra-research/hyper-tau-bench](https://github.com/sierra-research/hyper-tau-bench)
- **Tags:** mcp-agents, agent-evaluation, benchmark, agent-construction, customer-service
- **Repository:** https://github.com/sierra-research/hyper-tau-bench

## Is Hyper-τ-Bench open source?

Yes, Hyper-τ-Bench is open source under the MIT license.

## How much does Hyper-τ-Bench cost?

Hyper-τ-Bench is free to use.

## Can I self-host Hyper-τ-Bench?

Yes, Hyper-τ-Bench can be self-hosted (the source is available under the MIT license).


---

Open-source benchmark (MIT, Python) from Sierra Research that tests whether an AI coding agent can BUILD a working customer-service agent — not just operate one: the developer agent gets a sandbox with business records, a codebase and a production API, and must design, wire and ship a functioning agent scored against simulated customers.

## Provenance

- MIT independently WebFetch-verified 2026-09-11 (20★, 4 commits, active; supports Codex, Claude Code, OpenCode harnesses; public leaderboard). Companion 41-page paper submitted to arXiv 2026-09-04. Pass-rate figures (23.9% best automated vs 82.2% expert reference) are Sierra's reported results, not independently reproduced.
- Surfaced via the 2026-09-11 viral-posts brief; curated from the GTM Stacker signal registry (2026-09-11 pass); license independently WebFetch-verified 2026-09-11.

## Why it matters for a GTM stack

"An agent will build your support agent" is the pitch behind a wave of GTM tooling. This is the first open yardstick for that exact claim, and the current numbers argue for buying expert-built or building carefully — useful evidence in any build-vs-buy conversation about agentic customer service.
