Every week a new model tops the leaderboard. GPT-5.6, Grok 4.5, Fable 5, Opus 5, the conversation quickly turns into "Which model should we standardise on?"
Public benchmarks such as SWE-bench, Terminal-Bench, and the arena leaderboards are incredibly valuable. They measure broad model capability and do an excellent job of reducing a field of dozens of models to a shortlist worth evaluating.
The problem is that enterprise AI systems rarely solve generic benchmark problems. They classify domain-specific content, extract information from proprietary documents, generate structured outputs for downstream applications, and automate business workflows that are unique to the organisation.
Start Smaller Than You Think
Rather than relying on hundreds of synthetic prompts or public datasets, we started with just 18 carefully selected examples from one of our highest-volume business workflows.
Within an afternoon, we had enough signal to compare multiple frontier models against the task that actually mattered to us.
Three Complementary Benchmark Suites
The framework itself is deliberately simple. We evaluate every model using three complementary benchmark suites.

- Domain-specific suite: containing 30–50 representative examples labelled by someone who genuinely understands the business domain. This is by far the most valuable benchmark, yet it's the one most organisations never invest in.
- Smoke suite: a lightweight collection of prompts that runs against every new model and version to detect silent regressions before they reach production.
- Structured-output suite: Because a model can top every public leaderboard and still fail in production if it occasionally wraps JSON in prose, omits required fields, or drifts from the agreed schema.
The Real Asset Is the Gold Labels
The biggest lesson, however, had nothing to do with the models themselves.
The real asset wasn't the evaluation framework or the prompts—it was the Data gold labels. They define what "correct" means for your business, and unlike public benchmarks, they evolve as your products, policies, and customer expectations change.
That became obvious when we reviewed and refined our own gold labels. Without changing the prompts, the models, or the evaluation framework, our leaderboard completely reordered. A model that had previously ranked last suddenly moved to first.
Nothing about the models had changed; our understanding of the business had.
Model Selection Is a Business Decision
That reinforced an important point. Selecting an LLM isn't just a technical decision; "it's a business decision".
Public benchmarks are an excellent way to narrow the field, but they should never be the final authority. The final decision should come from an enterprise benchmark that reflects your own data, workflows, and definition of success.
Public benchmarks tell you what's generally good. Your enterprise leaderboard tells you what's right for your business.
