Somewhere around week four at Maesa, a designer turned a laptop toward me with a model leaderboard open and asked which one the team should be standardizing on for packaging copy. Fair question. I asked a duller one back: if the model you picked was wrong for this work, how would you find out?
Nobody in the room had an answer. Not because the team wasn't sharp. These are people who have shipped more retail packaging than I ever will, across 12+ brands, into every category you can name. They had a very high standard and they enforced it every day in line review. What they didn't have was a way to point that standard at a model and get a straight answer.
So that afternoon we stopped talking about models and started pulling files. Forty-odd naming and packaging jobs from the previous eighteen months. For each one: the original brief, the version that actually got approved, and whatever note explained why the earlier versions didn't. That folder became their benchmark. It is the least glamorous artifact from the whole engagement, and one of the two or three most used. Maesa now runs launches at roughly $280K less per launch, in 3 months instead of 9. Their VP of Marketing, Oshyia Savur, put the cumulative figure at "tens of millions" from the stage at Shoptalk. Some of that is agents. Some of it is process. A quiet amount of it is simply that they stopped arguing about which output was better and started checking.
The leaderboard is measuring somebody else's job
Every's Katie Parrott wrote about this on August 25, and she's pointing at the same gap from the enterprise side: teams spending enormous sums on AI without ever building what Mercor's CEO calls "offline evals" — a fixed set of real tasks you run a model against before you commit to it. Her examples are legal clauses and house style. Mine are naming conventions, retailer spec sheets, and whether a line of body copy sounds like the brand or sounds like a brand.
Public benchmarks are not lying to you. They're just answering a question you didn't ask. They measure competition math and code puzzles and reading comprehension, and a creative team looks at that and tries to reason its way from "scores well on hard problems" to "will write packaging copy that survives legal." There is no path between those two sentences. You cannot infer it. You have to test it.
Parrott makes one point I'd tattoo on the wall of every creative department: a high acceptance rate can hide the work. Every's own editing tool scores 85–90% on suggestions accepted, and she notes that number says nothing about how much cleanup the human still does after accepting. Creative teams get fooled by this constantly. "The AI got us 80% there" is a sentence I hear in week one of almost every engagement, and it almost always means the last 20% took longer than the whole job used to.
What actually goes in the set
Pick 40 to 60 finished jobs. Not your best work — your normal work. Roughly half should be routine, a third genuinely hard, and the rest the ones you barely got right, where the client came back twice and the third version squeaked through.
Each entry needs three things. The brief as it was actually given, mess included, not the tidy version you'd show a prospect. The output that got approved. And the reason the rejected versions were rejected, in one sentence, in the words of the person who rejected them. That third piece is the one teams skip and the one that carries all the signal. "Too safe." "Reads like the competitor." "Legal will kill the claim." That's your quality bar written down, probably for the first time.
Keep the failures. A test set made only of wins tells you nothing about a model's floor, and the floor is what wrecks a launch calendar.
Score it the way you already score work
Don't invent a rubric. You have one — it lives in the head of whoever signs off. Get it onto paper as four or five pass/fail criteria and use those. Pass/fail, not one-to-ten. A ten-point scale invites people to average their way to a shrug.
Then measure the thing nobody measures: residual work. For each output, how many rounds from here to shippable? That's the number that decides whether a model earns a place in the line. A model that lands at "pretty good" in one pass and needs three rounds of surgery is worse than one that lands at "adequate" in one pass and needs one.
Run the set when a new model ships, when you onboard a brand whose voice sits outside your usual range, and once a quarter regardless. The ground moves faster than most creative directors realize — Parrott cites Vercel data showing open-source models going from 28% to 62% of token share in two months. Whatever you standardized on in the spring is a decision worth re-testing, not a permanent choice.
Give it to the person who says no
Here's where teams get this wrong. The benchmark goes to the AI enthusiast, because they're the one who cares, and within a month it's a personal science project nobody else trusts.
It belongs to whoever owns quality. The executive creative director. The head of production. The person whose approval the work needs anyway. Their standard is the thing being encoded, so their hands should be on it, and the rest of the team should watch them run it at least once. That's the whole-team point I keep making in a different costume: an eval set that one person maintains is a preference. An eval set the department has watched get run is a standard.
It also settles arguments that otherwise burn weeks. When two people disagree about whether the AI version is good enough, you don't have a debate. You have a folder.
The cost of not having one
Across the engagements I've run, the average production-time reduction is about 90%. None of that came from picking the best model off a chart. It came from teams that could tell, quickly and without a meeting, whether a given output cleared their bar — and could therefore move.
The teams that stay stuck are running on taste-by-anecdote. Someone tried a model on a Tuesday, it wrote something great, and now it's the house standard for eight months. That's not a strategy. It's a memory.
Your last fifty approved jobs are sitting in a folder on a server right now. That's a better benchmark than anything on a leaderboard, and it took you years to build. Use it.

