Behind the Build 001: The tool nobody planned, that is used the most

There’s a task on most teams that everyone dreads and no one questions. For one person I work with, it’s checking product packaging before it ships. Not the design itself. The claims sitting on top of it.
-0:00
What we built: an agent that reads packaging drafts in whatever language they’re in, checks the claims on them against compliance guidelines, and flags what needs a second look before anything ships.

Her job is hard to put a single label on. Not quite a designer, not quite a marketer. She sits at the intersection of both, with a compliance lens over everything. Packaging usually isn’t reinvented from scratch each time, it’s adapted: same product, different market, different language. But adaptation is where things get quietly dangerous. A word that’s neutral in English can sound like a promise in another language. A sentence that fits the layout in one length turns vague or overstated once it’s shortened to fit another. None of that is a translation error exactly. It’s the gap between what a sentence says and what it starts to imply once it’s squeezed into a different shape.

I didn’t find this by looking for a use case. I found it by asking what part of her week she’d quietly rather not do. That question took a while to actually answer. A task done long enough stops feeling like a problem and starts feeling like just the job.

The first version of what we built wasn’t it. It tried to flag too much, and everything looked urgent, which is its own kind of useless. It also missed the plainest thing: it needs to actually read the design like a person would, not just the text pulled out of it, since a claim can be technically fine and still sit next to an image that changes what it means. Trimming it down, and getting the reading part right, took longer than building the first version did.

What it does now is small on purpose. Upload the draft, and it reads the design, checks the claims against the relevant guidelines, and sends back a short report, three flags at most, ranked by how much they matter.

It doesn’t decide anything. It can’t. Too much of this is judgment: what a market will tolerate, what a claim implies next to a specific image, what’s worth a fight with a supplier and what isn’t. She still signs off on every version. What changed is what lands on her desk first. Most of the time now it’s close to clean, one small thing to check, instead of a long pass through every version by hand.

The agent itself is not finished. It doesn’t look particularly polished yet, and it probably won’t for a while, since it keeps getting nudged forward in small ways rather than built out all at once. And yet it has already become one of the most used agents we’ve built, and likely the most liked piece of AI work in that company so far. Not because it does the most. Because it removed something that had been quietly wearing people down for years.

It’s stayed with me since. Not the tool itself, but how quietly it was hiding. Nobody had flagged it as a priority, nobody had a name for the problem, it was just Tuesday. Worth asking around your own team what that equivalent might be. You might not get a straight answer the first time you ask. But when you do, there’s usually something real underneath it.

Other interesting reads

The quiet cost of plausible work

AI has made it easy to produce work that looks finished. The question worth sitting with is who now carries the weight of deciding whether it actually is.

Which task rebuilds itself every Thursday

Multi-agent workflows are having a moment. For most small agencies, the useful question is much smaller and much closer to home.

An AI tool is not the same as an AI process

Most organisations have added AI to their workflows. Fewer have built one. The difference shows up quietly, but it shows up consistently.