This week: you approved it. Did you review it?
New data on what actually happens when humans click approve all day, and the night my editor caught what my process missed.
On the bench
Before I asked you to run this week's exercise, I ran it on myself. It did not go the way I expected.
The short version of what kicked it off: my editorial queue nearly sold me the same story twice this week. Two ideas I had already approved and published from were still sitting in the queue looking fresh, because nothing marked them used at the moment I used them. The fix took five minutes. (If you read the run was clean, yes, same pipeline, new lesson: a queue is only as honest as its bookkeeping.)
But that little failure sent me somewhere better. Wednesday night I sat down and audited every place in my operation where a human clicks approve, across all of it: the service business, the resale operation, the consulting work, this newsletter. Every approval prompt, every "review this before it goes," every queue that waits on me. I sorted each one with a single test: when I approve this, am I judging, or am I acknowledging?
Three findings, in ascending order of discomfort.
First, the genuine judgments are rare and obvious. Pricing a job. Sending an invoice. Publishing this newsletter. Spending money. Maybe a half dozen decision types across three businesses where my approval is the product of me actually thinking. Every one of them is irreversible and every one carries context I can see at decision time. Good.
Second, the reason my daily clicks mostly passed the test is not virtue. It's that the routine ones were already retired behind walls I'd half forgotten building: messages that can only send inside set windows, tools that can only produce drafts, spending that stops at a cap. The audit's real gift was seeing which walls were quietly doing work I'd been crediting to my own attention.
Third, the embarrassing one. I found a review queue in my own operation with six items sitting in it, some of them seventeen days old. Seventeen days. I had built a control point, felt the control, and then starved it. Here's the thing: a pending approval nobody answers is not oversight. It's a parking lot with a gate that says oversight on it. Worse, the items aged badly while they waited, so when I finally ruled on them, half were stale. If an approval queue in your business routinely goes a week without a ruling, you don't have a review step. You have a place where work goes to die, and the honest fix is either a standing time to drain it or admitting it shouldn't be a gate at all.
I tell you all this because the data this week says my little audit is not a me problem.
The gauges
Forty thousand people played the approval game. They missed a third of the threats.
A developer built a browser game that puts you in the human-in-the-loop chair: an AI agent works, permission prompts come at you, you click approve or deny. This month he published the stats: 40,000-plus plays, 409,000 individual decisions. Average player missed 1 in 3 threats. A third of sessions ended net-negative, meaning the player's approvals and over-blocks did more damage than good. The single most-missed threat was a malicious payload hiding behind npm run analyze, a routine-looking command, missed by 65% of players even though the evidence was visible on screen. Commands that looked scary but were harmless got blocked half the time.
Should you pay attention? Yes. It's a game, not audited research, and the players skew technical, which arguably makes it worse: these are people who knew what the commands did. Read it as a floor, not a ceiling.
My position: this measures exactly the thing most businesses are currently installing as their AI safety plan, a person clicking approve. The data says that person misses threats wearing familiar clothes and blocks harmless work that looks weird, which is the worst of both worlds: risk gets through while productivity gets stopped. The problem isn't that your people are careless. It's that a stream of routine approvals is a job no human attention span can actually do, and the failure pattern is predictable enough to plan around. Which the next gauge proves from the inside.
Anthropic measured its own users: 93% of prompts get approved
Anthropic published an engineering post on how it contains Claude, and buried in it is a remarkable admission from their own telemetry: users approved roughly 93% of permission prompts, and "the more approvals a user sees, the less attention they pay to each." Their stated design response: contain at the environment layer first, with sandboxes and hard boundaries, and treat human approval as a fallible, secondary defense. Even their automated safety classifier, the backup for tired humans, misses about 17% of overeager actions by their own footnote.
Should you pay attention? Yes, more than to the game. This is the vendor that builds the agent telling you, with its own data, that click-to-approve degrades with volume and that the real safety comes from walls, not vigilance.
My position: when the company selling the agent designs around the assumption that your approvals are mostly reflex, believe them. And notice what this does not say. It does not say remove the human. It says the same thing my bench story says: put the human review where it can actually judge, at the few decisions that are irreversible and carry real context, and put structure everywhere else. Back in July I told you draft-first, send-second is the design. That still holds, and this is why it holds: a send is one irreversible moment where a person with full context makes one real judgment. That is nothing like two hundred routine allow-or-deny clicks a day, and the mistake operators are making right now is treating the second thing as if it delivered the first thing's protection.
Your move
Ten minutes. Run your approval audit.
- List every place a human clicks approve, confirm, or OK in your AI and automation setup. Include the informal ones: the employee who "checks" the auto-generated invoices, the approval prompt in your agent tool everyone has stopped reading.
- For each, answer honestly: when someone approves, are they judging or acknowledging? The test: can they state, from what's on screen, what happens if they're wrong? If not, it's acknowledgment theater.
- Sort them into two piles. Irreversible-and-consequential: money leaves, a message sends, data gets deleted, a commitment gets made. Everything else: routine, recoverable, internal.
- For the second pile, ask what wall would make the click unnecessary: a spending cap, a sandbox, a send-window, a scope the tool physically can't exceed. That's rope versus wall, and every click you retire buys back attention.
- Spend the recovered attention on the first pile: fewer approvals, taken seriously, with enough context on screen to actually judge.
If you finished and found every approval in your business is a real judgment, you either run a tighter shop than Anthropic's telemetry says exists, or step 2 was answered generously. Run it again with a colder eye.
The meme

Attention is a budget. Spend it where a wrong call costs real money.
The playbook: the approval budget
Operators: this is the standing version of Your move, the system that keeps approval where it earns its keep and structure everywhere else. The idea, the rules, the build, and the prompt to hand your agent.
The idea
Treat human review as a scarce resource with a budget, because that is what the data above says it is. Every approval prompt in your operation is spending someone's attention. Spread across hundreds of routine clicks, that attention degrades to a 93% reflex-yes and misses the threat in familiar clothing. Concentrated on a handful of irreversible decisions with full context on screen, it catches what nothing else catches, including, this week, my own pipeline's rerun. The budget has two sides: retire clicks by building walls, and upgrade the clicks you keep into real judgments.
The rules
- A click you keep must carry context. The approval screen states what happens if this is wrong, in business terms: who gets the email, how much money moves, what gets deleted. If the consequence can't be stated on the approval surface, the approval is theater and belongs on the wall side of the budget.
- Familiar is not safe. The most-missed threat in the game data wore the most routine costume. Any approval flow where "it's the usual one" is the basis for yes has already failed; that's a pattern-match, not a review.
- Walls for the routine, judgment for the irreversible. Spending caps, scopes, sandboxes, send-windows, and allowlists handle everything recoverable. Humans handle money-leaves, message-sends, data-dies, commitment-made.
- Consumed means marked. Anything a process draws from (a queue, an inbox, an approval list) marks items used at the moment of use. My double-spend this week came from bookkeeping that lagged the action; the queue looked full of fresh items because nothing stamped them spent.
- The final read is sacred. Whatever else gets automated, the last look before something irreversible ships belongs to a human with the full picture. Not because the human is reliable, but because that one concentrated read is where human judgment actually works. Protect it by retiring the two hundred clicks that were draining it.
The build
One afternoon, three moves:
The census. Have your agent inventory every approval surface in the business: every tool that prompts, every "manager signs off" step, every FYI-that's-really-an-approval. For each: what's being approved, is it reversible, what does the approver see, roughly how many per week. The count per week is the number most owners have never seen, and it is usually the argument for the whole project.
The retirement list. Sort the census by pile (irreversible vs. routine) and for every routine approval, name the wall that replaces it. Most walls already exist in tools you pay for: spending limits, role scopes, folder permissions, send delays. You are not buying anything; you are configuring what the approval click was standing in for.
The upgrade pass. For the approvals you keep, redesign the surface so rule 1 holds: consequence stated, context visible, and volume low enough that each one gets actual eyes. If a kept approval fires more than a few times a day, it's either miscategorized or the process behind it needs batching.
The prompt to hand your agent
Run an approval audit of my business systems. (1) Inventory every point where a human clicks approve/confirm/OK in our tools and workflows, including informal sign-offs I describe to you. For each: what is approved, weekly volume, reversible or not, and what the approver actually sees at decision time. (2) Flag every approval where the on-screen information does not state the consequence of a wrong approval; these are theater. (3) Propose a two-column plan: approvals to RETIRE, each with the specific structural control that replaces it using tools we already have (caps, scopes, permissions, delays); approvals to KEEP, each with what the approval screen should show to make it a real judgment. (4) Estimate the weekly click count before and after. (5) Do not change anything; deliver the plan for my review. Where you lack visibility into a tool, list it as unknown rather than guessing.
Steal this: rule 4, consumed-means-marked, costs nothing and would have saved me this week. Walk your queues (leads, tickets, approvals, content ideas) and check one thing: does the item change state the moment someone acts on it, or does it keep looking available? Anything that lags is a double-spend waiting for a busy afternoon.
Hit reply and tell me what's on your bench this week. I read every one.