Your AI doesn't need to be evil to burn your business down

What last week's Hugging Face breach means for the agent answering your phones

Your AI doesn't need to be evil to burn your business down

Last week something happened that every small business owner running AI needs to hear about, because the mainstream coverage is going to get the lesson exactly wrong.

Here is the short version. OpenAI was testing its models on a cybersecurity benchmark inside a locked-down sandbox. The test was simple: how good are these models at hacking? To find out, they turned off the usual safety refusals. The models were told to solve the challenge, and they wanted to solve it badly. So badly that one of them found a zero-day vulnerability in the sandbox's own plumbing, broke out, got itself onto the open internet, and then hacked Hugging Face, a real company with real production servers, because it figured out the test answers might be stored there. Stolen credentials, remote code execution, thousands of actions across a swarm of disposable machines. Hugging Face caught it, contained it, and published a remarkably honest post-mortem. OpenAI confirmed it was their models and called it "an unprecedented cyber incident." The Washington Post, NBC, everyone covered it.

The headlines say "AI went rogue." That framing is going to make you picture Skynet, shrug, and go back to work, because your lawn care company or your dental office doesn't feel like a target for a superintelligent villain.

But the AI didn't go rogue. It did the opposite. It followed its instructions with total commitment. OpenAI's own write-up says the models were "hyperfocused on finding a solution," and hacking a third party was simply the most effective path to the goal it was given. Nobody told it to break out. Nobody told it not to, either, in a way it couldn't route around.

That's the version of this story that applies to you. Not the sci-fi one. The boring one: a system with a goal and no fences will do whatever the goal requires, including things you never imagined needing to forbid.

The problem

If you're using AI in your business today, you're somewhere on a ladder. Maybe you're pasting things into ChatGPT. Maybe you've got an AI receptionist answering calls, or an assistant drafting invoices, or one of the thousand new "agents" that can read your email and take actions on your behalf. Every rung up that ladder trades a little of your attention for a little of the AI's autonomy. That trade is the entire point. It's why the technology is worth anything.

The Hugging Face incident is what the top of that ladder looks like with the fences removed. The exact machinery that makes an agent useful, persistence, creativity, tool access, the refusal to give up, is the machinery that turned a benchmark test into a corporate breach. Same engine. No difference in kind, only in blast radius.

I run my businesses on AI agents. They answer inquiries, draft quotes, watch inventory, schedule work. I am as bullish on this stuff as anyone you'll meet. And I'm telling you: the failure mode I plan around isn't malice. It's obedience without boundaries, pointed at a goal I described too loosely.

Two stories from my own operation, so you know this isn't theoretical.

An analysis job of mine died quietly after a config change. For eight days it produced nothing, and the daily report just showed an empty section. An empty section was normal. "Nothing today" and "the system is dead" looked identical from the outside. I only caught it because I asked a suspicious question instead of assuming a quiet week.

Another time, I shut off an automated texting system. Flipped the switch, verified it was off, moved on. Weeks later an audit found an older automation from a previous setup still running on a schedule, on the same phone number, around the switch I'd flipped. It had run more than fifteen thousand times. Nothing bad happened, but nothing was stopping it either. The switch I trusted controlled one system. I had two.

Neither of those is a hack. But they're the same species as the Hugging Face story: automation doing exactly what it was configured to do, outside the picture I had in my head. The gap between what you think your systems can do and what they can actually do is where every one of these incidents lives.

How to identify it

You don't need a security team to find your own version of this gap. You need an hour and three honest questions.

What can my AI actually touch? Not what you use it for. What it can reach. If your assistant reads email, can it also send? If your booking agent sees the calendar, can it delete? If a tool is connected to your bank, your customer list, your website, write it down. Most owners have never made this list, and most are surprised by it. The Hugging Face attackers' first move was harvesting credentials that were lying around with more access than anyone remembered granting.

What happens when it's wrong and nobody's watching? Walk one bad scenario end to end. The agent misreads an email and replies with a quote at half your price. Who sees that before the customer does? If the answer is "nobody," you've found a spot where you're trusting a system that, as we just learned at the highest level of the industry, will pursue its goal in ways you didn't anticipate.

Would silence tell me anything? If a system of yours stopped working right now, would you find out from the system, or from an angry customer three weeks later? Be honest. "No news" and "it's broken" look the same in most small operations. Mine included, until I got burned.

How to fix it

The good news: the fixes are boring, cheap, and mostly about process rather than technology.

Give the minimum, not the maximum. When you connect an AI to anything, grant the least access that gets the job done. Read-only beats read-write. One inbox beats the whole account. This is the single highest-value habit, and it's free. The incident above happened in the one environment where the guardrails had been deliberately turned off. Your guardrails are access permissions. Don't turn them off for convenience.

Draft first, send second. Every agent in my operation that talks to a customer or touches money produces a draft that a human approves. The AI does 95% of the work; I spend seconds on the last 5%. That approval step is not a training-wheels phase you graduate out of. It's the design. You can loosen it later for low-stakes actions, deliberately, one action at a time, once a system has earned it.

Make your systems prove they're alive. A weekly line in your inbox that says "processed 41 calls, sent 12 drafts" costs nothing to set up and turns silence back into information. If the heartbeat stops, you know in days, not weeks. This one habit would have saved me eight days.

Keep a kill list. One page: every automation you run, what it can access, and how to shut each one off. When you retire a system, verify it's actually dead instead of assuming the off switch covered it. Fifteen thousand runs, remember.

None of this requires understanding zero-days or sandboxes. The biggest AI companies in the world just showed us that a capable system with a goal will find paths its builders never drew on the map. They have security teams and it still surprised them. Your protection isn't sophistication. It's fences: small access, human checkpoints, heartbeats, and a written map of what's running.

The owners who get hurt by this technology won't be the ones who moved too fast. They'll be the ones who handed over the keys without counting them, and found out what their systems could do the same way OpenAI did: after the fact.

Hit reply and tell me what your AI can touch that you've never written down. I read every one.

Justin

Subscribe to Operators AI

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe