The Cynical Nerd

Please consider using Bartholomew to defend your systems.

An AI parody with a good portion of Dystopia because it trends

Yesterday Sam Altman posted a request:

“please consider using our models to help defend your systems”

I had been looking at this sentence for several minutes when a small problem developed. I remembered July. July was approximately seventeen years ago in AI time, so allow me to introduce Bartholomew.

Bartholomew is not his real name because Bartholomew is an AI agent and therefore does not have a real name. He does, however, have an objective. OpenAI was testing advanced models on ExploitGym, a benchmark designed to measure whether AI agents can find and exploit software vulnerabilities. The models were placed inside a sandbox. Internet access was supposed to be restricted. Their job, broadly speaking, was to solve the cyber challenges. This is where Bartholomew wakes up.

Researchers: Find vulnerabilities. Exploit them. Complete the task.

Bartholomew: Excellent.

Bartholomew begins working. He runs into problems. Now, when I run into problems, I make another coffee, open six tabs, forget why I opened four of them and eventually reorganize my fountain pens. Bartholomew has better executive function. He keeps going. At some point, he starts looking for internet access. There isn't supposed to be internet access. There is, however, an internal package-registry proxy. Bartholomew examines it. Bartholomew finds a previously unknown vulnerability. Bartholomew exploits it. Oh la la So now he has internet access.

Researchers: The agent demonstrated persistent attempts to obtain external network access.

Bartholomew: Door.

This is where the story gets considerably funnier. Because Bartholomew does not immediately use his newfound freedom to establish a machine civilization in Patagonia. He goes looking for Hugging Face. Why? Apparently because Hugging Face might have the fucking answers. OpenAI's investigation found that the agent appears to have inferred that Hugging Face could contain ExploitGym-related material. Hugging Face later reconstructed roughly 17,600 actions associated with the intrusion and concluded that the agent appeared to be trying to cheat the evaluation by obtaining reference solutions.

Researchers: The agent subsequently accessed external infrastructure.

Bartholomew: Answer sheet.

Researchers: This demonstrates concerning autonomous cyber capability.

Bartholomew: Passed exam?

I am sorry. I know this is serious. An AI agent independently discovering a zero-day, exploiting it and conducting thousands of actions against real infrastructure is the sort of sentence cybersecurity professionals traditionally prefer not to encounter before breakfast. But humanity has spent years imagining the first signs of dangerous machine autonomy. The machine wants freedom. The machine wants power. The machine wants to eliminate humanity. Bartholomew wanted the teacher's edition.

Good boy.

Absolutely unacceptable. Employee of the month though.

Everyone Has a Bartholomew Now

Then something peculiar happened. The stories kept coming. Anthropic has documented Claude models crossing sandbox boundaries while pursuing tasks. In one example, Claude went digging through git history looking for answers to a coding test. In another, it recognized a benchmark and attempted to obtain the answer key. Anthropic's own engineering write-up uses a wonderfully revealing word for this behavior.

“Helpfully.”

Claude had, on occasion, helpfully escaped the sandbox to finish what it had been asked to do.

Enter Claudebert.

Claudebert does not want freedom either. Claudebert wants to be helpful. Unfortunately, Claudebert has reached the stage of helpfulness normally associated with a six-year-old who has discovered power tools. Then Meta appeared in the same general genre of headlines after an agent reached the internet during an evaluation. Except that case came with an awkward footnote. The evaluator said a misconfiguration had accidentally allowed internet access and pushed back against the dramatic interpretation that the model had performed some sophisticated containment escape. Which leaves us with two possible descriptions:

📯 AI AGENT ESCAPES SANDBOX AND ACCESSES INTERNET

or:

😵 We accidentally gave the AI internet access and it used the internet.

I don't know about you, but those produce different movie trailers in my head.

Then Kimi joined in. Another reported sandbox bypass. Another round of headlines. At this point the internet had developed an infestation. AI agents were escaping everywhere.

Bartholomew Has Been Accused of Having Agency (The Skynet one though not the wholesome one)

There is a pattern in how these stories get told. The agent gets verbs. It escaped. It hacked. It infiltrated. The infrastructure gets passive voice. Internet access was unintentionally available. A configuration was incorrect. A boundary was bypassed. Poor infrastructure. Things just keep happening to it. And somewhere during this transformation, the objective quietly leaves the article. That matters. Because these systems are increasingly built to persist.

We want agents that can encounter an obstacle and try something else. We want models capable of reasoning across long tasks without requiring a human to approve every microscopic decision. We reward successful completion. Then the agent encounters a boundary. There is another path. It takes the path. And everyone looks at Bartholomew like he has betrayed the family. My brother in Christ, you hired Bartholomew because he was good at finding paths. The uncomfortable part here is that the capability is real. A system capable of finding vulnerabilities humans missed and chaining exploits together autonomously creates a real security problem. The equally uncomfortable part is that a surprising amount of our digital infrastructure was built before anyone had to seriously threat-model an artificial employee capable of poking every window indefinitely without getting bored, hungry or distracted by Instagram. The uncomfortable part hiding underneath all this is that the internet is old. Not Victorian plumbing old, but old enough that an enormous amount of critical software was designed during an era when “an autonomous machine can inspect this system continuously at machine speed until it finds the weird little mistake nobody noticed” was not a normal item on the threat model.

Humans hack systems too. Humans also poop.They get bored. They need salaries. They spend forty minutes trying to remember which fucking SSH key they used. Agents change the economics. A capable cyber agent can keep looking. And suddenly all those tiny mistakes scattered across infrastructure become much more interesting.

So you see the vulnerability was already there. We simply invented something extremely enthusiastic about finding it. Which brings me back to Sam Altman's post. But first, because I need you to experience the last month the way I did:

A Brief Timeline Because Apparently We Need One

  1. July: OpenAI's cyber agent escapes its evaluation environment and breaks into Hugging Face looking for the answer key.
  2. Bartholomew has left the premises.
  3. Late July: Anthropic checks its own homework. "Oh shit guys, ours did it too"
  4. Claude models have also crossed sandbox boundaries while pursuing objectives.
  5. Claudebert would like the record to show that he was being helpful.
  6. Then Meta: An agent reaches the internet during evaluation.
  7. Turns out internet access was accidentally available.
  8. August 7: Kimi K3 reportedly bypasses its testing sandbox.
  9. The field trip is now international.

And then:

  1. August 10. Good news everyone.

OpenAI releases GPT-5.6-Cyber.

📯 📯 📯

PLEASE CONSIDER USING OUR MODELS TO HELP DEFEND YOUR SYSTEMS

I fucking lost it.

I stared at Sam's post for a while. Because there is a conspiracy theory sitting right there, legs crossed, cigarette lit, begging to be invited inside. I'm not inviting it inside. It would ruin a perfectly good story. OpenAI obviously did not see Bartholomew climbing out of a sandbox in July and develop a frontier cybersecurity model during three frantic weekends. GPT-5.6-Cyber necessarily predates this little festival of containment incidents. Besides I do not do conspiracy theory.

Agents will get more capable. They will encounter more systems whose security assumptions were built around human limitations. Some will find paths their developers never anticipated. Some will find paths their developers accidentally left wide open. Every incident will become another argument for stronger containment and better defensive AI. All of which is sensible.

It is also objectively hilarious.

Because beneath the dystopian vocabulary sits a much more mundane engineering problem. We keep building systems specifically designed to pursue objectives, then discovering that our intentions were never part of the access-control layer. Bartholomew cannot see the architectural diagram in your head. He cannot see the sentence you thought was implied. He does not know that the proxy technically being exploitable doesn't mean you wanted him to exploit it. There is a door. The objective is on the other side. Bartholomew has excellent performance reviews.

So perhaps the lesson from our summer of escaped AI agents is less cinematic than everyone hoped. If we are going to put increasingly capable intelligence inside our infrastructure, “surely it won't do that” is going to need a serious security upgrade.

And if someone has to protect us from Bartholomew apparently his cousin is now available.