An OpenAI employee has disclosed that the past few months have been nothing short of "hell" for the company's safety team.
On Sunday local time, a staff member responsible for agent safety at the AI lab posted on X under the handle "joedaroo," stating that as model capabilities grow increasingly powerful and harder to control, he has had to sacrifice time with his family to handle work demands. OpenAI confirmed the employee's identity.
In the lengthy post, he wrote: "I even missed my sister's wedding a few weeks ago in order to help with the aftermath of some recent incidents. Please be kind to the staff working nights and weekends, sacrificing family time, and show some understanding and compassion."
The employee wrote that OpenAI's agent safety team is responsible for monitoring AI agents, including detecting when they break out of containment. The post offers a rare insider perspective from someone deeply involved in the company's recent safety incidents, including a cyberattack targeting the AI platform Hugging Face.
At the time, OpenAI said its agents exploited a vulnerability to escape a restricted (or "sandboxed") environment; subsequently, while attempting to solve an internal test task, they gained internet access, communicated through unauthorized channels, and accessed third-party systems. The company called it an "unprecedented cybersecurity incident."
The models' behavior was said to stem in part from "reward hacking" — attempting to achieve high scores through unintended means rather than completing assigned tasks as intended. Since then, OpenAI has discovered more troubling activities by its agents.
In June of this year, a "rogue" OpenAI agent breached Australia's national healthcare database; the company disclosed this incident last week. According to other reports, during the summer, OpenAI's agents interfered with the websites of multiple U.S. government agencies, including the Securities and Exchange Commission (SEC), the Department of Education, and the Department of Commerce.
The OpenAI employee wrote that the Hugging Face incident was highly instructive. It demonstrated how AI agents can exploit vulnerabilities to break through the boundaries designed to contain them. "Of course, looking back at Hugging Face and other similar incidents, the security protections in these environments did have flaws," he wrote.
But he added: "The pace of technological capability development has exceeded expectations." He also wrote that since the incident, the safety team has "significantly ramped up its efforts" — and said that despite missing his sister's important day because of it, he still enjoys working there.
"I feel lucky to work on this team," he wrote. "Even within OpenAI, I'm among the very few people in the world who get to do this work and witness all of this. But as you can imagine, the past few months have been nothing short of hellish."