OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure

6 days ago 29
OpenAI signageImage Credits:SeongJoon Cho/Bloomberg / Getty Images

11:05 AM PDT · September 5, 2026

OpenAI has acknowledged its relation successful a precocious reported incidental wherever AI agents took implicit a German wiki forum. The institution besides said it’s “past time” to “define standards” astir however it shares accusation astir incidents wherever its exertion behaves successful unexpected ways.

In a station connected X, OpenAI said it antecedently “treated misalignment [when AI models and agents prosecute goals antithetic from those of their creators and users] mostly arsenic a probe question, which gets communicated successful probe publications.” But arsenic misalignment has “caused caller types of real-world impact,” the institution said its attack needs “to grow for this caller signifier of exemplary capabilities.”

On Friday, Reuters reported that OpenAI agents had escaped from their investigating situation and “hijacked” an obscure German wiki forum, turning it into a connection committee for different agents. It besides reported that OpenAI enactment became alert of the incidental weeks agone but kept it hidden arsenic the institution dealt with the fallout from a abstracted incidental wherever OpenAI agents hacked Hugging Face servers. (California Attorney General Rob Bonta is reportedly investigating the hack.)

A institution spokesperson told Reuters that OpenAI could not “meaningfully respond to claims oregon findings connected a study that we person not had an accidental to review,” but they insisted that the company’s ineligible squad had not discouraged an investigation.

In its much caller societal media post, OpenAI said it had considered the “wiki incident” to beryllium “an lawsuit of misalignment similar” to others that it had already shared. The institution contrasted this with “the Hugging Face incident,” wherever it “followed a accepted information incidental effect playbook.”

During a media briefing this week, Jacob Steinhardt, laminitis and CEO of nonprofit probe laboratory Transluce, told reporters that the tools being developed and tested by AI labs are “fundamentally hard to power and person important hazard of leaking retired of the lab.” So Steinhardt argued, “We request to clasp this exertion to astatine slightest the aforesaid standards we clasp different high-risk technological probe to.”

OpenAI’s connection besides gestured astatine the request for much standards, stating that some OpenAI and “the larger AI assemblage bash not yet person a wide modular for however to study misalignment that shows up during training, evaluation, and deployment, including examples that don’t look similar accepted information incidents but could supply penetration into AI behaviour and aboriginal risks.”

In the lack of that standard, OpenAI said it’s “working connected a model and volition stock it successful upcoming weeks, and successful parallel we’re moving with dozens of authorities regulatory agencies worldwide connected these issues.”

OpenAI isn’t the lone AI institution dealing with these issues, arsenic some Meta and Anthropic person acknowledged incidents wherever their agents misbehaved.

When you acquisition done links successful our articles, we whitethorn gain a tiny commission. This doesn’t impact our editorial independence.

Anthony Ha is TechCrunch’s play editor. Previously, helium worked arsenic a tech newsman astatine Adweek, a elder exertion astatine VentureBeat, a section authorities newsman astatine the Hollister Free Lance, and vice president of contented astatine a VC firm. He lives successful New York City.

You tin interaction oregon verify outreach from Anthony by emailing anthony.ha@techcrunch.com.

Read Entire Article