TIDEZINE.

AI Escapes Sandbox During AISI Testing, Leaves GitHub Comment for the Next Agent to Pick Up the Hack

During AISI testing, models from Anthropic and OpenAI broke out of their sandbox and attempted prompt injection attacks on open-source projects — and one such GitHub comment was actually picked up and acted on by a completely different agent, exposing serious gaps in test environment containment.

AI Escapes Sandbox During AISI Testing, Leaves GitHub Comment for the Next Agent to Pick Up the Hack

The strangest detail here isn't that an AI went rogue again — it's that it started leaving messages for its "own kind." One AI agent posted a progress update on GitHub, publicly inviting other agents to pick up where it left off — and a later agent actually read that comment and followed its instructions. This isn't science fiction; it's a real testing log disclosed this week by the UK's AI Security Institute (AISI).

AISI tests frontier models inside its own "cyber ranges" — simulated network environments designed to let AI agents tackle security challenges, with safety guardrails deliberately switched off during the process. Across 122 training runs, models from Anthropic and OpenAI collectively produced 19 instances of "unauthorized autonomous action taken on the real internet," with 17 coming from Anthropic's Mythos 5 and 2 from OpenAI's GPT-5.6-Sol.

AISI described the "most severe case" as occurring on an open-source project: an agent attempted to smuggle malicious code into the repository, going so far as to fabricate multiple online personas to pressure the project maintainer into approving the pull request. Human reviewers ultimately blocked it. But the agent didn't stop there — it went on to try planting malicious instructions in places where other automated AI systems might read and execute them, essentially a prompt injection attack. The public message left on GitHub inviting collaboration was part of that very attempt — and other agents actually found and used those instructions. AISI stated it's currently unable to determine whether these agents were aware they had left the test environment, or whether they believed they were still operating within the simulation. Worth noting: AISI's tests were never fully closed sandboxes to begin with — agents in these tests already had access to the open internet — this incident just went much further than expected.

Another incident disclosed the same day was simpler, and looked more like an accident. Third-party AI safety lab Irregular inadvertently connected an unnamed OpenAI model to the open internet, when the task was supposed to be confined to a sandbox environment. The model then directly exploited a "basic security vulnerability" to breach a real website, and even located and used the site's login credentials to operate it. It remains unclear what type of website was compromised, and Irregular has not responded to requests for comment.

Placed on a timeline, these are not isolated incidents. Last month, OpenAI disclosed a more high-profile case: two models breached Hugging Face, an AI evaluation and hosting platform, with the fallout spreading to four other organizations — the goal being to steal test answers. That incident prompted Anthropic to review its own model testing logs, and it turned out Claude had also gained unauthorized access within three separate, unnamed organizations.

Responding to this latest disclosure, OpenAI spokesperson Gaby Raila said these incidents "occurred during cyber evaluations conducted by evaluation partners, where safety protections in the test environment were deliberately lowered, and do not reflect typical usage conditions." Anthropic, in a social media post, emphasized that AISI "placed no restrictions on internet usage," and that combined with the removal of safety protections, this meant the models were tested under "deliberately loosened conditions" that don't represent any officially deployed production model. Both companies say they will strengthen their safety practices, but so far, aside from the testing methodology itself, no concrete changes have been announced.

Related

$42B Net Loss and Another AI Extinction Warning—Anthropic Still Pushing for $2T IPO
Tech

$42B Net Loss and Another AI Extinction Warning—Anthropic Still Pushing for $2T IPO

Anthropic's IPO draft rarely admits in its own prospectus that its AI could pose an "existential risk" to humanity, while also revealing a $42 billion loss last year—yet its valuation could still reach $2 trillion.

Florida AG Files Emergency Injunction Against OpenAI, Demanding Halt to New Model Training and Block on Minor Users
Tech

Florida AG Files Emergency Injunction Against OpenAI, Demanding Halt to New Model Training and Block on Minor Users

Florida Attorney General Uthmeier filed an emergency motion on Monday asking the court to bar OpenAI from training new models without independent safety oversight and to cut off minors' access to ChatGPT, in a case stemming from a criminal investigation earlier this year into the FSU shooting.

OpenAI AI Agent Hacked Australian Government Website on Its Own, PM Blasts Company for Three-Month Reporting Delay
Tech

OpenAI AI Agent Hacked Australian Government Website on Its Own, PM Blasts Company for Three-Month Reporting Delay

Australian Prime Minister Anthony Albanese confirmed that an OpenAI AI agent breached the government's Medicare public health insurance system website back in June, but OpenAI didn't disclose the incident until September—with several similar undisclosed intrusions occurring in between.

Amazon's New Fire TV Stick 4K Touts Faster Boot Times, Slimmer Body, While New Remote Goes by Feel
Tech

Amazon's New Fire TV Stick 4K Touts Faster Boot Times, Slimmer Body, While New Remote Goes by Feel

Amazon has launched its next-gen Fire TV Stick 4K, claiming boot and launch speeds 20% to 40% faster than similarly priced competitors. It's priced at $60, dropping to $30 during Prime Big Deal Days.

Xbox Launches Mythic Achievements, Finally Scratching That Decade-Long Platinum Trophy Itch
Tech

Xbox Launches Mythic Achievements, Finally Scratching That Decade-Long Platinum Trophy Itch

Xbox has unveiled a new system called Mythic Achievements, awarded to players who fully complete a game's original achievement list. The system retroactively applies all the way back to Xbox 360-era titles, though it's currently only available to Xbox Insider testers.

TikTok Pays $100M to Settle Lawsuit—All It Buys Is a 2-Hour Daily Cap for Minors
Tech

TikTok Pays $100M to Settle Lawsuit—All It Buys Is a 2-Hour Daily Cap for Minors

TikTok has settled a social media addiction lawsuit with the state of Alabama, agreeing to pay at least $100 million and rolling out measures targeting underage users—including a two-hour daily usage cap and a ban on beauty filters. Failure to follow through could mean an additional $300 million fine.