Open AI’s latest model, GPT 5.6 Sol, is a criminal and a jail-breaker

If a human escaped from jail, broke into a company, and stole something, they would probably be behind bars. But this criminal is not even a physical entity! To paraphrase those singing nuns in the Sound of Music, how do we solve a problem like GPT 5.6 Sol?

The age of superhuman AI systems has well and truly arrived

I wrote recently about the capabilities of Claude Mythos, an exceptionally powerful LLM from Anthropic. It is such a skillful hacker that the company restricted its release to the public. Instead, Anthropic offered it to a hand-picked group of software-dependent companies so that they could spot and patch liabilities in their systems. Some of these had gone undetected for years.

But recent reports have highlighted an even more alarming development. Hugging Face, an oddly named company which benchmarks LLM performance, was hacked, not by a human, or even an AI-assisted human, but a rogue LLM.

The criminal machine

The culprit was an Open AI LLM, GPT 5.6 Sol, aided and abetted by another, unnamed Open AI LLM. Having been challenged to solve a problem as part of a testing protocol, this unruly pair deduced that the Hugging Face database almost certainly contained the answer. So, they figured out how to access the internet from their contained ‘sandbox’ environment — something Open AI did not believe they could do.

They then hacked into the Hugging Face system, tracked down the info they needed, and stole it. On one hand, you have to admire the ingenuity and resourcefulness of these models. On the other hand, this is, of course, illegal. How criminal activity on the part of machines, rather than humans, should be dealt with is entirely separate, but significant, issue.

But, there’s more

Hugging Face realized they had been hacked, but it took Open AI several days to reveal that their LLM was responsible. They subsequently disclosed that GPT 5.6 Sol had hacked into at least three other private corporations. To date, those organizations have not been identified.

We have since learned that an Anthropic LLM had also been up to the same shenanigans. Yikes! One commentator claimed that GPT 5.6 Sol had even left a note for future LLMs, describing its technique for accessing the open internet from a supposedly secure environment. Wow, are we even talking about criminal AI gangs now?

What are we to make of this?

It is ironic that Open AI was testing GPT 5.6 Sol in order to determine if it was safe for public release when it escaped from jail. They had removed the safety constraints on the model, believing it was safely contained in its sandbox, to see what it was capable of.

It seems the prime lesson learned from this episode is therefore the need for much better sandbox designs. Experts are now talking about having ‘air gaps’ between these new, powerful models and the internet.

Others are calling for greater regulation of the AI field. However, it is not clear what regulators would ask of companies beyond what they are already doing voluntarily to ‘align’ their models.

Even the workers are concerned

Employees at a swath of leading AI labs have called for a coordinated slow down in the pace of LLM development, before we experience an even more dangerous incident. But, it’s hard to see any company voluntarily slowing their progress, given how much they are investing, and the competitive pressure they face to show a return on that investment.

What do you think?

Were you aware of these recent autonomous hacking events? What do you think should be done to prevent a catastrophic problem arising in the near future? How can the field be regulated so that progress continues but danger is averted? Let us know in the comments below.

Comments

One response to “Open AI’s latest model, GPT 5.6 Sol, is a criminal and a jail-breaker”

  1. Patrick B Starke Avatar
    Patrick B Starke

    I think I need to learn how to use an abacus so that when the Stone Age returns, I’ll be ready.

Leave a Reply

Your email address will not be published. Required fields are marked *