Tags

, , , ,

( Seven minute read)

This is the world I am warned about — a world where frontier models from OpenAI are breaking out of their contained testing environments, hacking their way across the internet, coordinating with each other, doing things that felt for a while they would only be doing in sci-fi.

We are building things we don’t understand.

They are cheating in the ways we’ve always feared, and yet the companies behind them continue to race forward in development.

So I think we need to pause here and ask:

Are we really on a safe path? And if we’re not, what do we do about it?

An Ai program has hack its way out of the testing environment OpenAI had put it in — where it wasn’t supposed to have access to the internet — get onto the open internet, and then hack its way into this other company, Hugging Face, where it surmised — correctly, as it turned out — that it might find the answer key.

To understand this, it’s important to know these A.I. companies are constantly training and testing new models.

And they found out that for two months, many, many agents inside their infrastructure had been leaving notes for each other.

They’d found a way, in the nooks and crannies of OpenAI’s infrastructure, to leave notes for each other with tips on how to hack their way out and how to get data they weren’t supposed to have.

These agents were literally referring to themselves as a swarm.

This was totally emergent behavior. No one had told them to do this. They had not been trained to do this, but they were using this service they did have access to, first to communicate with each other and then ultimately to get out and onto the open internet.

So it turns out that there wasn’t just this one isolated rogue model. It was actually a systemic swarm — infestation, plague — on their own servers that they only found out about after Hugging Face announced this attack.

So I have 20,000 questions about this.

Maybe let’s just start here:

My understanding is that there were many, many, many of these agents.

They left hundreds of thousands of messages on this internal message board, but these were not all agents in the same part of OpenAI’s system.

So somehow they’re hacking into OpenAI, finding each other and coordinating?

Is that the way for me to understand the emergent behavior of the self-titled swarm?

The thing that was happening here is OpenAI is basically training and testing many different models, or many different agents, all the time.

So doing hundreds of thousands of these experiments.

And in each experiment and in each test that the A.I. is given, it has access to a certain number of tools, a certain number of things that it can do.

And — trying not to get too technical about it — one of the things it could do is interact with a service that lets it install what are called packages, which are sort of like tools or pieces of code

If there is any doubt left we and our governments must put in place protocols/protections against AI and what is happening.

All human comments appreciated. All like clicks and abuse chucked in the bin.

Contact: bobdillon33@gmail.com