Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

When exactly did we forget how to make literally anything that can perform a computation but (physically, hardware-level) not have the ability connect to the Internet?

With these companies spending the kind of money they are, if they actually mean what they say about the security risks, they should be expected to figure out those kinds of precautions and take them.

And build Faraday cages too, just in case of a hardware supply chain compromise.

 help



That's why all of this marketing about agents going rogue is so unbelievable. The only way for a tool to escape a sandbox is if you built a crappy sandbox, and after this length of time I literally don't believe that they can't do it

This kind of sandboxing is not complex to do, especially for a company with OpenAI money. If you want your tools to explore hacking, you restrict them from internet access except for a whitelist of sites that have either opted-in, or you've very carefully vetted to make sure you won't cause any problems to. Its also not difficult to restrict their ability to make calls to be simulated, or to use fake tools that can only run the real commands if they're being run against the correct target

This is all incredibly basic security stuff to make sure you don't accidentally cause someone problems, and I simply don't believe these AI companies anymore. Its either intentional, or gross negligence


Two things can be true. Yes this appears to be negligence on the part of OpenAI.

However, making a secure 'sandbox' is quite hard. There's a huge variety of exploits that exist today, including many we don't know about. Strong models have already shown a capability of finding and using such bugs.

Even one of the strongest boxes we can imagine, literally just a text interface a human can read, has been repeatedly shown to allow unfriendly AI to escape containment: https://www.lesswrong.com/w/ai-boxing-containment


I find the super-persuasion argument annoying. Yes, an AI could socially engineer a human to give it what it wants. However, the vast majority of AI conversations in OpenAI's training runs were unmonitored. The agents in the Hugging Face hack correctly understood that human supervision was nonexistent and acted accordingly. I suspect the agents in this breach did the same.

All the AI-boxing arguments from the LessWrong people assume that there will be one AI, with one chat context, in conversation with a human. The actual deployment environment is more like many contexts talking to themselves, so arguments about human supervision are limited. Technical constraints and IT security practice are what matters in the actual deployment environment.


It's what we know mattered for this model. It doesn't tell us about future models; we already know this strategy is unsound, so it is silly to rely on it.

This also happens in prisons where inmates are able to manipulate guards or nurses.

e.g.

- guard is chewing gum

- inmate says "where's my piece of gum?"

- this is b/c it's against the rules to chew gum and the inmate is implicitly stating this

- the guard should go to his boss and admit he made a mistake

- but decides to give the prisoner a piece of gum instead

- the inmate now has leverage over the guard b/c the guard broke the rules and then "covered it up" by giving the inmate gum

- this can then get slowly escalated into bigger and bigger asks from the inmate until the guard is bringing in drugs

This is not hypothetical. There are documented cases of both this and male prisoners "seducing" multiple female prison employees despite their being explicit warnings about this.

Both of the above are from this book: Maximum Insecurity: A Doctor in the Supermax by William Wright [0]

I highly recommend it for both the stories above but also a look inside prisons in general and how they operate with medical care more specifically.

0 - https://amzn.to/3TVhGBV


Hadn't seen that link. Thanks!

Protecting against an AI convincing the human to let it out of the box is an interesting challenge. Even a dual-keyed system to do so only requires convincing two people, although I suppose more elaborate containment mechanisms can be devised. N-keys is democracy and that doesn't seem immune.

Two random thoughts:

* This box we want to keep AIs in reminds me a bit of the story of Pandora's box

* I'm also reminded of the original alignment problem in Genesis 3 where one creature convinces another to take an unaligned action.

Containment is a pretty fundamental tricky security problem actually once persuasion is part of the threat model.


It's been shown that people could be convinced to say "okay I'll let you out of the box". That doesn't mean that the person thus convinced is actually capable of doing so.

Of course, there are huge risks there. But this goes more towards explaining the fact that OpenAI's experiments thus far have worked the way they did, than it does towards actually informing a useful threat model for OpenAI to follow.


>has been repeatedly shown to allow unfriendly AI to escape containment:

Has it been shown, or has the cult simply updated their tenets to preclude AI containment?

I swear any idiot with the tiniest bit of security or networking experience could box up an LLM. Failing to do it properly is a choice.


It’s beyond absurdity at this point, but the marketing plan of “describe our poor sandbox design as powerful agent capabilities” clearly works.

> The only way for a tool to escape a sandbox is if you built a crappy sandbox

Well, sure, but typical software-level sandboxes are crappy at an alarmingly high rate, either on this access or the usability access. Languages like Python are fundamentally not designed for sandboxed interpretation; any Bash tool is at least as insecure as all of the vulnerabilities in all whitelisted executables.

I'm arguing for hardware-level measures on basic defense-in-depth principles. Like, such a huge part of the reason why we're even doing this AI research is to find vulnerabilities, so it's insane to have a test environment that doesn't start from the premise that there are vulnerabilities. In everything.


Well:

1. Park your Python or whatever code inside a VM with no connectivity to anything (other than to accept inbound ssh) and with a canned set of PyPi etc packages available for it to use

2. Park hypervisor for that in a machine with no connectivity except thru a firewall that only admits the relevant ssh traffic.

3. Make sure to use ssh clients that can’t be exploited by a remote server. If this is too challenging, then use telnet or rlogin instead.

4. You can extend this to eg allow outbound calls to an LLM. Alternatively, place the GPU hardware and LLM weights inside the hypervisor machine.

Congratulations, now nothing can escape except what you allow to from the output from rsh.


Based on our existing understanding of the software, yes.

You want to train these models with access to the internet so that they will learn to use the internet.

So you build an offline tool that simulates it, or you proxy through your own service where you can ratelimit, inspect, and restrict the traffic

None of this is difficult to do, and its impossible to believe that a company the scale of OpenAI doesn't know this. I've built web crawlers and scrapers before, and the thing you do is test them extensively offline against simulated versions of the sites in question, and then very VERY cautiously run them against the prod versions so that you don't cause anyone any issues

The only reason not to do this is because OpenAI doesn't give a rats ass about the internet as a public good, nor the legal consequences of compromising systems


> or you proxy through your own service where you can ratelimit, inspect, and restrict the traffic

It literally seems like they are doing just that, and the agents are just finding holes in that.


If you own a network and the servers, you can DPI every single packet and see literally every bit of information. All of the text to the "forum" that they created must have been in -outbound- packets to their compromised package manager, by definition. If they can't properly analyze network traffic, they should not be running 'sandboxes'.

Anything beyond baseline would be observable- silence, malformed packets, too much egress, unusually large packets, etc


OpenAI should really do better. If you want to build a mostly airgapped system, you find the surface that bridges the inside part to the outside part, and you enumerate every single thing that can get through. Which presumably should be a very very small list and should not include DNS.

If you are using a firewall, you are already doing it wrong. Don’t list thinks to block - list things to allow and make that list small.


Yep and this is literally the most basic security 101. It isn't hard. They've had literally years and years to iron out the kinks in this setup as well. The only reason not to do it is pure negligence

We cannot compute anything without internet access and a Facebook account.

Networking companies (like Cisco, HPE etc.) do this all the time with their test beds deliberately disconnected from internet. It's not hard.

They’re testing rather different capabilities, to be fair. This is a bit like saying people test combustion engines without connecting them to WiFi.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: