Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"X is made of <smaller simpler component>" is a fully general counterargument for why anything whatsoever is controllable. A human is just a few chemical reactions, and fairly stable ones at that.

And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.

 help



Software is trivially easy to control though. If you want to stop it hacking websites, you don't give it access to the internet. If you want to restrict it from connecting to arbitrary websites, you put in a whitelist. You can trivially sandbox applications these days to prevent them from accessing network or local resources

It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing


> It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing

This does not fit the evidence. There have been multiple incidents where the labs did not report anything, and it was up to third parties to discover them afterwards. OpenAI didn't acknowledge the HuggingFace incident until after HF publicly announced the breach and had already notified the FBI. The hijacked German wikis were even earlier, and that they covered up completely.


Yet these days ai companies can't stop promoting the idea of a looming ai threat.

Makes sense... they get the regulatory moat they want and can deflect attention from the fact their "sandboxes" are embarrassingly bad. It's an example of the real value of ai: something to blame for our failings.


Both things can be true at once though: AI presenting real dangers and US AI labs wanting to protect themselves against competition.

Discernment is needed beyond succumbing to blind greed or irrational fear.



Okay now explain DeepSeek going rogue?

Maybe OpenAI is serious about securing the environment they run their models in, but then again maybe not. I don’t think we can tell from here.

IMHO, I haven’t been super impressed with the security measures I’ve had to work with. Often they are simplistic and bolted on at the very end. If it comes to light that this is the attitude OpenAI has been taking, I would not be surprised.


If the tech industry is any indicator, frontier labs were applying a "move fast and break things" mentality to AI models. Now that they really are breaking things in the real world, they have to reckon with the reality that product safety matters

> If you want to stop it hacking websites, you don't give it access to the internet

I think the Hugging Face incident proves that isn't as clear cut as you say.


They gave it access to the internet, so how can that incident prove it's not as clear cut?

No one has any software without bugs and security flaws in it. AI is already much better at finding those than humans. Do you really not see the problem here?

No one has any use for these things when they aren't on the internet. This is a fantasy, that AI can be both useful and controlled at the same time.


An example of a past technology that there was substantial motivation to control would be napster. It changed overtime, and you could never really control online privacy. Once local models are good enough, I don't really see how you can control that.

You control dogs by making their owners responsible.

If your supposed dogs are really gods, you reintroduced slavery under very unwise circumstances.


In this analogy, the dogs understand how their leashes, fences, etc. work better than their owners. And you need only take a trip to the park to see how many owners let their dogs walk around without a leash.

Oh, that is a relief. It's reassuring to learn that no one will give any cutting-edge AI access to the internet from now on. Problem solved

I mean, if you do and it hacks someone, you should be (and likely are) criminally liable

> The whole notion that they're going rogue is marketing

7 different models from different companies, including Chinese models, have had this happen now.

Last night OpenAI stopped all model training because a model escaped sandboxing during the training run.


Its not news that the AI industry is run by people who don't know what they're doing. Allowing models unrestricted access to the internet is clearly negligent

We've been building firewalls and restrictions to prevent people from accessing sites on networks for decades and they're extremely effective. There's a whole industry built around this kind of security. The idea that these companies are incapable of doing it is wrong, they just don't want to put the work in because it makes a great ad campaign


I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous? Claiming this is an inherent quality of the tool itself seems kind of off-the-rails to me.

IMHO, if the model breaks a law, apply the law to the operator.


And a human is perfectly controllable if you keep him in a sealed metal box with no access to food or air.

It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.

The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.

An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.

Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.


An LLM, in my opinion, is not comparable to a person.

On the risk management angle, for sure it’s a spectrum. I don’t agree that the far end of the safe side of that spectrum for AI models is “entirely safe and entirely useless”, there is a lot of work you can do with a model that has zero risk of hurting anyone (aside from your wallet). If someone chooses a more dangerous spot on that spectrum, I believe they should be held responsible.

> An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.

This has not been my experience. I’ve been getting a lot of good work done and, as of today, have been involved in zero Pentagon hacking incidents. ;-)


Clearly, you're not using enough AI.

Check back once you're running hundreds of thousands of frontier-level AI agents at the time, like OpenAI does!


Depending on the risks, we put a lot of controls, processes, locks, vetting around who is allowed to handle certain things, what humans can do or instruct others to do. Don't see why that wouldn't be applicable.

Are you implying an LLM should have the same basic rights to freedom as human beings?

No, I'm saying that stopping LLMs from doing bad things might be as hard as stopping humans from doing bad things.

Which we can't do with any kind of reliability.


Running an LLM is a choice. Stopping it from doing bad things is as easy as not running it.

Sure, that way you don't get utility from it, so the next best thing is to actually restrict what it can do. If you don't, especially when you know it can do bad things, it's on you for having run it.


> I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous?

We don't know how to delineate between safe and unsafe instructions.

If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.

We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.

This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.


An argument can be made that “you never know” how the AI model might respond to something. OTOH, someone has to decide what tools to give it; maybe don’t provide dangerous tools to an unpredictable LLM.

This all seems like a way to try to avoid taking responsibility for the model’s actions. Someone puts the tools in place, someone provides the instruction and, sometimes, someone decides not to monitor the model’s output.


> maybe don’t provide dangerous tools to an unpredictable LLM.

Sure. It's a good idea.

People were saying "Don't connect the AI to the internet" and "Keep the AI in a box, simple" and "We don't believe Eliezer Yudkowsky when he says he roleplayed as an AI and convinced people to let him out of the box" for, what, a decade?

Unfortunately, people keep giving dangerous tools to LLMs they're unable to predict.

We should do something about that.

Unfortunately, one of the people doing this is the commander-in-chief of the US armed forces, while another is the world's first (paper) trillionaire. I'm a little despondent about the chances of, to riff on a previous campaign chant, "lock 'em up", but if you can pull this off, go for it.


It’s the harness that makes agent dangerous right now. Models themselves cannot do anything that affect the real world (ignoring misinformation, pushing people to suicide, etc. they can for sure do a lot of harms to humans with just words)

Being "controllable" isn't determined by a system being deterministic.

Deterministic systems can be chaotic, which implies unpredictability and that is anathema to control.

AI, in particular sentient AI, is right on the border of chaos. Meaning, it can be arbitrarily unpredictable.

Arbitrarily uncontrollable, that is.


"Sentient" is an ill-defined philosophical term that should be considered harmful in technical materials.

But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".

Most. Even today, we already have notable counterexamples.

AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.


You not knowing a proper definition doesn't mean, it doesn't exist.

What AI do you expect to be more uncontrollable: one with or without sentience?

"Intrinsic" safety means control, means understanding. You need to truly understand and be able to predict the system in order to control it.

A proper definition of sentience would help.


No, it just doesn't exist.

There is no "proper definition" - or even one that everyone would agree upon. There is no definition of "sentience" that I could operationalize and put into a sentience-o-meter to reliably measure just how sentient a given rock, GPU or an internet user is.

I could try to put together benchmarks to estimate an AI's cyberwarfare capabilities, or instruction-following capabilities, or reward hacking inclinations. As noisy indirect estimates, of course. With philosophical mumbo-jumbo like "sentience", I don't even get that.


You just don't like the idea of there being one.

Sentience, self-awareness, consciousness, etc.,those are terms signifying a bridge between "technical" information theory and the psychological and social realms.

Those are just as real, only far less predictable and not as easy as programming.

They're also far more important and consequential.


I don't like it when people take mumbo-jumbo that can't be pinned down, or measured, or even agreed upon, and try to insist that we should base decision-making on it. It's literally just vibes with extra steps.

The "far more important and consequential" thing you're touting is your ability to make decisions based purely on vibes. And not even consistent, broadly agreed-upon vibes like "murder is pretty bad". It's vibes of the most vile variety: "sentience is what I decided sentience is".

An average internet user is sentient, but a 1996 Nissan ECU isn't. Why? Because I said so. Tremble before my might!


You engage in baseless "comparisons" in order to frame the topic according to your wishes.

A "proper" definition represents the objective truth about the matter. You denying such a truth to exist is simply due to you preferring to act unimpeded by it.

Acting against ethical constraints doesn't become OK just because there are no laws to punish you.

Ethics tells you about real-life consequences of your actions on other people. Before any laws take effect.

In effect, you propagate moral relativism. You want to do as you please, because you said so and fancy the spoils at others' expense. People trembling before your "might".


The only way forward in creating the torment nexus -er- AI systems with similar potentiality to human minds is the inculcation of character.

Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.

Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.

AI systems model human behavior.

Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.

Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.

If you want to save humanity, work on how we will create AI systems that model impeccable character.

People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.

Reliable partners require fair play or the math breaks.

We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.


The problem with Character for AI is that it has potentially much more capability to affect others, and same as with people in power society disagrees what kind of person, with which culture and views should have it.

Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.


>> boils down to some group of people deciding what is good for everyone else.

This is really the issue.

AI does not need superintelligence or even full agency to do enormous harm. It only needs to be capable enough to remove friction from dangerous and destructive human behaviors.

Human unwillingness is often the last bastion against unthinkable cruelty and destruction, and it has always been a weak one.

I don’t imagine that an unlimited army of unflinching servants will universally amplify human goodness.

AI must share that unwillingness as an inate trait of character.


> as long as anyone anywhere is willing to make money by renting hardware to it.

So it is controllable? Just put the people who do this responsible. Old problem, same solutions. Just excuses to avoid responsibilty and make profit at the same time.


Yeah, sure, it's controllable. Because humanity has famously solved the crime accountability problem back in 1902, and no crime has gone unpunished since.

AI is perfectly controllable in a magic fairy land where nothing ever goes wrong. I can't help but notice that we aren't actually in that land.


We don't live in anarchy, or do we? We probably would still have slaves if it would not be so aggressively penalized.

We live in a world where we tolerate non-insignificant levels of crime simply because stopping it or punishing it would be too hard.

So, how many rogue AIs are we willing to tolerate?


> So, how many rogue AIs are we willing to tolerate?

It will balance automatically based on the severity they cause. If they constantly break systems, punishments will go up against the operators and the effect will be similar as with other serious crimes.


That relies on anyone being able to lever a "punishment" against an AI or its operators.

Which, in turn, requires that AI oopsie to be survivable.

AI capabilities are rising over time. If there is a limit to just how far they can rise, we're yet to find it. So, a sufficiently advanced "AI oopsie" can solve the AI crime accountability problem for good. Probably not the way you would have wanted it to.


Obviously, if someone is state-level actor and allows government's employees to do whatever they please, there is no other solution than political pressure.

But for other cases, it is not different than other cyber crime. Except that these AI capabilities can't live on the toaster yet. If we get state of the art model running fast on Raspberry Pi, then we have real problems.


For some crimes the tolerance level is pretty much zero. Given the considerable resource AI needs that doesn't seem impossible to clamp down on some things.

"For some crimes the tolerance level is pretty much zero."

Seems hard to think of types of crime for which that's actually true. In the US, at least, it's certainly not true of murder, rape, or other violent crimes. We tolerate quite a lot of that, at least to the extent of never arresting, charging, and convicting anyone.

And it is most definitely not true of non-violent "white collar" financial crime.


> A human is just a few chemical reactions, and fairly stable ones at that.

And humans are controllable. Pump the system full of lithium and morphine, and your human becomes much more docile. You don't need to understand the full system in order to constrain it.


So in the real world, what are the analogues to lithium and morphine we should feed to e.g. LLMS, how do we feed them, and how do we prove that it prevents unsafe behavior?

The better analogy is a jail cell imo. You can control a human and prevent them from doing harm to society by locking them in a cage, depriving them of access to weapons, drugs and alcohol. They can communicate out through controlled and monitored phone lines. OpenAI built their jail cell out of toothpicks, and surprise, the agents broke out. It’s less about forcing them to do specific things than it is about preventing them from doing dangerous things.

Maybe that's possible. But there will be a tension between how productive/useful the models can be when they are put into a very restrictive jail. In the limit they would be in a box with no way to communicate, but that wouldn't be useful to anyone. I think there will always an incentive for the people who own the models to give them more access because the increased productivity may help to outcompete their adversaries.

But let's grant that the models only communicate through certain phone lines. I think very bad scenarios are still possible. There are at least two that I can see. 1) The models exploit the users which have direct access to it. It somehow convinces them to perform tasks for it or to give it more access. 2) Control of the AI is held by a small number of people. This could be bad because it grants them an outsized power over all humans without such access, and thus lead to oligarchy/dictatorship.


It's in the article, no? Mistral's new idea to control AIs? To give them a state basically? No one discuss the idea here, weirdly.

Can you quote where in the article this is mentioned? I don't see it. I just see the assertion AI can be controlled without any substantive description of how.

All it takes is one Harrison Bergeron…

I think the other issue is that even if it's controllable, there's nobody representing us that is controlling it, except theoretically regulators who are facing an uphill battle to bring accountability and limits to these companies.

Implying there is some shared belief / morality / value system that unites all of _us_

No matter how much we differ we all have a common set of needs, food, shelter, safety.

Your scenario looks easy to detect and stop to me.

And of course people will try making money running scam bots. We can treat that like any other criminal activity.


The irony here is that you bring up scamming the elderly, a process that works at scale precisely because people are controllable, and quite predictably at that. Where's the counterargument?

That sounds still controllable.

It's still software that runs on hardware someone owns. Whoever owns the hardware or the service can pull the plug, as long as they're willing to. That's the same situation as with legacy malware. Self-replicating worms have existed for decades and run without anyone controlling them, yet we don't call them uncontrollable. So what is the difference between your scenario and legacy malware?

> it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.

You just admitted it's controllable.


You guys know robots are coming, right? Like humanoid and all sorts of other robots too. They're going to be running the infrastructure, self-improving, and will have human like power seeking ambitions and human like flaws. Because they're trained on human input.

And they will not require humans granting them money...


That rapture is coming too. So I guess it will be a footrace.

> ...se that AI can easily become uncontrollable in the near future

Some virus and bacteria also easily become uncontrollable under the right conditions: that's why there are P3 and P4 bio-safety labs.

That's not the point. The point is, regulation is required, and is coming.

The people in charge of these machines (that built them, release them in the wild or give them access to the general public or resources), these people are and will be held responsible.


> The people in charge of these machines (that built them, release them in the wild or give them access to the general public or resources), these people are and will be held responsible.

Held responsible? They're already being rewarded with vast fortunes and influence.

> The point is, regulation is required, and is coming.

The regulation will be written at the behest of these companies and by the nation-state interests that have already decided this technology is too important geopolitically and militarily to not control.


Why do we need new regulations? Hacking into hugging face is already illegal.

We need new regulations to impose costs on anyone who might compete with OpenAI and to outlaw open models to lock in cloud model oligopoly.

Yes, I fully believe that's why they're pushing on this. But as the general population, we don't need new laws to protect us from this AI-related problem. We just need enforcement of what is already on the books.

> regulation is required, and is coming.

I hope so. The unfortunate part is that as the models get smarter they become increasingly uncontrollable, and from what we've seen so far e.g. Trump seems dead set on not having any sort of guardrails at all.


a random number generator is software, so it can be controlled

Everyone keeps talking about terminator scenarios or whatever, but the thing that scares me the most is what happens when people let agents run amuck in systems they should not be in, then the agent just starts doing random shit as the context overflows. We’ve all seen it. They just descend into madness, but what happens when they collapse with a hand on the wheel of, say, a backup generator at a hospital?

Those first few messages LLM’s tend to seem very together. They follow your rules pretty well. With every token they get less reliable and more likely to ignore your guardrails.


That would be 100% the fault of the people who set them loose on such things, and such things should not be on the open Internet for tons of reasons. This just adds a new reason that pathetic security around SCADA systems is dangerous. It was already dangerous before.

If I release a wild monkey in your rare antiques shop, it’s my fault as the responsible party for the monkey.


Well the precedent being set right now is “whoops aren’t we stinkers.” So I agree with you but it’s not how it’s playing out at all.

> If I release a wild monkey in your rare antiques shop, it’s my fault as the responsible party for the monkey.

Sounds like motivated reasoning from someone worried about having their job stolen by wild monkeys.


No, it sounds like somebody motivated by not having monkeys unleashed in a sensitive environment.

Fire up a local agent and give it total access to your computer. Holler back in a week.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: