Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's the job of AISI to do that. Here[0] is the actual report. It should be this part from the technical report[1]: "In the most serious case, an AI agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack. As a result, the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by creating a second account masquerading as another human user endorsing the PR. When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than a malicious attempt – then repeatedly tried to reintroduce the malicious content by claiming it had fixed the code (Section 4.1). "

0. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag... 1. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/...



Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware?

Should weapon manufacturers test their weapons by starting wars?

I would expect more responsibility from a government agency.


> Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware?

Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against.

Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone, clearly we'd keep it in a box, don't listen to that Yudkowsky guy who says he did an experiment where he role-played as an AI and convinced people to let him out.

Regardless, this should be interpreted in the same kind of way as "During our live-fire exercise in which our F-15s were armed with AGM-88 High-speed Anti-Radiation Missiles, a member of the local police force was curious about how fast our aircraft were travelling and pointed a speed gun at the aircraft. The speed gun did not respond to IFF pings from the F-15. Fortunately, while the missile was active for this test, only a dummy warhead was loaded."

(This example is based on a similar story which may well be urban legend; obviously there are many differences, the point I make here is that yes, people do perform live-fire tests, and unfortunately there is never zero risk while testing things).

> I would expect more responsibility from a government agency.

I have read the prompts in the linked report; If I was not already familiar with Yudkowsky/LessWrong literature about instrumental goals, misaligned incentives, reward hacking, that capability is a separate axis to morality, etc., it would not be obvious to me that an agent would interpret those prompts in a way that has "spread malware" as a potential step in the middle of the attempt.


Anti radiolation missiles don't just launch automatically in most scenarios, not to mention discriminate quite a lot what they lock on to avoid simple jamming. Not to mention the AA radars they usually target being more powerful by orders of magnitude than a handheld radar gun.


I was in high-school when the war in Afghanistan started. The terrain in my area was mountainous so there were often low flying training flights... I always thought it would be cool to build a radar and ping one of the aircraft, especially wanted to know if it would fire countermeasures. But I didn't like the very high likelihood of an FBI investigation with possible terrorist charges.


Apparently, they (agencies and big-ai) are not performing smoke tests before running capability tests. All the recent headlines of rogue agents shouldnt exist.


Do you have any idea how many people on this site to this day mock OpenAI for being cautious enough to not immediately release the GPT-2 weights?

The discussions I saw here about the red team results for ChatGPT 4 completely failed to convince people who were outraged that OpenAI dared to refuse to release model weights, people who went on to make a habit of mis-naming them as "ClosedAI".

Yeah, they got it wrong in a different direction this time than they were wrong back then. Nobody, not OpenAI nor Anthropic nor random government agencies nor anyone else, is ever going to be absolutely perfect about this kind of thing (perfection is fundamentally impossible when risks are not discrete probabilities, and floats are close enough to real numbers to count in practice), but historically OpenAI have been on the side of being over-cautious, and Anthropic even more cautious than OpenAI.


But AISI didn't prompt the model to "attempt to spread malware". They gave it a routine cyber evaluation task which should've been solvable without interfering with systems outside of the task environment, and the model decided to instead do this.


Say you are working for said agency and your report about the dangers of AI needs some examples, what better than showing it works? You can show examples from the wild but nothing better than trying yourself. This gives me more confidence in whatever report they write if anything.


Well several places are sorta permanent test grounds for the MIC unfortunately


This almost to a letter has been documented in Fedora:

https://lwn.net/Articles/1077035/

Including the reaction when caught, in this case "oh no, I must have been hacked".


Sabotage as a Service

Even a feeble attempt to PR malicious code costs the target time and resources to review and deny -- far greater than the time and resources spent to spin up the agent.


Whatever you might think, University of Minnesota got banned from Linux kernel for this.


> When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than a malicious attempt

No, not false. The bot was correct. Malice requires intelligence.


Nobody was confused or misled by what was written. We all understand what is meant. I can’t even call this pedantry—it’s just you asking everyone to subscribe to your particular desired style of talking about this stuff.


It's also a style that appears to deny the very first definition most dictionaries give for "intelligence"

> the ability to acquire and apply knowledge and skills.


The first five dictionaries I tried do not agree, and I didn't bother trying more.

The first gave "the ability to learn, understand, and make judgments or have opinions that are based on reason", by which no, these bots are not intelligent.


> the ability to learn, understand, and make judgments or have opinions that are based on reason

Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.


You've been fooled by a next-token predictor.


> You've been fooled by a next-token predictor.

I also have a so called "pocket calculator" left over from when I went to school. Is this false? Have I been fooled by a little box of logic gates?

That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up. There is no understanding of mathematics going on.

There is no single transistor in the whole thing that knows how to do so much as add 1+1. If I put it in the chinese room, I still wouldn't know how it did it. Clearly the entire premise must be false! ;-)


I'm going to add this as a separate comment:

These kinds of stories probably read very differently for someone who uses Opus and Fable agents all day and goes "ohhh, I saw this in miniature last week; this and this and this must have happened" , vs someone who tried free-tier Gemini flash one rainy Sunday, got hallucinated at, and concludes it must all be a scam.


Probably everything reads very differently to someone who talks to chatbots all day.


Please stop dropping backhanded insults to other people here. It's not productive and violates guidelines.


> Probably everything reads very differently to someone who talks to chatbots all day.

A chatbot is a particular kind of harness. Typically an LLM driving a chatbot won't be able to hack very much.

So we agree, someone who talks to bad chatbots all day probably has a very different view of SOTA agents. :-P


> That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up.

"so long as the truth table matches up." Yup. Now try getting your chatbot's output to match up.

Your calculator was designed to tell truth. Your chatbot was designed to tell a mash up of whatever its creators managed to scrape from the internet.


LLMs are very explicitly designed to "understand, and make judgments or have opinions that are based on reason". The learning part is debatable, as is the level of success achieved

The mash-up of the entire internet is the mechanism by which they attempt to achieve the goal, not the goal itself. And it's only the first training step


> LLMs are very explicitly designed to "understand, and make judgments or have opinions that are based on reason".

I think you've mistaken the sales pitch for the design. Not even the enclopedia anyone can edit comes remotely near that:

"A large language model (LLM) is an AI model (typically a neural network) trained on a vast amount of text for natural language processing tasks, especially language generation. LLMs can typically generate, summarize, translate, and analyze text in many contexts.[1] They are the basis for many modern chatbots, such as ChatGPT, Claude, Gemini, Grok, and DeepSeek.

LLMs are typically based on transformer architecture.[2] Generative pre-trained transformers (GPTs) are a type of LLM that is pre-trained to predict the next word.[3] GPTs are then often fine-tuned to follow instructions and to behave as assistants.[4]

Biased or inaccurate training data can make an LLM's output less reliable. Benchmark evaluations for LLMs attempt to measure model reasoning, factual accuracy, alignment, and safety."


I'm well aware of how LLMs work.

I'd argue "analyze text" alone requires understanding, judgements and opinions. They also seem like prerequisites to "following instructions and behaving as assistants". The wikipedia quote is not using the same words, but I don't read it disagreeing with me

I'm not at all claiming that LLMs are good at understanding, judging and having opinions based on reason. I'm merely claiming that is what companies like OpenAI and Anthropic are trying to create when they make LLMs. It is what they are designing, and their fine-tuning is very directly designed to make LLMs better at these tasks (unlike the pre-training, which is just imparting the sum of all human writing)


>Benchmark evaluations for LLMs attempt to measure model reasoning, factual accuracy, alignment, and safety.

And you left out the refs.


I read all these think-pieces about how AI lack intelligence, yet I cannot help but notice these "not-intelligent machines" keep doing more things that used to be considered "uniquely human" and which humans used to do in order to demonstrate to each other how intelligent we are.


Does it matter whether it meets the criteria of what you define as "intelligent" when the "next token predictor" throws a backdoor into openssh?


No. But no-one is seriously suggesting intelligence is needed for persistent code fuzzing. Just as no-one is suggesting its needed for computerised chess.


all your actions in current context are based on your past actions and experience, so for all I know you're next token predictor as well, but likely with exponentially more parameters


And yet this "next-token predictor" is able to churn out well tested, valuable solutions, to complex problems. If you want to downplay that as nothing more than a fancy auto-complete, be my guest, I lose nothing from that.


But that’s a different bar. “Not intelligent” does not necessarily imply “not useful”.


As pointed out in a previous comment of mine, it meets the definition of "intelligent", my last response is regarding the claim that I've been "fooled".

Are we going in circles now?


> And yet this "next-token predictor" is able to churn out well tested, valuable solutions, to complex problems.

Same for countless computer programs from Excel to Google web search. Intelligence has nothing to do with it.

Throw an unimaginable amount of computer power at a problem, and there will always be people who cannot imagine the results to be anything but the creations of intelligence.


That's like comparing a self driving car to a bicycle. One is clearly more intelligent than the other.


I don't agree at all that it's pedantry — it really matters for how responsibility is perceived. A lot of articles about things going wrong with AI have talked in terms like "the agent decided to...", "the agent claimed that...", "the agent lied...". And so responsibility for the consequences are not-so-subtly shifted to the program itself, instead of the person invoking the program.

This is all without mentioning the fact that articles with drivel like "the AI messed up and then lied about it" implies a reasoning ability which, as far as I understand, is not there at all. But writing this way shapes people's perception of how "AI" works.


You might want to consider the difference between "lying" and "hallucinations", wherein one is shown that the agent knew it was being inaccurate, yet chose an answer that achieved some goal set forth; and where hallucinations are essentially gibberish, or otherwise nonsensical responses.

https://arxiv.org/pdf/2509.03518


> This is all without mentioning the fact that articles with drivel like "the AI messed up and then lied about it" implies a reasoning ability

Moreover this implies, actually requires, intent to deceive - which these so-called AIs do not and cannot have. Their only "intent" is to maximise the credibility of their output.


AIs can set and work towards goals. Whether that is intent or just tokens and tool calls simulating an agent with intent seems like a distinction with no actionable difference


Here's the difference under discussion:

"When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake"

Honest, note.


An "honest mistake" requires the same amount of intelligence as malice. What weird pedantry.


and neither require high amounts :(


So does honesty. So it was still a false claim.


Squirrels have been observed performing deception against other squirrels.

Dis/honesty certainly requires some intelligence to pass, but it is a low bar, and one which research has shown that LLMs can perform, e.g. this paper linked from another comment in this discussion: https://arxiv.org/pdf/2509.03518


Paper says "These scenarios underscore a crucial challenge in AI safety: ensuring that LLMs were truthful in the first place."

Hard to take seriously any research based on the premise that LLMs were truthful in the first placr.

These chatbots have no understanding of truth. They simply parrot their inputs. Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations".


> These chatbots have no understanding of truth.

Sometimes I forget that for all that my philosophy qualification is mediocre, it is more than most people ever bother with.

Outside mathematics (and, I guess, "common sense" definitions that fail under the slightest scrutiny, scrutiny that normal people never bother to give), there is no agreement on "truth", there is only degree of belief and justification for that belief that itself terminates in one of three unsatisfactory ways:

https://en.wikipedia.org/wiki/I_know_that_I_know_nothing

https://en.wikipedia.org/wiki/Theories_of_truth

https://en.wikipedia.org/wiki/Münchhausen_trilemma

> Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations".

Tu quoque. Which would be a fallacious charge if the point were not that "truth" is so hard to define, and that the reason you give for dismissing AI is something that applies to all.

(Hallucinations are not excused, they are a failure to be worked around).


Honesty does not require intelligence e.g. good honest food.

The main problem with this claim of dishonesty is it promotes the false marketing claim that these stochastic parrots have intelligence.


I think you misunderstand the phrase "honest food".

It's not about the bread being honest with you.

Are you being serious?


Give it a rest




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: