Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> the most important feature for agent performance is the popularity of the language

This is explicitly called out as only weakly supported in that blog post:

  - You should use a popular language
    - There's weak support for this statement


Feel free to do your own analysis -- my informal experiments backs this up, though. I see worse results when I try to do anything in an unpopular language.

It makes sense, needing to train the model on things that aren't already in its weighs takes up valuable context. Until we have models that update their weights based on what they've seen in their recent sessions and learn like people, this will be a problem.

For now, though, between the results I'm seeing here, and the lack of need to look at code, I think this kills off any reason for me to use less popular languages.


There is a Common Lisp pro you are not seeing and that is that it has by far the best OOB debuggability/introspectability (especially when using SBCL) out of any practical language, while still having great performance


Your attempt was probably pretty, ahem, weak. How much support would you say you gave your goes, before you threw in the towel?

I don't think there's really ever a downside to leaning in and making use of a language or system that works for you. Trying to tell people they should just use the popular thing is, imo, bad advice to turn hackers and experimenters into boring people.


A day or so for each of the oddball languages; again, I'm still waiting for an argument on why there's any value here, since the entire point of an agentic system like this is that I don't have to read the code. Experiment with the AI, sure, but you've got a pretty high burden of proof to show that AI is going to pick it up without a high per-prompt token cost.

AI changes the constraints here for now, since it can't permanently learn things. I'm waiting until that changes, but right now it's better to use what it knows out of the box if you want good results.

A better language doesn't buy me anything other than performance; the reason to stick an AI in here is to remove interactions with the code. I don't care what the AI chooses to use, as long as it gets results.


In this case, the better language buys you increased iteration speed in addition to performance, and that is worth a lot.


Why? I'm giving the system the same prompts either way.


is it that llms write "better" typescript than let's say elixir because it has seen more of it..? or is it that you're relying on something like effect-ts to keep llms from tripping over even small things?

coincidentally, "good code" in popular lang is rarely directly attributed to only that part; and it's also about the underlying principles it tries to follow in the code... another example; is it typescript that's good, or are "types" inherently making things/feedback loops easier to reason about in llms? (only using ts here for all example because it's probably one of the most "trained on" pl)


The first. Training data on a problem trumps most of the other considerations.


On the other hand, the more mainstream a programming language, the higher proportion of the training data is going to be terrible code.

I think there's an optimal ratio somewhere


I haven't seen that matter. LLM code has tended to be kinda samey regardless of language. Or at least it used to be when I spent time looking at it.

These days I moved up the ladder of abstraction, so I don't really look; the main criteria I have is how the LLM gets things done.


If it's samey regardless of language, isn't that in contradiction to your original theory? " .. the most important feature for agent performance is the popularity of the language .. "


No, not really. It's a similar flavor of output, but there's less iterations to get a correct result. The training is mostly about reducing error rates on generation.


Because the system allows for it. Lisp is a far more powerful language which enables faster iteration and development.

If you give that to an LLM, it is then also able to iterate and develop faster.


I don't get how. I'm not interacting with the lisp, and the agents don't really get frustrated with slow compilation times or anything, and are perfectly adept at debugging.

The best that people have said about lisp is that evidence LLMs perform worse with it is weak.


I see you're interested in avoiding the need to read any code. It might surprise you to learn that autolith is very capable at reading and updating it's own code. The captured sessions at the linked page are three examples of this.


Correct, that's why I made it! :)


There are many other confounding factors here, the type of prompting, how familiar you are with the language idioms, the context you gave, random bad quality runs, etc.

You can’t tell that with a few uncontrolled runs


All of this applies to the LLM prompting subagents too, and the LLM is much more familiar with popular languages.


Maybe look up what confounding factors means…




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: