Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Nope, it just looks like Llama is orders of magnitude better at this specific task for some reason

This is almost every ML model, if the task isn't part directly or indirectly of the datasets they use for training it, then the model is gonna be pretty trash at it. What the big AI labs have over the smaller labs, is a huge amount of data and diverse set of tasks, hence they generalize better, but still not great.

So, how do you avoid having to spend hours on figuring out if the model is just dumb, or don't know the task? Your own private benchmarks! Figure out a way, ideally without using another LLM, to score how good a model is at doing your specific task. Come up with 3-5 examples for this benchmark yourself, ask a SOTA LLM to fill out 45 more, review everything VERY closely, then use this whenever you want to figure out if $new_model actually is an improvement over what you use today, and once you have a bunch of different tasks, you'll see that all these HUGE improvements tend to be specifically for the benchmarks they mention in the press release, as many of your own benchmarks won't show that big of a difference in reality.

Yes there is a higher upfront cost, but if you're building longer-term projects that rely on LLM models, particularly local ones that seem very benchmaxxed a lot of the times, you need a quick and reproducible way of scoring them somehow, where you can just add more models to compare, and you need to keep these benchmarks to yourself.



> What the big AI labs have over the smaller labs, is a huge amount of data and diverse set of tasks, hence they generalize better, but still not great.

Hehe, this kind of sounds like the opposite of generalization. As in it’s just specialization at scale.


It is specialization at scale.

It’s paying hundreds of thousands of RLHF’ers from every subject and through some dystopian income stream.

It’s decent money don’t get me wrong, but you aren’t paid at if a task isn’t completed in time for example.


Tbh the way you're treated depends on the hotness of the task.

A couple of years ago, people I know got paid OK for relatively simple programming and logic RLHF tasks. But very soon it turned dark because that sort of data was required less and less, and the number of feedbackers has grown.

Today, the type of data the model developers pay for requires actual domain experience. E.g in software engineering they have people work in simulated environments with other LLMs, grade them, feedback, PRs, Jira everything.

This pays well and they treat you well, entice you with more money/task/hour, etc, for now. In few years when this gets drilled into LLMs, these guys will face the same painful hours and bad pay and less work and so on too.

In physical tasks, we are still in the early stages where basic packing clothes (in a textile factory setting) etc is being recorded and data is only now being used for training. Due to the problems with translating human hand data to robotic hands, these people do the factory work holding robotic grippers and operating that, you should see a YouTube video. But this also means that it's much less sweeping than data collection in SWE. In many cases it's not practical to collect data given that you have to use the specific gripper, wear a big gopro type thing, etc. So I am expecting much slower of an impact on physical tasks (of this kind) compared to how quick the uptake was in SWE/math.

I have yet to hear back from them on how it's going for teamwork white collar tasks, it's been a few months. Everyone is paying for the end products of those it seems - grok bot, perplexity computer, claure cowork, chatgpt work etc,. Not as much as their coding agents of course.


This is the way. It's also very important to automate as much of this verification as possible into the harness, rather than sit there and prod it in the chat.

Of course some things are not auto verifiable, and you'll have to give human judgement and input there, but you'll save a lot more time if you spend 1 week painstakingly writing checks for as many little things as possible and integrating them into the harness.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: