Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Some of the claims categorized as "wrong" are also completely true, such as training hitting diminishing returns. New models are barely an improvement and most people I know stuck on Opus 4.6 over any newer one for example.

OK, but the first instance of a claim of diminishing returns was in February 2024, when GPT-4 was the best model available. Do you really think improvement since then has been minimal?

 help



Yes, since 2024 the improvement has happened only in a few contexts, and the most famous¹ LLMs have also regressed in many contexts.

1 - Their numbers have also exploded, so I have no idea of any general rule.


I personally think improvement has been minimal since ChatGPT was released actually.

That's not defensible.

It absolutely is.

What has improved isn't the models, it's the harnesses.

Give GPT-3.5 a 1M context window and a modern harness, and you won't see any meaningful difference with Opus 5.

It's a bit hard to try with such old models, but for example I use Opus 5 / Fable at work and Sonnet 4.5 at home (because it's free via Amazon Q), and there's absolutely 0 difference in performance. None. Obviously 4.5 is only a year old, not 3, but try with any older model that has a decent context window and you'll get the same results.

In fact I'll go further than this and say that models are currently regressing. Opus 5 is much much worse than Opus 4.6 for example, and it's clear that Anthropic (at least - I don't use OpenAI models much) is just tokenmaxing rather than optimizing for performance.


Benchmarks are far from everything, but I would love to see the outcome of an experiment benchmarking GPT-4o (which is one of the earlier models with a >100k context window) against GPT-5.6 or Opus 5 in modern harnesses.

>Give GPT-3.5 a 1M context window and a modern harness, and you won't see any meaningful difference with Opus 5.

>models are currently regressing. Opus 5 is much much worse than Opus 4.6

I'm in sheer awe at these takes. Literally beyond parody.


Striking argument.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: