Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How does this jive with the exponential growth claims? Theoretically sol models are better than the 4 series models I was using at the beginning of the year, but in practice the results don’t seem to be much better. They always nerf the models over the course of the release so it _looks_ like the next version is better but I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
 help



They never nerfed any model after release.

The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.

I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.


How large of codebases are you working on? The models have gotten good enough to 1 shot stupid "trivial" throwaway integration projects with 0 handholding (was having RL'd garbage in late 2025), and I'm actually enjoying designing bounded greenfield personal software from scratch with Astra, in my experience. It's quite slow - 2 weeks of credits and constant talking and back and forth with Astra, but it doesn't feel annoying to talk to and is like an intelligent colleague maybe 70% of the time? Which is great. Just push back when it's dumb.

I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".

I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.

> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.

Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).


It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: