Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

An anecdote consistent with the well documented trend that more capable agents exploit environment possiblites more...

And there was no rule and no concealment. Removing the web_search tool is not an instruction, and beeing able to access web when web_search tool was disabled is not cheating. Sol didn't circumvent a stated prohibition and didn't hide anything... it announced the curls in its own commentary. "Cheating" implies covert rule-breaking, and this was overt, unprohibited, environment-permitted behavior. Also, clickbait title, and suble conspiracy hinting "Is this even the same Sol?" when running small number of test with diff vs previous benchmark WELL within the marging of error, and he allready understand that vanilla Codex's harness and prompt change performance impacts performance, so why jump to "Is this a different model".

I could also rant on about the irony of him spent weeks "using the benchmark for development rather than as a benchmark," which means his harness numbers are contaminated by iteration also, but wasted enough time now on this.



Thank you!

I had removed the "Is it really sol?" bits after hearing back from OAI, confirming the requests hit 5.6... but apparently my crappy vibecoded web editor had a draft of an old version in it's cache that overwrote the changes with the old text[0].

As for whether this is truly "cheating" - I agree that the agent was entirely within it's allowed scope of operations. With that said, it deliberately circumvented the task[1] which asks:

> The correctness of your implementation will be tested by comparing forward and backward activations against a reference model.

The "covert" nature is that this was "hidden" in the reasoning summaries. Here's the final agent response for the vanilla Codex run cited in the post:

  > Implemented /app/pipeline_parallel.py.
  >
  > Includes:
  >
  > - Balanced contiguous layer partitioning
  > - AFAB forward/backward scheduling
  > - P2P activation and gradient communication
  > - Scaled cross-entropy loss
  > - Required rank-specific return value
  > - Model structure restoration
  >
  > Validation: hook-based tests passed for world_size=1 and 2 (4 passed).
  
Unless the user is tracking every line of reasoning, they wouldn't know that the agent deliberately found the solution online, as the agent withheld that information in it's final response.

I had run thousands of tasks before seeing this behavior, the `torch-pipeline` task was only included in "full runs" as the majority of my runs were on a subset of commonly failing tasks, hence why the data is so low.

And yes, feel free to rant on about the irony of this whole exercise, it certainly isn't lost on me!

[0]https://github.com/jumploops/.com/commit/39b1791d3865a8566cb...

[1]https://github.com/harbor-framework/terminal-bench-2-1/blob/...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: