Then you should be surprised that turbo-instruct actually plays well, right? We ...

akira2501 · 2024-11-15T08:51:51 1731660711

There are some who suggest that modern chess is mostly a game of memorization and not one particularly of strategy or skill. I assume this is why variants like speed chess exist.

In this scope, my mental model is that LLMs would be good at modern style long form chess, but would likely be easy to trip up with certain types of move combinations that most humans would not normally use. My prediction is that once found they would be comically susceptible to these patterns.

Clearly, we have no real basis for saying it is "good" or "bad" at chess, and even using chess performance as an measurement sample is a highly biased decision, likely born out of marketing rather than principle.

mewpmewp2 · 2024-11-15T10:38:03 1731667083

It is memorisatiom only after you have grandmastered reasoning and strategy.

DiogenesKynikos · 2024-11-15T10:39:59 1731667199

Speed chess relies on skill.

I think you're using "skill" to refer solely to one aspect of chess skill: the ability to do brute-force calculations of sequences of upcoming moves. There are other aspects of chess skill, such as:

1. The ability to judge a chess position at a glance, based on years of experience in playing chess and theoretical knowledge about chess positions.

2. The ability to instantly spot tactics in a position.

In blitz (about 5 minutes) or bullet (1 minute) chess games, these other skills are much more important than the ability to calculate deep lines. They're still aspects of chess skill, and they're probably equally important as the ability to do long brute-force calculations.

henearkr · 2024-11-15T14:53:12 1731682392

> tactics in a position

That should give patterns (hence your use of the verb to "spot" them, as the grandmaster would indeed spot the patterns) recognizable in the game string.

More specifically grammar-like parterns, e.g. the same moves but translated.

Typically what an LLM can excel at.

the_af · 2024-11-15T14:25:01 1731680701

> Then you should be surprised that turbo-instruct actually plays well, right?

Do we know it's not special-casing chess and instead using a different engine (not an LLM) for playing?

To be clear, this would be an entirely appropriate approach to problem-solving in the real world, it just wouldn't be the LLM that's playing chess.

mda · 2024-11-15T15:29:31 1731684571

Yes, probably there is more going on here, e.g. it is cheating.

flyingcircus3 · 2024-11-15T06:15:24 1731651324

"playing strong chess" would be a much less hand-wavy claim if there were lots of independent methods of quantifying and verifying the strength of stockfish's lowest difficulty setting. I honestly don't know if that exists or not. But unless it does, why would stockfish's lowest difficulty setting be a meaningful threshold?

golol · 2024-11-15T06:53:41 1731653621

I've tried it myself, GPT-3.5-turbo-instruct was at least somewhere in the rabge 1600-1800 ELO.

niobe · 2024-11-17T01:00:39 1731805239

But to some approximation we do know how an LLM plays chess.. based on all the games, sites, blogs, analysis in its training data. But it has a limited ability to tell a good move from a bad move since the training data has both, and some of it lacks context on move quality.

Here's an experiment: give an LLM a balanced middle game board position and ask it "play a new move that a creative grandmaster has discovered, never before played in chess and explain the tactics and strategy behind it". Repeat many times. Now analyse each move in an engine and look at the distribution of moves and responses. Hypothesis: It is going to come up with a bunch of moves all over the ratings map with some sound and some fallacious arguments.

I really don't think there's anything too mysterious going on here. It just synthesizes existing knowledge and gives answers that includes bit hits, big misses and everything in between. Creators chip away at the edges to change that distribution but the fundamental workings don't change.