Something weird is happening with LLMs and chess

(dynomight.substack.com)

696 points crescit_eundo | 3 comments | 14 Nov 24 17:05 UTC | HN request time: 0.738s | source

Show context

codeflo ◴[15 Nov 24 10:52 UTC] No.42145710[source]▶

At this point, we have to assume anything that becomes a published benchmark is specifically targeted during training. That's not something specific to LLMs or OpenAI. Compiler companies have done the same thing for decades, specifically detecting common benchmark programs and inserting hand-crafted optimizations. Similarly, the shader compilers in GPU drivers have special cases for common games and benchmarks.

replies(3): >>42146244 #>>42146391 #>>42151266 #

darkerside ◴[15 Nov 24 12:21 UTC] No.42146244[source]▶

>>42145710 #

VW got in a lot of trouble for this

replies(10): >>42146543 #>>42146550 #>>42146553 #>>42146556 #>>42146560 #>>42147093 #>>42147124 #>>42147353 #>>42147357 #>>42148300 #

sigmoid10 ◴[15 Nov 24 13:08 UTC] No.42146560[source]▶

>>42146244 #

Apples and oranges. VW actually cheated on regulatory testing to bypass legal requirements. So to be comparable, the government would first need to pass laws where e.g. only compilers that pass a certain benchmark are allowed to be used for purchasable products and then the developers would need to manipulate behaviour during those benchmarks.

replies(3): >>42146749 #>>42147885 #>>42150309 #

0xFF0123 ◴[15 Nov 24 13:32 UTC] No.42146749[source]▶

>>42146560 #

The only difference is the legality. From an integrity point of view it's basically the same

replies(7): >>42146884 #>>42146984 #>>42147072 #>>42147078 #>>42147443 #>>42147742 #>>42147978 #

currymj ◴[15 Nov 24 14:54 UTC] No.42147443[source]▶

>>42146749 #

VW was breaking the law in a way that harmed society but arguably helped the individual driver of the VW car, who gets better performance yet still passes the emissions test.

replies(2): >>42147637 #>>42149872 #

1. jimmaswell ◴[15 Nov 24 15:15 UTC] No.42147637[source]▶

>>42147443 #

And afaik the emissions were still miles ahead of a car from 20 years prior, just not quite as extremely stringent as requested.

replies(1): >>42148188 #

2. slowmotiony ◴[15 Nov 24 16:11 UTC] No.42148188[source]▶

>>42147637 (TP) #

"not quite as extremely stringent as requested" is a funny way to say they were emitting 40 times more toxic fumes than permitted by law.

replies(1): >>42201013 #

3. linksnapzz ◴[21 Nov 24 04:09 UTC] No.42201013[source]▶

>>42148188 #

40x infinitesimal is still...infinitesimal.

↑