←back to thread

2127 points bakugo | 1 comments | | HN request time: 0.363s | source
1. zone411 ◴[] No.43166755[source]
Claude 3.7 Sonnet Thinking scores 33.5 (4th place after o1, o3-mini, and DeepSeek R1) on my Extended NYT Connections benchmark. Claude 3.7 Sonnet scores 18.9. I'll run my other benchmarks in the upcoming days.

https://github.com/lechmazur/nyt-connections/