←back to thread

GPT-5.2

(openai.com)
1019 points atgctg | 2 comments | | HN request time: 0s | source
Show context
josalhor ◴[] No.46235005[source]
From GPT 5.1 Thinking:

ARC AGI v2: 17.6% -> 52.9%

SWE Verified: 76.3% -> 80%

That's pretty good!

replies(7): >>46235062 #>>46235070 #>>46235153 #>>46235160 #>>46235180 #>>46235421 #>>46236242 #
1. fuddle ◴[] No.46236242[source]
I don't think SWE Verified is an ideal benchmark, as the solutions are in the training dataset.
replies(1): >>46236752 #
2. joshuahedlund ◴[] No.46236752[source]
I would love for SWE Verified to put out a set of fresh but comparable problems and see how the top performing models do, to test against overfitting.