GPT-5.2

(openai.com)

https://platform.openai.com/docs/guides/latest-model

System card: https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944...

Show context

josalhor ◴[11 Dec 25 18:24 UTC] No.46235005[source]▶

>>46234788 (OP) #

From GPT 5.1 Thinking:

ARC AGI v2: 17.6% -> 52.9%

SWE Verified: 76.3% -> 80%

That's pretty good!

replies(7): >>46235062 #>>46235070 #>>46235153 #>>46235160 #>>46235180 #>>46235421 #>>46236242 #

verdverm ◴[11 Dec 25 18:28 UTC] No.46235062[source]▶

>>46235005 #

We're also in benchmark saturation territory. I heard it speculated that Anthropic emphasizes benchmarks less in their publications because internally they don't care about them nearly as much as making a model that works well on the day-to-day

replies(5): >>46235126 #>>46235266 #>>46235466 #>>46235492 #>>46235583 #

Mistletoe ◴[11 Dec 25 18:43 UTC] No.46235266[source]▶

>>46235062 #

How do you measure whether it works better day to day without benchmarks?

replies(3): >>46235305 #>>46235348 #>>46235398 #

1. verdverm ◴[11 Dec 25 18:51 UTC] No.46235398{3}[source]▶

>>46235266 #

Internal evals, Big AI certainly has good, proprietary training and eval data, it's one reason why their models are better

replies(1): >>46235532 #

2. aydyn ◴[11 Dec 25 18:58 UTC] No.46235532[source]▶

>>46235398 (TP) #

Then publish the results of those internal evals. Public benchmark saturation isn't an excuse to be un-quantitative.

replies(1): >>46235607 #

3. verdverm ◴[11 Dec 25 19:03 UTC] No.46235607[source]▶

>>46235532 #

How would published numbers be useful without knowing what the underlying data being used to test and evaluate them are? They are proprietary for a reason

To think that Anthropic is not being intentional and quantitative in their model building, because they care less for the saturated benchmaxxing, is to miss the forest for the trees

replies(1): >>46236582 #

4. aydyn ◴[11 Dec 25 20:17 UTC] No.46236582{3}[source]▶

>>46235607 #

Do you know everything that exists in public benchmarks?

They can give a description of what their metrics are without giving away anything proprietary.

replies(1): >>46238542 #

5. verdverm ◴[11 Dec 25 23:00 UTC] No.46238542{4}[source]▶

>>46236582 #

I'd recommend watching Nathan Lambert's video he dropped yesterday on Olmo 3 Thinking. You'll learn there's a lot of places where even descriptions of proprietary testing regimes would give away some secret sauce

Nathan is at Ai2 which is all about open sourcing the process, experience, and learnings along the way

replies(1): >>46241985 #

6. aydyn ◴[12 Dec 25 08:18 UTC] No.46241985{5}[source]▶

>>46238542 #

Thanks for the reference I'll check it out. But it doesnt really take away from the point I am making. If a level of description would give away proprietary information, then go one level up to a more vague description. How to describe things to a proper level is more of a social problem than a technical one.

↑