\
  The most prestigious law school admissions discussion board in the world.
BackRefresh Options Favorite

Insane how gpt-6 leapfrogged fable in metrics

...
The Penis
  09/05/26
Do you think I can maek it on police force at 40?
Tall Man? Yes, you can!
  09/05/26
idk. lets launch a micro-saas. can still maek it. just carve...
The Penis
  09/05/26
Absolutely. I read somewhere that in the past before we had ...
Richard Ames
  09/05/26
That's something I think about a lot. The Great Sorting of c...
Tall Man? Yes, you can!
  09/05/26
It is also one of those things that the majority of people, ...
Richard Ames
  09/05/26
….or bomb OCI.
cowgod
  09/05/26
Let's face it, anyone who truly bombs OCI is more or less sc...
Richard Ames
  09/05/26
Gemini still beats everything
scholarship
  09/05/26
They are certainly behind Fable and Astra based on peak perf...
,.,....,..,.,.,,,,..,..,.,..,.,.,.,...
  09/05/26
...
harnessman
  09/05/26
lol
The Penis
  09/05/26
...
scholarship
  09/05/26
in what? gpt seems numba 1 for everything, esp i trid writin...
AZNgirl stopping WhiteLady killing HandsomeBaby
  09/05/26
The flash model is faster than anything else and works well ...
\"\'\"\'\"\'\'\"\"\"
  09/05/26
What is max iq to think that benchmarks mean anything at all...
NCAA Football 2006 for PlayStation 2
  09/05/26
You should give your crank theory about why the metrics are ...
The Penis
  09/05/26
Benchmark progress is almost definitionally guaranteed to ha...
NCAA Football 2006 for PlayStation 2
  09/05/26
Even when new benchmarks are created, the rank ordering is h...
\"\'\"\'\"\'\'\"\"\"
  09/05/26
Wow you’re a dumbass nigger lol
Arkan
  09/05/26
The criticism of benchmarks reminds me of the criticisms by ...
\"\'\"\'\"\'\'\"\"\"
  09/05/26
This argument would be reasonable if the benchmarks actually...
NCAA Football 2006 for PlayStation 2
  09/05/26
The benchmarks that are being used to evaluate these models ...
\"\'\"\'\"\'\'\"\"\"
  09/05/26
This is all over the place. RL training for obscure mathemat...
NCAA Football 2006 for PlayStation 2
  09/05/26
arc-agi-3 tests generalized tasks that the model wasn't trai...
The Penis
  09/05/26


Poast new message in this thread



Reply Favorite

Date: September 5th, 2026 1:48 PM
Author: The Penis



(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117605)



Reply Favorite

Date: September 5th, 2026 2:12 PM
Author: Tall Man? Yes, you can! (🧐)

Do you think I can maek it on police force at 40?

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117649)



Reply Favorite

Date: September 5th, 2026 3:15 PM
Author: The Penis

idk. lets launch a micro-saas. can still maek it. just carve out that special niche and make 500K MRR within months

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117802)



Reply Favorite

Date: September 5th, 2026 3:23 PM
Author: Richard Ames

Absolutely. I read somewhere that in the past before we had such effective methods for sorting by aptitude, that you'd end up with way smarter police officers (and other such jobs) because people would go into whatever profession their family was in, regardless of ability.

We've obviously had less of this for 50+ years now. And I think that is part of why the police and other important fields are increasingly full of dumbasses.

You can be the tip of the spear in reversing this, friend.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117808)



Reply Favorite

Date: September 5th, 2026 3:37 PM
Author: Tall Man? Yes, you can! (🧐)

That's something I think about a lot. The Great Sorting of college meritocracy in the 20th century has basically hyper-concentrated intelligence in the highest paying/most prestigious fields and I think is a very underrated reason in the decline of many societal structures/industries we used to take for granted

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117829)



Reply Favorite

Date: September 5th, 2026 4:30 PM
Author: Richard Ames

It is also one of those things that the majority of people, even smart people, don't consider when they look at the bigger picture and the problems we face.

TBH I think the flip side of it is that a smart person can massively excel in their life by joining a non-preftigious field. But only if they're willing to forego preftige.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117915)



Reply Favorite

Date: September 5th, 2026 4:41 PM
Author: cowgod

….or bomb OCI.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117932)



Reply Favorite

Date: September 5th, 2026 5:52 PM
Author: Richard Ames

Let's face it, anyone who truly bombs OCI is more or less screwed. Even the non-preftigious fields will tell them to eat shit because of phenotype alone.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50118042)



Reply Favorite

Date: September 5th, 2026 3:27 PM
Author: scholarship

Gemini still beats everything

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117815)



Reply Favorite

Date: September 5th, 2026 4:02 PM
Author: ,.,....,..,.,.,,,,..,..,.,..,.,.,.,...


They are certainly behind Fable and Astra based on peak performance but the 3.8 flash model is actually pretty good for most things and is very fast.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117861)



Reply Favorite

Date: September 5th, 2026 5:00 PM
Author: harnessman



(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117951)



Reply Favorite

Date: September 5th, 2026 4:04 PM
Author: The Penis

lol

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117870)



Reply Favorite

Date: September 5th, 2026 4:37 PM
Author: scholarship



(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117921)



Reply Favorite

Date: September 5th, 2026 4:11 PM
Author: AZNgirl stopping WhiteLady killing HandsomeBaby

in what? gpt seems numba 1 for everything, esp i trid writing programs and gemini sucks, only good thing its free unlimited it seems

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117882)



Reply Favorite

Date: September 5th, 2026 4:24 PM
Author: \"\'\"\'\"\'\'\"\"\"

The flash model is faster than anything else and works well for most things. It's even pretty good at coding anymore based on DeepSWE, which has traditionally been Gemini's weak point.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117906)



Reply Favorite

Date: September 5th, 2026 4:06 PM
Author: NCAA Football 2006 for PlayStation 2

What is max iq to think that benchmarks mean anything at all in 2026

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117874)



Reply Favorite

Date: September 5th, 2026 4:40 PM
Author: The Penis

You should give your crank theory about why the metrics are meaningless to open ai and anthropic! I'm sure they will pay you big money for figuring out something their top engineers and researchers couldn't. I'll bet you are so high iq you could even come up with better metrics and save the industry!

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117928)



Reply Favorite

Date: September 5th, 2026 4:55 PM
Author: NCAA Football 2006 for PlayStation 2

Benchmark progress is almost definitionally guaranteed to happen because the process of constructing a benchmark is a direct precursor to the process of constructing a training dataset used for hill climbing that benchmark

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117942)



Reply Favorite

Date: September 5th, 2026 4:59 PM
Author: \"\'\"\'\"\'\'\"\"\"

Even when new benchmarks are created, the rank ordering is highly correlated to past benchmark scores. This is prior to any sort of hill climbing.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117948)



Reply Favorite

Date: September 5th, 2026 4:50 PM
Author: Arkan

Wow you’re a dumbass nigger lol

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117937)



Reply Favorite

Date: September 5th, 2026 4:56 PM
Author: \"\'\"\'\"\'\'\"\"\"

The criticism of benchmarks reminds me of the criticisms by people like Gould who disputed the existence of general intelligence in humans. It fails for similar reasons. The benchmarks are highly correlated, even private benchmark scores. They are clearly loading on a g factor equivalent, so the arbitrary nature of a particular benchmark doesn't matter much.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117943)



Reply Favorite

Date: September 5th, 2026 5:05 PM
Author: NCAA Football 2006 for PlayStation 2

This argument would be reasonable if the benchmarks actually corresponded to general increases in practical abilities (let alone g equivalent "IQ" lol) and were not specifically engineered to reflect "improvements" that are selected due to ease and efficiency to train for and in some cases profitable as a product

If what you are saying was true, we'd see models get continuously smarter in terms of "g" or "IQ." That's not what we're seeing. We are seeing them get barely incrementally smarter in terms of overall "IQ," while their specific abilities that they're RL trained for, like coding and math, are what increase

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117960)



Reply Favorite

Date: September 5th, 2026 5:15 PM
Author: \"\'\"\'\"\'\'\"\"\"

The benchmarks that are being used to evaluate these models now rely on obscure pieces of human knowledge that often only a few people in the world could competently handle. They have saturated practically everything that is easy so now they create ridiculous things like FrontierMath that get saturated in a year and a half. They are being evaluated much more broadly and deeply than any IQ test.

There are deficiencies in the models, but the reality of increasing practical usefulness is impossible to dispute. There is a reason why Anthropic and OpenAI revenue is growing so fast.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50117971)



Reply Favorite

Date: September 5th, 2026 5:39 PM
Author: NCAA Football 2006 for PlayStation 2

This is all over the place. RL training for obscure mathematical ability is the oppose of "broadly and deeply." I don't dispute that models are getting massively better at the tasks that they're being RL trained for. I'm saying that "benchmarks" are misleading and largely pointless because we know by definition that the models *must* be getting better at tasks they're being specifically RL trained for. But there is no corresponding way to "just RL train" for general intelligence/IQ

Also you guys (AI evangelists) always fall back on this trope where whenever someone points out the hard and real limitations of LLMs, you say oh but they're increasingly practically useful and profitable. Yes, no one disputes that, myself included. But that's just "moving the goalposts" away from the substance of the very real limitations of LLMs

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50118006)



Reply Favorite

Date: September 5th, 2026 5:53 PM
Author: The Penis

arc-agi-3 tests generalized tasks that the model wasn't trained on and gpt-6 saturated the metric and beat humans on it so lol at you. you are arguing outdated points.

(http://www.autoadmit.com/thread.php?thread_id=5900885&forum_id=2),#50118045)