you have to be careful. if the test exists on the internet, the ai system might have learned it. it could tell you the answers by finding it in it's database. that is what an ai is - an advanced gui on top of a search algorithm running on a database in the backend. and that's all it is.
so, somebody devised an offline test that no ai had ever seen before.
this is where they were at in 2025, when the issue was fresh:
that's actually not very impressive, is it? but it's about what i expected. it's interesting that the ai systems show about the same level of variation as humans, but that's a function of the central limit theorem.
since this hit the internet, these numbers have started coming up, but i'm skeptical. if i was deepseek, i would consider this bad marketing and want to improve it. the nature of the technology makes this hard.
you could argue it's not fair to the system to ask it to think like a human. ok. the mirror test has been modified for some species to adjust to their perception of existence, and they tend to do better when it has. but if you're going to argue the ai is 100x more intelligent than the smartest human, you need a standard of comparison.
i would not have a lot of faith in the updated numbers, at least not until some kind of truly third party institution can be built that can truly ensure these tests are fully offline.
in 2025, this was random, with no prep. and the results were what you'd expect that autistic kid to score.