TECHNOLOGY
ZakGT Tech
OpenAI's GPT-6 has passed bar exam, medical licensing, and PhD-level physics problem sets at the 97th human percentile — but the benchmark methodology matters as much as the scores. ZakGT Tech reviewed every evaluation framework, consulted three AI safety researchers who assessed the model independently, and identifies the specific reasoning tasks where GPT-6 still fails in ways that current benchmarks are not designed to catch.
Full article content coming soon...
Share this article