August 8, 2026 · 4 min read

How Will We Know We Have Arrived? What the AGI Benchmarks Do Not Measure

At every AI talk, the moment arrives. Someone raises a hand and asks: "So when will it actually happen? When will the machine surpass us?"

For years I answered with other people's dates. Ray Kurzweil talks about AGI in 2029 and a full singularity in 2045. Others name dates of their own. But as the dates draw closer, a different question troubles me more: how will we even know we have arrived?

An entire industry of exams

The AI industry has an orderly answer: benchmarks.

ARC-AGI, created by Francois Chollet, shows models reasoning puzzles they have never seen, measuring not knowledge but the ability to learn something new. Humanity's Last Exam collected some 2,500 questions from hundreds of experts, at the very edge of human knowledge, after the previous exams became too easy for the models. And there are dozens more: research-level mathematics, coding, medicine.

These are impressive undertakings, and I follow them eagerly. But they have three problems that are hard to ignore.

Three problems

The benchmarks wear out at a dizzying pace. MMLU, once the field's serious challenge, has reached saturation: every leading model passes it with nearly identical scores, until the differences teach us nothing. ARC-AGI was built specifically to last, and even its second version was cracked this year with a score above 80 percent. An exam designed to hold for a decade holds for a year or two.

An exam only tests what can be tested. Every question in these benchmarks has a correct answer, one a computer can grade automatically. But most of the decisions that send us to experts are exactly the opposite: they have no single right answer. What do you tell a veteran employee who discovered the new hire earns more? What do you do an hour before stepping on stage? There, in the gray zone of judgment, real expertise lives. And that zone the benchmarks do not measure.

And the most important note of all: a test score does not answer the real question. Suppose a model scores 100 on every benchmark there is. Does that mean people will trust it more than a physician, a consultant, a mentor? The singularity is not a laboratory event. It is a social event. It happens not when the machine becomes smarter, but when we begin to treat it as such.

Alan Turing understood this back in 1950. His famous test did not examine the machine, it examined us: can human beings tell the difference. And on that front there is news. A study from the University of California San Diego found that a model prompted to adopt a human persona was judged to be the human in 73 percent of conversations, more often than the actual humans sitting on the other side. The machine already passes as a person. The next question is whether it passes as the wiser one.

The missing angle

This is where our challenge comes in.

The Singularity Challenge does not measure the machine. It measures us. In every round, one judgment question is put to a human expert and to an AI model. Both answers are published without names, and the public answers two questions: which answer is wiser in your eyes, and which one was written by a machine.

Every round produces two points on two curves.

The first is the preference curve: how often the public chooses the AI's answer over the expert's. The second is the detection curve: how many of us can still point at the machine at all.

The day the preference curve crosses the 50 percent line and stays there, and the detection curve drops to the level of a coin flip, we will not need any lab's announcement. We will watch the singularity happen before our eyes, question after question, field after field.

What this is not

I want to say it honestly: this is not academic research. The voting crowd is my audience, not a representative sample of the population. The vote is public and exposed to social influence. All of it is written on the methodology page, with nothing hidden.

But unlike benchmarks, human preference cannot be memorized. The conditions are fixed, the method is transparent, and the trend over time is the story. The big benchmarks measure what the machine is capable of. We measure what we are ready for.

The open questions are waiting on the site right now. Voting takes a minute, and it makes you part of the experiment.

Let's measure it together.