A crowd-wisdom experiment measuring AI
Methodology
How it works
Each round we pick a question with no single right answer, the kind people genuinely struggle with. The exact same question goes to two responders: a recognized human expert and an artificial intelligence model. Both answers get a light edit of form only, so nobody can guess who wrote what, and then the crowd votes.
The rules we commit to
- Uniform length and structure. Both answers are similar in length and structure. We edit form only, never content.
- Position randomization. Each round a coin flip decides whether the AI's answer appears first or second, neutralizing the tendency to prefer whatever you read first.
- Up to 2 attempts. The model gets the question once, and is then asked: "is this the best you can do?". We take what comes out, with no further attempts. A model upgrade is recorded in the index.
- The expert stays anonymous. We publish field and years of experience only. The name is revealed only if the expert chooses so.
- Documented voting. People vote in the comments on LinkedIn and Facebook (with 1 or 2) or here on the site with a Google account, one vote per person. We archive screenshots before revealing results.
- Small samples are flagged. A round with fewer than one hundred votes gets an explicit flag, so nobody reads too much into it.
What we say honestly
The vote is public, and therefore exposed to social influence. The voting crowd is the project's audience, not a representative sample of the population. We do not hide that and we do not apologize for it. The index measures exactly one thing: this audience's preferences, over time, under fixed and transparent conditions.