The ranking
Click any row to read the judges' reasoning and the evidence they cited. Sort by any column.
Ranks joined by a bracket in the left margin are statistically tied: repeat
grading of an unchanged transcript moves the overall score by about 1.54 points, so gaps
smaller than 4.3 points do not separate two leaders.
Insight — 45% of composite
Technical depth — 35%
Clarity — 20%
What is being scored
Each transcript is graded on three dimensions, each on a 1–100 scale, supported by
15 sub-criteria that a judge may also mark not observed when the format gave no
opportunity to demonstrate them. The composite is
0.45×Insight + 0.35×Technical + 0.20×Clarity.
- Insight (45%) — causal and counterfactual reasoning, originality, strategic
tradeoffs, understanding of incentives, calibration of confidence, engagement with opposing
arguments, and whether the speaker recomputes when a premise changes.
- Technical depth (35%) — mechanism, magnitudes and denominators used correctly,
operational and industry specificity, and movement between strategy and implementation.
- Clarity (20%) — directness, structure, concrete language, economy. Deliberately
the smallest weight: the rubric explicitly refuses to reward fluency, charisma, or a confident
delivery, because a polished non-answer is the failure mode being guarded against.
The pipeline
- A roster of 40 was built by 12 parallel sector scouts pooling 240 candidates, merged, then
attacked by four adversarial critics checking fame, availability, factual accuracy, and coverage
bias. Contested cases were settled by measuring actual long-form supply, not by argument.
- Transcripts are verbatim automatic captions of publicly posted recordings, pulled
programmatically with timestamps. No model paraphrased or summarised them at any point.
- Every transcript passed deterministic quality gates before grading: out-of-vocabulary rate,
speech-recognition repetition loops, type-token ratio, words per minute, and minimum length.
- Speaker and company names were replaced with
[SUBJECT] and [COMPANY],
including speech-recognition manglings of the surname found by phonetic matching. This step is
covered by tests, after an early version replaced the contraction “that’s” with
[SUBJECT] 34 times in one transcript.
- Two judges graded every transcript independently: Claude Fable 5.1 at maximum reasoning
effort, and OpenAI GPT-6 Astra at maximum reasoning effort. Model identity was asserted from
each call's telemetry rather than assumed.
How reliable is a score?
One unchanged transcript was graded five times by each judge under identical conditions. The
spread that produced is pure method noise.
- Repeat grading moves the composite by about 1.54 points of standard deviation.
- So two leaders differing by less than 4.3 points
are not distinguishable. The bracket in the rank column marks those groups.
- Coverage and venue-difficulty judgements were perfectly stable across repeats
(standard deviation 0.00), so the judges read the same conversation the same way every time.
The two judges disagree in a specific, correctable way.
Across blinded grades the raw means were Fable 54.7 and
Astra 63.4 on insight, a consistent offset rather than
genuine disagreement about who is impressive. Each judge's distribution is therefore recentred on the
pooled distribution before averaging, so only real disagreement moves a leader.
Correlation between judges on the composite: r = 0.924.
Mean absolute gap: 8.6 points.
What this cannot tell you
- Blinding removed the name, not the identity. Judges recognised the speaker anyway in
100% of blinded transcripts, from products, projects, and context. Defeating that
would mean stripping the technical content the study exists to measure. Every transcript was
therefore also graded unblinded, and the Halo column reports the gap: how many points a
leader gains once the judges are told who they are. The published composite is the blinded one.
- Captions have no speaker labels. Judges separated the subject's speech from the
interviewer's by context and reported their confidence and the subject's estimated share of the
talking. Both are shown in each transcript card.
- Venue difficulty is reported, never corrected for. A leader who only sits for friendly
interviews will score lower on insight, because a soft conversation cannot demonstrate reasoning
under pressure. Adjusting for that would mean inventing a score for a conversation that never
happened.
- Sampling is not exhaustive. A handful of appearances per leader is a sample of a public
speaking record, not the whole of it. Leaders marked low confidence have too few transcripts for
their rank to be trusted.
- Judges are language models. Two frontier models with different training agreeing on a
ranking is meaningful evidence, and it is not the same thing as being correct.
Who was excluded, and why
Excluded names are recorded rather than silently omitted, because silence looks like an
oversight. 12 people famous enough for consideration were left out, most for lack of
retrievable long-form public speech: John Ternus, Liang Wenfeng, Lei Jun, Hock Tan, Masayoshi Son, Shou Zi Chew, Jony Ive, Michael Truell, Larry Page, Amjad Masad, Ali Ghodsi, Matt Garman.