More than seventeen years after Bitcoin’s launch, the question of whether Satoshi Nakamoto was a single individual or part of a coordinated group remains one of the protocol’s most consequential unsolved mysteries. So we asked five of today’s leading artificial intelligence (AI) models to settle the debate.
Key Takeaways
Kimi K26 was the only model favoring a group, at 50% versus 45% for a sole creator.ChatGPT 5.6 Sol alone reconciled its probabilities to a consistent 54% solo estimate.Gemini Pro, Grok 4.3 and peers show Satoshi debates will likely persist beyond 2026.To explore how leading artificial intelligence (AI) models assess this question with structured probabilistic reasoning, we ran the same prompt across a panel of frontier systems, including Grok 4.3, Claude Fable 5, ChatGPT 5.6 Sol, Gemini Pro, and Kimi K26. Each model was asked to build a simple Bayesian scenario tree, assign probabilities to the most likely explanations, and estimate the chances that Satoshi operated alone versus as part of a group.
The prompt given to each AI model was:
“Using a simple Bayesian scenario tree, identify the three most likely scenarios for whether Satoshi Nakamoto was a sole individual or part of a group (such as a small coordinated team or collective), and assign each a probability. Ensure the three scenarios plus an “other scenarios” category total 100%. Based on those weightings, estimate the overall probability that Satoshi was a sole individual versus part of a group. Respond in one paragraph of no more than 100 words, distinguishing evidence-based assumptions from speculation in one sentence.”
“Explain the reasoning behind your specific prediction in one paragraph of no more than 100 words, citing the primary factors, assumptions, and logical steps that led you to this conclusion.”
Kimi K26 Thinking answer:Kimi K26 reasoning:
I weighted sole authorship at 45% because the whitepaper and code exhibit a unified voice, consistent C++ style, and narrow expertise window typical of one polymath, yet the 2008–2010 round-the-clock forum activity and simultaneous cryptography, economics, and coding mastery suggest a small coordinated team (35%). An institutional collective (15%) is less likely given the absence of leaks or bureaucratic artifacts, while “other” captures unknown unknowns (5%). These priors update on the evidence of sustained secrecy and stylistic coherence, yielding a 45% solo versus 50% group probability.
Claude Fable 5 High Intelligence answer:Claude Fable reasoning:
I anchored on stylometric and behavioral evidence: Satoshi’s writings show a consistent voice, spelling conventions, and idiosyncrasies (British spellings, double-spacing after periods), and the original codebase reflects a single programmer’s style — favoring a sole author. However, the project’s breadth (cryptography, economics, C++ engineering, peer-to-peer networking) exceeds typical individual expertise, and gaps in activity suggest possible consultation, so I reserved meaningful weight for informal collaboration. I assumed public artifacts (emails, forum posts, code) are representative and not deliberately obfuscated. Combining these via rough likelihood judgments — not formal priors — produced the 50/25/15/10 split.
Grok 4.3 Expert Mode answer:Grok 4.3 reasoning:
ChatGPT 5.6 Sol High Intelligence answer:ChatGPT 5.6 Sol reasoning:
Gemini Pro Extended Thinking answer:Gemini Pro reasoning:
The 70% probability for a sole individual rests on three primary factors: the original codebase’s stylistic consistency, the unified “voice” across emails, and timestamp data indicating a single person’s sleep schedule. The core assumption is that a group would inevitably leak diverse coding habits or overlapping active hours. Logically, the deep, multidisciplinary expertise required in cryptography and economics points strongly to an established cypherpunk (60%) rather than an isolated unknown (10%). However, this same extreme complexity justifies the 25% team probability, acknowledging that a highly disciplined, small collective could theoretically mask their collaboration behind one persona.
Weighing the EvidenceThe gap also exposes how loosely “Bayesian” gets applied. Most models skipped the actual math: four of five failed to reconcile their own scenario percentages back to a clean two-way sole versus group split, quietly dropping their “other” category rather than accounting for it. ChatGPT 5.6 Sol was the exception, showing its allocation work end-to-end. Gemini Pro took the opposite approach, splitting “sole individual” into two sub-scenarios and combining them, a framing choice that likely explains its outlier 70% figure more than any unique evidence it cited.
What the models do agree on is more telling than where they diverge. All five pointed to the same two pillars, consistent writing voice and consistent coding style, as their evidence base, and all five drew an identical line between that evidence and pure speculation about specific identities or institutional backers. That convergence suggests the real disagreement among these AI systems isn’t about the underlying facts of the Satoshi mystery. It’s about how each one structures uncertainty and assigns weight to evidence that, seventeen years on, still resolves nothing definitively.



















