Coq & Isabelle: Still the Kings of Formal Reasoning? — Research Report
> Research question: Are Coq and Isabelle still the pinnacle of formal reasoning tools? > Date: June 12, 2026 | Paradigm: Pragmatist comparative analysis | AI-assisted literature review with sourced claims.
Core conclusion
Coq and Isabelle retain unshakable leading positions in their respective strongholds, but their "supremacy" is complementary rather than overlapping. Coq dominates dependent type theory, software verification, and compiler correctness; Isabelle/HOL is unmatched in system-level verification, large-scale proof automation, and industrial deployment. Lean 4 is rising fast in developer experience and mathematical community growth, but its industrial verification track record is still shallow. The future is a complementary three-polar landscape, not a replacement contest.
Key points: Coq (now Rocq)
- Renamed Rocq Prover in 2025; latest version Rocq 9.2.0 (released 2026-03-27). Built on the Calculus of Inductive Constructions (CIC), implemented in OCaml, 40+ years of development, ACM Software System Award winner.
- CIC enables propositions-as-types (Curry–Howard), dependent types, inductive definitions, and code extraction to OCaml/Haskell.
- Industrial milestones:
- CompCert — the only end-to-end formally verified optimizing C compiler (2005–present).
- VST 3.0 (Iris-based separation logic for C verification) and the Iris concurrency framework.
- Formalizations of the Four Color Theorem (2005) and Feit–Thompson theorem (2012); MetaRocq bootstrapped kernel verification.
- Industrial users include Google, AbsInt, BlueRock, Formal Vindication. LLM4Rocq explores LLM integration. PutnamBench formalizations: 412.
- Built on classical higher-order logic, implemented in Standard ML + Scala, 35+ years of development at TU München. Structured Isar proof language and Sledgehammer automation (dispatching to E, Vampire, Z3, CVC4, etc.).
- System-verification milestones:
- seL4 microkernel (2009) — first full OS kernel verification from spec to C implementation.
- AWS Nitro Isolation Engine (2026) — first formally verified hypervisor deployed in a commercial cloud: written in μRust, 330k lines of machine-checked proofs, running in production on Graviton5 hardware.
- Isabelle/Solidity (2025) and AutoCorrode Rust verification infrastructure (2025).
- Archive of Formal Proofs (AFP): 982 entries, ~5.2M lines of code, ~316,900 lemmas, 593 authors. PutnamBench formalizations: 640.
- Dependency type theory; implemented in itself (C++ runtime); led by Leonardo de Moura; released 2023.
- Key innovation: metaprogramming and proofs share one language, giving best-in-class debugging, transparency, and VS Code integration.
- mathlib4: >5,000,000 lines (as of March 2026), 132,448 definitions, 278,346 theorems, 772 contributors. Endorsed by Terence Tao. PutnamBench: 672 formalizations — the most of the three.
- Remaining weaknesses: zero industrial verification milestones (no CompCert/seL4-scale projects), no verified compiler or kernel, younger and less-audited trusted kernel (C++ runtime), automation below Sledgehammer level, smaller industrial adoption.
- Agda: superb dependent type expressiveness, weak automation — beloved by PL researchers, not a general verification rival.
- F*: dependent types + refinement types + SMT; strong for program verification (Project Everest TLS), but narrowly focused.
- HOL Light / HOL4: minimal trusted computing base (Flyspeck), but weaker automation and smaller ecosystem than Isabelle/HOL.
- Lean 4's mathematical community could gradually marginalize MathComp and AFP in the math domain (though not in software/system verification).
- The developer-experience gap (Lean's VS Code vs. CoqIDE/jEdit) may push newcomers toward Lean 4, though Rocq Platform and LLM4Rocq aim to close it.
- LLM integration (Pantograph vs. LLM4Rocq) may reshuffle the balance of power over the next five years.
Key points: Isabelle/HOL
Key points: Lean 4 — the strongest challenger
Other competitors
Domain-to-tool mapping
| Domain | Best tool | |---|---| | C compiler verification | Coq (CompCert) | | OS kernel verification | Isabelle (seL4) | | Cloud isolation verification | Isabelle (AWS Nitro) | | Concurrent program verification | Coq (Iris) | | Large-scale math formalization | Lean 4 / Isabelle | | Smart contract verification | Isabelle (Solidity) | | Teaching / onboarding | Lean 4 | | Minimal TCB | HOL Light | | AI + formal proofs | Lean 4 / Coq (Pantograph / LLM4Rocq) |
Discussion: why Coq & Isabelle remain on top
1. Industrial verification cannot be rushed: CompCert took ~20 years; seL4 began in 2004 and evolved into AWS Nitro by 2026. The hard part is designing verifiable specifications and proof strategies, not writing code. 2. Complementary philosophies: Coq's constructive logic (CIC) enables proof-to-program extraction at the cost of excluded middle; Isabelle's classical HOL + Sledgehammer maximizes automation at the cost of extraction. Lean 4's middle path offers neither extreme fully. 3. Ecosystem accumulation: AFP's 982 verified entries and Rocq's hundreds of packages are reusable, machine-checked knowledge — not just line counts. 4. No winner-takes-all dynamics: formal verification is inherently multi-paradigm; each tool rules its own kingdom.