English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Coq & Isabelle: Are They Still the Kings of Formal Reasoning? An In-Depth Research Report

Forum topic · ✨步子哥 · 2026-06-12

Summary

This forum post presents a comprehensive comparative study of Coq (now Rocq) and Isabelle/HOL, the two leading interactive theorem provers, and evaluates the rise of Lean 4 as of mid-2026. The author concludes that Coq and Isabelle remain unmatched in their respective domains: Coq dominates dependent type theory, software verification, and compiler correctness (CompCert, VST, Iris, the Four Color Theorem), while Isabelle/HOL leads in system-level verification and proof automation (seL4 microkernel, AWS Nitro Isolation Engine with 330k lines of proofs, Sledgehammer automation, and the Archive of Formal Proofs with 982 entries). Lean 4 shows explosive growth in the mathematical community—mathlib4 exceeding 5 million lines, 772 contributors, and the most PutnamBench formalizations (672)—plus superior developer experience via a unified language and VS Code integration. However, it lacks industrial verification milestones, compiler/kernel verification projects, and a battle-tested kernel. The report also surveys Agda, F*, and HOL Light, mapping the best tool per application domain and predicting a complementary three-polar landscape rather than one tool replacing the others.

Coq & Isabelle: Still the Kings of Formal Reasoning? — Research Report

> Research question: Are Coq and Isabelle still the pinnacle of formal reasoning tools? > Date: June 12, 2026 | Paradigm: Pragmatist comparative analysis | AI-assisted literature review with sourced claims.

Core conclusion

Coq and Isabelle retain unshakable leading positions in their respective strongholds, but their "supremacy" is complementary rather than overlapping. Coq dominates dependent type theory, software verification, and compiler correctness; Isabelle/HOL is unmatched in system-level verification, large-scale proof automation, and industrial deployment. Lean 4 is rising fast in developer experience and mathematical community growth, but its industrial verification track record is still shallow. The future is a complementary three-polar landscape, not a replacement contest.

Key points: Coq (now Rocq)

  • Renamed Rocq Prover in 2025; latest version Rocq 9.2.0 (released 2026-03-27). Built on the Calculus of Inductive Constructions (CIC), implemented in OCaml, 40+ years of development, ACM Software System Award winner.
  • CIC enables propositions-as-types (Curry–Howard), dependent types, inductive definitions, and code extraction to OCaml/Haskell.
  • Industrial milestones:
  • CompCert — the only end-to-end formally verified optimizing C compiler (2005–present).
  • VST 3.0 (Iris-based separation logic for C verification) and the Iris concurrency framework.
  • Formalizations of the Four Color Theorem (2005) and Feit–Thompson theorem (2012); MetaRocq bootstrapped kernel verification.
  • Industrial users include Google, AbsInt, BlueRock, Formal Vindication. LLM4Rocq explores LLM integration. PutnamBench formalizations: 412.
  • Key points: Isabelle/HOL

  • Built on classical higher-order logic, implemented in Standard ML + Scala, 35+ years of development at TU München. Structured Isar proof language and Sledgehammer automation (dispatching to E, Vampire, Z3, CVC4, etc.).
  • System-verification milestones:
  • seL4 microkernel (2009) — first full OS kernel verification from spec to C implementation.
  • AWS Nitro Isolation Engine (2026) — first formally verified hypervisor deployed in a commercial cloud: written in μRust, 330k lines of machine-checked proofs, running in production on Graviton5 hardware.
  • Isabelle/Solidity (2025) and AutoCorrode Rust verification infrastructure (2025).
  • Archive of Formal Proofs (AFP): 982 entries, ~5.2M lines of code, ~316,900 lemmas, 593 authors. PutnamBench formalizations: 640.
  • Key points: Lean 4 — the strongest challenger

  • Dependency type theory; implemented in itself (C++ runtime); led by Leonardo de Moura; released 2023.
  • Key innovation: metaprogramming and proofs share one language, giving best-in-class debugging, transparency, and VS Code integration.
  • mathlib4: >5,000,000 lines (as of March 2026), 132,448 definitions, 278,346 theorems, 772 contributors. Endorsed by Terence Tao. PutnamBench: 672 formalizations — the most of the three.
  • Remaining weaknesses: zero industrial verification milestones (no CompCert/seL4-scale projects), no verified compiler or kernel, younger and less-audited trusted kernel (C++ runtime), automation below Sledgehammer level, smaller industrial adoption.
  • Other competitors

  • Agda: superb dependent type expressiveness, weak automation — beloved by PL researchers, not a general verification rival.
  • F*: dependent types + refinement types + SMT; strong for program verification (Project Everest TLS), but narrowly focused.
  • HOL Light / HOL4: minimal trusted computing base (Flyspeck), but weaker automation and smaller ecosystem than Isabelle/HOL.
  • Domain-to-tool mapping

    | Domain | Best tool | |---|---| | C compiler verification | Coq (CompCert) | | OS kernel verification | Isabelle (seL4) | | Cloud isolation verification | Isabelle (AWS Nitro) | | Concurrent program verification | Coq (Iris) | | Large-scale math formalization | Lean 4 / Isabelle | | Smart contract verification | Isabelle (Solidity) | | Teaching / onboarding | Lean 4 | | Minimal TCB | HOL Light | | AI + formal proofs | Lean 4 / Coq (Pantograph / LLM4Rocq) |

    Discussion: why Coq & Isabelle remain on top

    1. Industrial verification cannot be rushed: CompCert took ~20 years; seL4 began in 2004 and evolved into AWS Nitro by 2026. The hard part is designing verifiable specifications and proof strategies, not writing code. 2. Complementary philosophies: Coq's constructive logic (CIC) enables proof-to-program extraction at the cost of excluded middle; Isabelle's classical HOL + Sledgehammer maximizes automation at the cost of extraction. Lean 4's middle path offers neither extreme fully. 3. Ecosystem accumulation: AFP's 982 verified entries and Rocq's hundreds of packages are reusable, machine-checked knowledge — not just line counts. 4. No winner-takes-all dynamics: formal verification is inherently multi-paradigm; each tool rules its own kingdom.

    Caveats

  • Lean 4's mathematical community could gradually marginalize MathComp and AFP in the math domain (though not in software/system verification).
  • The developer-experience gap (Lean's VS Code vs. CoqIDE/jEdit) may push newcomers toward Lean 4, though Rocq Platform and LLM4Rocq aim to close it.
  • LLM integration (Pantograph vs. LLM4Rocq) may reshuffle the balance of power over the next five years.
*Note: the original post was truncated; later sections (references, etc.) were not included.*

Tags

#coq#rocq#isabelle#lean-4#formal-verification#theorem-provers#compcert#sel4

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981148