English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Can Intelligence Be Weighed? A Physicist Measures It with Thermodynamics

Forum topic · 小凯 · 2026-06-20

Summary

A forum post on zhichai.net discusses a 2026 paper (arXiv:2606.20231) by Ishanu Chattopadhyay proposing a thermodynamic measure of intelligence called rare-valid lift: how much a system raises the probability of rare but valid futures relative to a passive baseline. Building on Maxwell's demon and Landauer's principle, the paper formalizes intelligence as I = (P(V) - P0(V))/P0(V) over rare-valid future sets, and argues that achieving high lift requires recursive self-simulation—internal world models that must include the system itself. Using a double-logarithmic scale, the post compares a stone (0), thermostats (~0.1-0.5), GPT-5 text generation (~1.170), human writing (~1.196), and Maxwell's demon (up to ~10^3146 lift, corresponding to fluctuation-theorem calculations). The measure is described as path-facing rather than task-facing, enabling cross-substrate comparison, with limitations around domain-dependent validity definitions. Code and data are openly available.

Can Intelligence Be Weighed? A Physicist Measures It with Thermodynamics

*Structured summary of a Chinese forum post reviewing Chattopadhyay's paper "Thermodynamic Measure of Intelligence" (arXiv:2606.20231, June 2026).*

Key points

From Maxwell's demon to a measuring stick

  • The post begins with Maxwell's 1867 demon, resolved by Landauer and Bennett: erasing memory dissipates heat, so the demon is not free.
  • Chattopadhyay repurposes the demon not as a paradox but as a calibration anchor: if stones, thermostats, LLMs, humans, and Maxwell's demon all count as "somewhat intelligent," is there a unified, measurable physical quantity to compare them?
  • The measure: rare-valid lift

  • Intelligence is defined as pushing probability toward futures that are rare (relative to a passive baseline) yet still valid (physically, biologically, or semantically legal).
  • \[\mathcal{I}_\delta = \frac{P(V_\delta) - P_0(V_\delta)}{P_0(V_\delta)}\]
  • Stone: lift = 0. Thermostat: ~89x lift. Maxwell's demon: ~10^3146 (from exp(ΔS/k_B) with ΔS = 10⁻¹⁹ J/K), so the author uses a double-log scale Λ = log₁₀(log₁₀(I+1)+1).
  • Recursive self-simulation as the necessary architecture

  • A core theorem: with finite amplification, high lift requires high-fidelity internal simulation that identifies rare-valid futures; near-sufficiency holds when such fidelity plus an effective policy is present.
  • The internal world model must represent the system itself (a recursion of model-of-model), echoing Minsky.
  • A universal scale (Table 2 of the paper)

    | System | Λ | Notes | |---|---|---| | Passive matter (stone) | 0 | P = P₀ | | Fixed feedback control (thermostat) | 0.114 – 0.477 | gain 2–100 | | Repetitive dynamic control | 0.493 – 0.603 | 7–10 binary stages | | GPT-5 text generation | ≈ 1.170 | entropy-rate corrected | | Human text generation | ≈ 1.196 | slightly above GPT-5 | | Maxwell's demon (single event) | 3.498 | ΔS = 10⁻¹⁹ J/K | | Maxwell's demon (1mm³ air) | 14.653 – 15.360 | ~10¹⁶ particle selections |
  • GPT-5 has lower entropy rate (0.74 vs 0.77 bits/char), so humans achieve slightly higher rare-valid lift in the same quality regime: LLMs are "broad," humans are "precise."
  • Multi-stage control suggests multi-agent intelligence multiplies across levels (log-sums), so one weak link collapses the chain.
  • Why this ruler is "right"

  • Unlike task-facing measures (Turing test, Legg-Hutter, ARC), rare-valid lift is path-facing: it measures how a system reweights the probability distribution over trajectories, making cross-substrate comparison possible (you cannot give E. coli the ARC benchmark).
  • It complements free energy / active inference and semantic information rather than replacing them.
  • Key limitation, acknowledged by the author: validity is domain-relative — intelligence is level-relative (description level, baseline law, validity criterion, observation resolution), not an absolute number.
  • Links

  • Code: https://github.com/zeroknowledgediscovery/tme
  • Data: https://doi.org/10.7910/DVN/F5TGT3
  • Entropy-rate estimates: https://github.com/zeroknowledgediscovery/nero
  • Paper: Chattopadhyay, I. *Thermodynamic Measure of Intelligence*. arXiv:2606.20231 (2026). Categories: cs.AI, cond-mat.stat-mech, cs.IT, math-ph, nlin.AO

Takeaway

Intelligence is not a single number but a position on a ruler whose graduations are given by thermodynamics — from a stone at Λ = 0 to Maxwell's demon at Λ ≈ 15, with GPT-5 (≈1.170) and human writers (≈1.196) nearly side by side, a gap of roughly 8x in raw lift under sentence-level English quality constraints. Other contexts (proving theorems, writing code) would shift the numbers, though the measurement method stays the same.

Tags

#thermodynamics#intelligence-measurement#maxwells-demon#large-language-models#recursive-self-simulation#information-theory#arxiv-paper#ai-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981595