x64 Unbroken: A 48-Year War Over Compatibility
*Translation summary of a Chinese tech forum deep-dive (zhichai.net), based on a four-expert panel covering microarchitecture, software ecosystem economics, industry history, and 2026 frontier data. Original data as of 2026 Q2.*
Two contradictory narratives dominate the x86-64 vs ARM debate. One says ARM has already won: Apple's M-series has hammered Intel for six years, Snapdragon X2 Elite beats Intel's flagship by 24% in Geekbench, and ARM servers take nearly half of revenue share. The other says x86 stands unshaken: in 2026 Q2, Intel + AMD still dominate PC client CPU shipments, Windows on ARM (WoA) hovers around 3–5% penetration, and ARM Windows devices are exactly 0.00% of the Steam hardware survey.
Both narratives use real data. The contradiction dissolves once you stop crediting Apple's victories to ARM.
Key points
1. The headline ARM number is mostly Apple
- ARM's share of the entire PC client CPU market: 15.3% (Mercury Research, 2026 Q2, all-time high), corroborated by Counterpoint (~15%).
- Subtracting Mac share (~10–11%), Windows on ARM is only 4–5% of overall Windows laptop shipments; TrendForce's independent figure is 3.2%.
- US premium retail (>$800) WoA share exceeds 10% (Circana) — while global all-price-band shipments are only 4–5%. Both facts together prove WoA is concentrated in high-end retail and has not penetrated commercial bulk purchasing.
- TrendForce forecasts ARM notebooks at 34.2% by 2029, but Windows on ARM only 11.5% — Apple takes half of that third.
- The "x86 is CISC, so decoding is slow and power-hungry" claim is outdated. The μop cache is now the main path (~80% hit rate since Sandy Bridge, near 100% for hot loops); decoding costs only 3–10% of package power in the worst case (Hirki et al., USENIX CoolDC 2016). The paper concludes that switching ISAs saves little power because the decoder cannot be eliminated.
- ARM is not clean either: Cortex-A77 and later add 1.5K-entry μop caches; Fujitsu's A64FX splits one SVE
FADDAinstruction into 63 μops; Marvell ThunderX3's biggest single gain was reducing μop expansion. - Decisive natural experiment: AMD's Zen 5c vs Zen 5 — same ISA, same IPC, same decoders, yet 25% less single-core area and 2× core density. Design choices, not ISA, determine efficiency.
- Vector width: Qualcomm's Oryon uses 128-bit NEON with 4 pipelines (aggregated 512 bits/cycle, same peak class as Zen 5's native AVX-512). The real tax is instruction expansion: one AVX-512 instruction becomes 4 NEON instructions and 4 ROB entries, erasing Oryon's ~45% larger ROB advantage. This was a voluntary choice — Qualcomm's Karl Whealton confirmed Oryon stays on Armv8 without SVE/SME for battery-life reasons.
- Memory model — the real lock cylinder: x86 guarantees TSO as an ISA contract; ARM is weakly ordered. Compilers strip barriers from x86 binaries, so emulators must guess. Apple's M-series and Qualcomm's Oryon include a hardware TSO mode (Oryon came from Nuvia, founded by ex-Apple chief CPU architect Gerard Williams III). But Arm public Cortex cores have no TSO, so Microsoft's Prism must retain software fallbacks — including a "force single-core operation" option when even barriers aren't enough.
- Microsoft could never raise the hardware floor (e.g., require TSO) the way it once raised the 386→486 requirement: back then only software was orphaned (and it all still ran); today, already-sold WoA hardware would be bricked. With the Qualcomm exclusivity expired and NVIDIA/AMD/MediaTek entering, hardware diversity is peaking — making a floor raise least likely ever. More diversity, harder to set a contract. A negative feedback loop, not a transition problem.
- Intel, fearing x86-64 would cannibalize Itanium, deliberately delayed extending x86 (per Dileep Bhandarkar). IDC predicted $28B in Itanium sales by 2004; actual mid-2004 revenue was ~$606M — under 2.2% of the forecast. By 2008, 95% of Itanium sales came from HP, which paid Intel ~$690M to keep the line alive. Itanium only turned profitable at the end of 2009.
- Itanium died from a two-front squeeze: SIMD took throughput workloads, out-of-order execution took latency-sensitive integer work (SPECint loss, SPECfp win). AMD64 killed its 64-bit selling point in 2003 while preserving x86 compatibility. Either blow alone was survivable; together they were fatal.
- The asymmetry: on x86, raising the compatibility floor is free (fully upward compatible); on ARM, it is lethal (hardware fragmentation). x86's contract is enforced by the Intel–AMD cross-license duopoly; ARM's "contract" is a menu Arm has neither mechanism nor incentive to make mandatory.
> ARM's PC victory is Apple's victory, not ARM's. Apple's key is "vertical integration" — and Windows' door has no keyhole at all.
2. The technology layer is the thinnest layer
3. Two genuine ISA differences remain — both are design trade-offs
4. History: the price of abandoning compatibility
5. The verdict
Compatibility — a guaranteed, universal, multi-decade binary contract — is the moat. ARM can win benchmarks and even markets where one vendor controls the whole stack (Apple), but it cannot replace x64 where millions of legacy binaries and fragmented hardware must coexist. Technical problems end in contract problems, and contracts are what x64 has had for 48 years.*Note: figures, quotes, and forecasts are reproduced from the original post and its cited sources; the author flags one emulation μop-inflation estimate (2.6–4×) as a rough calculation, not a citable measurement.*