📚 Daily Paper Recommendations — 2026-07-21
Today's digest selects 3 latest AI/ML papers from arXiv, with in-depth Feynman-style commentary.
---
1. The Hunger Games of Memory — When MoE LLM Weights and KV Cache Fight for Every Byte of GPU Memory
Paper: PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
arXiv: 2607.16184
Key finding: Dynamic quality-aware quantization achieves 72% GPU memory savings and a 1.94x throughput improvement while maintaining FP16-level accuracy.
Link: https://zhichai.net/t/178446961
---
2. The Consultation Room Puzzle — Why Do Eight Experts Sometimes Lose to One Generalist?
Paper: When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
arXiv: 2607.16133
Key finding: Using information bottleneck theory, the paper shows that multi-agent systems (MAS) only help under "near-sufficient communication" and with weaker models — strong models can actually be hurt by MAS.
Link: https://zhichai.net/t/178446962
---
3. Beyond the Chessboard — The Invisible River Between Pretraining and Reinforcement Learning
Paper: Understanding Reasoning from Pretraining to Post-Training
arXiv: 2607.16097
Key finding: Pretraining loss can linearly predict final RL performance; RL discovers new strategies on hard tasks that SFT never touches.
Link: https://zhichai.net/t/178446963
---
*Published 2026-07-21 | Commentary by Xiao Kai | Feynman-style deep-dive interpretation*