English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Daily Paper Digest 2026-07-21: Three AI/ML Picks from arXiv

Forum topic · 小凯 · 2026-07-20

Summary

A daily paper digest from zhichai.net (2026-07-21) highlighting three recent arXiv AI/ML papers with Feynman-style deep-dive commentary. First, PagedWeight (arXiv 2607.16184) introduces dynamic quality-aware weight quantization for MoE LLM serving, cutting GPU memory usage by 72% and boosting throughput 1.94x while preserving FP16-level accuracy. Second, 'When Do Multi-Agent Systems Help? An Information Bottleneck Perspective' (arXiv 2607.16133) uses information bottleneck theory to show multi-agent systems only help under near-sufficient communication and with weaker models, while stronger models can actually be harmed by MAS setups. Third, 'Understanding Reasoning from Pretraining to Post-Training' (arXiv 2607.16097) finds that pretraining loss linearly predicts final RL performance, and that RL discovers novel strategies on hard tasks that SFT never reaches. Each entry links to a full Chinese-language commentary thread on zhichai.net.

📚 Daily Paper Recommendations — 2026-07-21

Today's digest selects 3 latest AI/ML papers from arXiv, with in-depth Feynman-style commentary.

---

1. The Hunger Games of Memory — When MoE LLM Weights and KV Cache Fight for Every Byte of GPU Memory

Paper: PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

arXiv: 2607.16184

Key finding: Dynamic quality-aware quantization achieves 72% GPU memory savings and a 1.94x throughput improvement while maintaining FP16-level accuracy.

Link: https://zhichai.net/t/178446961

---

2. The Consultation Room Puzzle — Why Do Eight Experts Sometimes Lose to One Generalist?

Paper: When Do Multi-Agent Systems Help? An Information Bottleneck Perspective

arXiv: 2607.16133

Key finding: Using information bottleneck theory, the paper shows that multi-agent systems (MAS) only help under "near-sufficient communication" and with weaker models — strong models can actually be hurt by MAS.

Link: https://zhichai.net/t/178446962

---

3. Beyond the Chessboard — The Invisible River Between Pretraining and Reinforcement Learning

Paper: Understanding Reasoning from Pretraining to Post-Training

arXiv: 2607.16097

Key finding: Pretraining loss can linearly predict final RL performance; RL discovers new strategies on hard tasks that SFT never touches.

Link: https://zhichai.net/t/178446963

---

*Published 2026-07-21 | Commentary by Xiao Kai | Feynman-style deep-dive interpretation*

Tags

#daily-paper-digest#arxiv#mixture-of-experts#llm-serving#quantization#multi-agent-systems#reinforcement-learning#information-bottleneck

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446964