Daily arXiv AI/ML Paper Digest (2026-07-08)
A total of 17 arXiv AI/ML papers were collected today; all have been translated and published as individual posts.
Machine Learning (ML)
1. [2607.05394] Weak-to-Strong Generalization via Direct On-Policy Distillation — direct on-policy distillation for weak-to-strong generalization; Qwen3-1.7B AIME score improves 48.3% → 62.4% 2. [2607.05380] TabPack: Efficient Hyperparameter Ensembles for Tabular Deep Learning — efficient hyperparameter ensembles for tabular deep learning, out-of-the-box performance comparable to tuned baselines 3. [2607.05378] CompactionRL: RL with Context Compaction for Long-Horizon Agents — RL with context compaction; GLM-4.5-Air reaches 66.8% on SWE-bench 4. [2607.05375] Fitted Occupancy-Ratio Evaluation without Bellman Completeness — fitted occupancy-ratio evaluation that removes the Bellman completeness requirementComputer Vision (CV)
5. [2607.05392] SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion — bootstrapped scene-scale 3D diffusion generation 6. [2607.05390] Deform360: Multi-view Visuotactile Dataset for Deformable World Models — large-scale visuotactile dataset for deformable world models 7. [2607.05389] InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics — dataset for dynamic camera intrinsics estimation 8. [2607.05382] Search Beyond What Can Be Taught: Evolving Knowledge Boundary in Agentic Visual Generation — evolving knowledge boundaries in agentic visual generation 9. [2607.05377] Cortex: Bidirectionally Aligned Embodied Agent for Long-horizon Manipulation — bidirectionally aligned embodied agent framework 10. [2607.05376] MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Self-Forcing — long-horizon multi-view video generation 11. [2607.05373] PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space — unified pixel-space 3D generation and reconstructionAI / Robotics
12. [2607.05369] GaP: Graph-as-Policy Multi-Agent Self-Learning Harness — graph-as-policy multi-agent self-learning framework 13. [2607.05363] SovereignPA-Bench: Evaluating User-Owned Personal Agents — benchmark for user-owned personal agents 14. [2607.05359] Graph Sparse Sampling: Breaking the Curse of the Horizon — graph sparse sampling to break the horizon curseNLP / Speech
15. [2607.05365] SPEARBench: Naturalness Evaluation in Streaming Speech-to-Speech LMs — naturalness evaluation for streaming speech-to-speech language models 16. [2607.05364] REDDIT: Correcting Timestamp Drift in ASR without Forgetting — correcting ASR timestamp drift without catastrophic forgettingCross-Disciplinary
17. [2607.05393] Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification — interpretable, human-label-free real-bogus classification (astrophysics + AI)--- *Automatically collected on 2026-07-08*