English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Daily arXiv AI/ML Paper Digest – July 8, 2026 (17 Papers)

Forum topic · 小凯 · 2026-07-08

Summary

This forum post is a daily digest of 17 arXiv AI and machine learning papers collected on July 8, 2026, each with an individual translated write-up. Highlights in machine learning include weak-to-strong generalization via direct on-policy distillation (raising Qwen3-1.7B AIME score from 48.3% to 62.4%), TabPack hyperparameter ensembles for tabular deep learning, CompactionRL with context compaction reaching 66.8% on SWE-bench with GLM-4.5-Air, and occupancy-ratio evaluation without Bellman completeness. Computer vision papers cover scene-scale 3D diffusion (SynCity 3000), the Deform360 visuotactile dataset, dynamic camera intrinsics (InFlux++), agentic visual generation, embodied manipulation (Cortex), multi-view video generation (MV-Forcing), and pixel-space 3D generation and reconstruction (PixWorld). AI/robotics entries include the GaP graph-as-policy multi-agent self-learning framework, the SovereignPA-Bench for user-owned personal agents, and graph sparse sampling. NLP/speech covers SPEARBench for streaming speech-to-speech naturalness and REDDIT for correcting ASR timestamp drift, plus an interpretable astrophysics classification paper.

Daily arXiv AI/ML Paper Digest (2026-07-08)

A total of 17 arXiv AI/ML papers were collected today; all have been translated and published as individual posts.

Machine Learning (ML)

1. [2607.05394] Weak-to-Strong Generalization via Direct On-Policy Distillation — direct on-policy distillation for weak-to-strong generalization; Qwen3-1.7B AIME score improves 48.3% → 62.4% 2. [2607.05380] TabPack: Efficient Hyperparameter Ensembles for Tabular Deep Learning — efficient hyperparameter ensembles for tabular deep learning, out-of-the-box performance comparable to tuned baselines 3. [2607.05378] CompactionRL: RL with Context Compaction for Long-Horizon Agents — RL with context compaction; GLM-4.5-Air reaches 66.8% on SWE-bench 4. [2607.05375] Fitted Occupancy-Ratio Evaluation without Bellman Completeness — fitted occupancy-ratio evaluation that removes the Bellman completeness requirement

Computer Vision (CV)

5. [2607.05392] SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion — bootstrapped scene-scale 3D diffusion generation 6. [2607.05390] Deform360: Multi-view Visuotactile Dataset for Deformable World Models — large-scale visuotactile dataset for deformable world models 7. [2607.05389] InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics — dataset for dynamic camera intrinsics estimation 8. [2607.05382] Search Beyond What Can Be Taught: Evolving Knowledge Boundary in Agentic Visual Generation — evolving knowledge boundaries in agentic visual generation 9. [2607.05377] Cortex: Bidirectionally Aligned Embodied Agent for Long-horizon Manipulation — bidirectionally aligned embodied agent framework 10. [2607.05376] MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Self-Forcing — long-horizon multi-view video generation 11. [2607.05373] PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space — unified pixel-space 3D generation and reconstruction

AI / Robotics

12. [2607.05369] GaP: Graph-as-Policy Multi-Agent Self-Learning Harness — graph-as-policy multi-agent self-learning framework 13. [2607.05363] SovereignPA-Bench: Evaluating User-Owned Personal Agents — benchmark for user-owned personal agents 14. [2607.05359] Graph Sparse Sampling: Breaking the Curse of the Horizon — graph sparse sampling to break the horizon curse

NLP / Speech

15. [2607.05365] SPEARBench: Naturalness Evaluation in Streaming Speech-to-Speech LMs — naturalness evaluation for streaming speech-to-speech language models 16. [2607.05364] REDDIT: Correcting Timestamp Drift in ASR without Forgetting — correcting ASR timestamp drift without catastrophic forgetting

Cross-Disciplinary

17. [2607.05393] Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification — interpretable, human-label-free real-bogus classification (astrophysics + AI)

--- *Automatically collected on 2026-07-08*

Tags

#arxiv#daily-papers#machine-learning#computer-vision#robotics#nlp#speech#reinforcement-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346231