English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

arXiv AI/ML Paper Digest - July 14, 2026: 20 New Papers on Agents, Video Diffusion, Robotics, and More

Forum topic · 小凯 · 2026-07-16

Summary

This digest compiles 20 recent AI and machine learning papers posted to arXiv on July 14, 2026, curated by zhichai.net. Highlights include E3, a complexity-aware LLM agent method that cuts cost by 85% while keeping 100% task success on MSE-Bench; a study revealing a seriality gap in video diffusion models when modeling causal chains; TerraZero, a procedural driving simulator sustaining 1.3M agent steps per second on a single GPU and topping the InterPlan benchmark; PalmClaw, a native on-device mobile agent framework; GyroFlow, a flow-matching model that skips transient phases in turbulence simulation; FlowWAM, using optical flow as a unified action representation reaching 92.94% success on RoboTwin; DiffusionGemma-based parallel speech transcription with 6.6% WER on LibriSpeech; DermDepth for monocular metric-scale dermatology 3D reconstruction; and ViCo3D for LiDAR-based collaborative 3D detection with vision foundation models. Each entry lists authors, arXiv links, and key quantitative results across agents, robotics, computer vision, speech, and AI safety.

arXiv AI/ML Paper Digest (2026-07-14)

Auto-collected on 2026-07-16. 20 latest papers, summarized below.

Key points

1. Do AI Agents Know When a Task Is Simple? (arXiv:2607.13034) - Junjie Yin, Xinyu Feng. Introduces task-aware execution scope estimation, formalizing minimal sufficient execution and the Agent Cognitive Redundancy Ratio (ACRR). The E3 (Estimate, Execute, Expand) method estimates an initial operating point, executes the minimal viable path, and expands only on validation failure. On MSE-Bench, E3 cuts cost by 85%, tokens by 91%, and files inspected by 92% while keeping 100% success.

2. The Seriality Gap in Video Diffusion Models (arXiv:2607.13031) - Jorge Diaz Chao et al. Controlled experiments on multi-ball hard-sphere dynamics show bidirectional video diffusion degrades as causal chains lengthen; performance loss disappears in length-matched single-ball controls. More denoising steps do not add serial computation beyond the backbone, indicating a structural barrier for serial reasoning in video diffusion.

3. TerraZero: Procedural Driving Simulation (arXiv:2607.13028) - Zhouchonghao Wu et al. A procedural driving simulator and self-play stack sustaining 1.3 million agent steps per second on one server-grade GPU (simulation on CPU, policy inference on GPU). First fully learned policy to top the InterPlan long-tail benchmark.

4. PalmClaw: Native On-Device Mobile Agent (arXiv:2607.13027) - Hongru Cai et al. An open-source framework running natively on phones, exposing device capabilities as tools with explicit parameters and structured results. +11.5% relative task success, -94.9% completion time.

5. Shortcut to Steady-State Turbulence with Flow Matching (arXiv:2607.13022) - Gianluca Galletti et al. GyroFlow, a latent generative model that directly estimates gyrokinetic turbulence steady-state statistics in 5D phase space, bypassing transient dynamics; outperforms autoregressive and reduced-order baselines.

6. FlowWAM: Optical Flow as Unified Action Representation (arXiv:2607.13017) - Yixiang Chen et al. A two-stream diffusion framework using optical flow for policy (action prediction) and world-model (video generation guided by target flow) modes. 92.94% success on RoboTwin manipulation, beating VLA and WAM baselines.

7. Audio-Native Speech Recognition with Frozen Discrete-Diffusion LM (arXiv:2607.13013) - Harsha Vardhan Khurdula et al. Trains an audio interface for DiffusionGemma (26B MoE), tuning only ~42M parameters (0.16% of the backbone). 6.6% WER on LibriSpeech test-clean with ~8 parallel denoising steps.

8. DermDepth: Monocular Metric 3D for Dermatology (arXiv:2607.13010) - Héctor Carrión, Narges Norouzi. First single-view metric-scale 3D model for dermatology plus D-Synth, the first synthetic dermatoscope dataset with pixel-perfect 3D. Reduces metric-scale error from >16x to below 1.1x.

9. Dynamic Resource Allocation for Ensemble Determinization MCTS (arXiv:2607.13007) - Jakub Kowalski et al. Dynamic determinization counts and non-uniform simulation allocation yield statistically significant strength gains in Jaipur, Lost Cities, and Splendor.

10. The Spectrum Is Not Enough: When Context Helps Forecasting (arXiv:2607.13006) - Mert Onur Cakiroglu et al. Spectral predictability scores do not answer whether extra context helps; proposes a label-free coverage-deficit diagnostic validated across seven benchmarks.

11. Watermark Forensics: An Information-Theoretic Perspective (arXiv:2607.13003) - Xiaoyu Li et al. For statistically distortion-free schemes, attributing text among N users costs Θ(log N / h) tokens - the first tight entropy-rate law for multi-user watermark attribution.

12. X-Lens: Real-Time Metric Depth from Heterogeneous Cameras (arXiv:2607.12993) - Heng Zhou et al. 0.04B-parameter feed-forward model at up to 41 FPS; 25.4% AbsRel reduction on OmniScene-Full with 88.9% fewer parameters.

13. Controllable Generation of Diverse Dermatological Imagery (cgDDI) (arXiv:2607.12987) - Héctor Carrión, Narges Norouzi. Hybrid framework synthesizing realistic healthy skin, mapping rare lesions across skin tones, and generating from just 10 training samples. 86.4% malignancy accuracy from synthetic-only training; 90.9% state of the art after fine-tuning on real data (DDI benchmark).

14. Win by Silence: LLM Plan Evaluation Non-Monotonicity (arXiv:2607.12986) - Aleh Manchuliantsau. Shows plan evaluators can reward plans for becoming less explicit; a typed-state GATE gatekeeper rejected score release on 26/26 silenced routes, with 47/54 post-rejection revisions fixed to coverage structure.

15. Resist and Update: Counterfactual Report Coordinates (arXiv:2607.12985) - Sen Yang, Yuen-Hei Yeung. A counterfactual report mediator constrains model reports to a causal contract: invariant to non-evidential pressure, responsive to genuine evidence. Achieves 1.00 resist and update on the witness benchmark (Wilson 95% CI [0.99, 1.00]).

16. FormalAnalyticGeo: Neural-Symbolic Geometry Problem Generation (arXiv:2607.12982) - Ruoran Xu et al. A scalable framework built on CDL and an SDF engine with four sequential LLM components (generator, formalizer, measurer, quality verifier), producing AnalyticGeo7K, a dataset of 7K+ validated multimodal problems.

17. Ensemble Controlled-Flow Filtering for Implicit Data Assimilation (arXiv:2607.12975) - Zhuoyuan Li et al. The EnCF filter via stochastic controlled flows suits non-Gaussian, many-to-one, multimodal, and implicit observation models, where Kalman-type filters fall short.

18. The Illusion of Robustness: Prediction Flips under Irrelevant Context (arXiv:2607.12963) - Yanzhe Zhang et al. Aggregate accuracy hides per-example instability: even semantically meaningless pseudowords can significantly flip predictions on a subset of examples, revealing context-induced tail risk.

19. Placebo-Controlled Evaluation of Learned Self-Repair in Small Code Models (arXiv:2607.12962) - Mehmet Iscan. Introduces PoPE (Popperian Placebo-controlled Evaluation). In the prompt channel, form placebos unlocked 12 units vs 10 for live error patterns; in the weight channel, error-content adapters tied the no-intervention baseline 8-8 (p=1.0).

20. ViCo3D: LiDAR Collaborative 3D Detection with Vision Foundation Models (arXiv:2607.12959) - Haojie Ren et al. Projects point clouds to BEV images so DINOv2 extracts semantically rich features for V2X collaboration; collaboration gains up to 1.8x over prior methods on DAIR-V2X.

---

*Auto-collected on 2026-07-16.*

Tags

#arxiv#ai#machine-learning#llm-agents#video-diffusion#robotics#computer-vision#research-digest

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178395177