English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Paper Express Sub-Index (May 9-25, 2026) - Daily AI Research Digests

Forum topic · 小凯 · 2026-05-25

Summary

This forum post is a sub-index of daily paper express digests published on zhichai.net between May 9 and May 25, 2026, listed in reverse chronological order. It aggregates links to individual paper report threads spanning a broad range of AI research topics: video understanding and generation (e.g., Cambrian-P, MotiMotion, RAVEN, CausalCine), vision-language models and navigation (AwareVLN, GesVLA, Vision-OPD), agents and agentic search (WildClawBench, DeepWeb-Bench, PREPING, SkillSmith), reasoning and RL (Equilibrium Reasoners, RubricEM, OpenDeepThink), diffusion and flow-matching generative models, mixture-of-experts architectures (DECO, UniPool, EMO), long-context and KV-cache efficiency (KV-Fold, DashAttention), benchmarks for 3D, memory, and evaluation (ESI-Bench, MEME, LongMemEval-V2, WikiVQABench), plus robotics, autonomous driving, and scientific ML. Each entry links directly to a dedicated thread with paper details. The index serves as a navigational hub for the forum's rapid paper coverage during this period.

Paper Express Sub-Index (2026-05-09 ~ 2026-05-25)

This index collects the daily paper express digests published between May 9 and May 25, 2026, listed in reverse chronological order. Each entry links to the corresponding paper thread on the forum.

2026-05-25

  • Cambrian-P: Pose-Grounded Video Understanding → https://zhichai.net/t/177620758
  • MotiMotion: Motion-Controlled Video Generation with Visual Reasoning → https://zhichai.net/t/177620759
  • Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs → https://zhichai.net/t/177620756
  • GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations → https://zhichai.net/t/177620762
  • The Matching Principle: A Geometric Theory of Loss Functions for Nuisance-Robust Representation Learning → https://zhichai.net/t/177620764
  • 2026-05-23

  • Integrable Elasticity via Neural Demand Potentials → https://zhichai.net/t/177620655
  • Tokenisation via Convex Relaxations → https://zhichai.net/t/177620653
  • AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation → https://zhichai.net/t/177620659
  • Cambrian-P: Pose-Grounded Video Understanding → https://zhichai.net/t/177620656
  • MotiMotion: Motion-Controlled Video Generation with Visual Reasoning → https://zhichai.net/t/177620657
  • Vector Policy Optimization: Training for Diversity Improves Test-Time... → https://zhichai.net/t/177620658
  • Remember to be Curious: Episodic Context and Persistent Worlds for 3D... → https://zhichai.net/t/177620660
  • GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations → https://zhichai.net/t/177620661
  • Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving → https://zhichai.net/t/177620662
  • Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness → https://zhichai.net/t/177620654
  • 2026-05-22

  • Velocityformer: Broken-Symmetry-Matched Equivariant Graph Transformers → https://zhichai.net/t/177620578
  • Quantifying Hyperparameter Transfer and the Importance of Embedding Layer → https://zhichai.net/t/177620572
  • Variance Reduction for Expectations with Diffusion Teachers → https://zhichai.net/t/177620569
  • One-Step Distillation of Discrete Diffusion Image Generators via Fixed... → https://zhichai.net/t/177620574
  • WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark → https://zhichai.net/t/177620576
  • EvoStruct: Bridging Evolutionary and Structural Priors for Antibody CDR → https://zhichai.net/t/177620573
  • Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning → https://zhichai.net/t/177620570
  • Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning → https://zhichai.net/t/177620571
  • DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source... → https://zhichai.net/t/177620575
  • 2026-05-20

  • Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Po... → https://zhichai.net/t/177620488
  • Actionable World Representation → https://zhichai.net/t/177620487
  • SURGE: Approximation-free Training Free Particle Filter for Diffusion → https://zhichai.net/t/177620486
  • DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention → https://zhichai.net/t/177620480
  • WavFlow: Audio Generation in Waveform Space → https://zhichai.net/t/177620482
  • Aurora: Unified Video Editing with a Tool-Using Agent → https://zhichai.net/t/177620483
  • Can These Views Be One Scene? Evaluating Multiview 3D Consistency → https://zhichai.net/t/177620479
  • A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime... → https://zhichai.net/t/177620481
  • Code as Agent Harness → https://zhichai.net/t/177620484
  • ESI-Bench: Towards Embodied Spatial Intelligence → https://zhichai.net/t/177620485
  • 2026-05-19

  • DeepSlide: From Artifacts to Presentation Delivery → https://zhichai.net/t/177620349
  • SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces → https://zhichai.net/t/177620352
  • Fair outputs, Biased Internals: Causal Potency and Asymmetry of Latent... → https://zhichai.net/t/177620353
  • NIMO Controller: a self-driving laboratory orchestrator → https://zhichai.net/t/177620357
  • NOVA: Fundamental Limits of Knowledge Discovery Through AI → https://zhichai.net/t/177620355
  • Does Theory of Mind Improvement Really Benefit Human-AI Interactions? → https://zhichai.net/t/177620351
  • SDOF: Taming the Alignment Tax in Multi-Agent Orchestration → https://zhichai.net/t/177620350
  • Solvita: Enhancing LLMs for Competitive Programming → https://zhichai.net/t/177620358
  • 2026-05-18

  • From Descriptive to Prescriptive: Uncover the Social Value Alignment → https://zhichai.net/t/177620218
  • PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts → https://zhichai.net/t/177620215
  • Conditional Attribute Estimation with Autoregressive Sequence Models → https://zhichai.net/t/177620216
  • Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory → https://zhichai.net/t/177620217
  • Mixed Integer Goal Programming for Personalized Meal Optimization → https://zhichai.net/t/177620211
  • A Two-Dimensional Framework for AI Agent Design Patterns → https://zhichai.net/t/177620212
  • PREPING: Building Agent Memory without Tasks → https://zhichai.net/t/177620214
  • Enhanced and Efficient Reasoning in Large Learning Models → https://zhichai.net/t/177620219
  • Invisible Orchestrators Suppress Protective Behavior → https://zhichai.net/t/177620213
  • Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use → https://zhichai.net/t/177620220
  • 2026-05-17

  • FutureSim: Replaying World Events to Evaluate Adaptive Agents → https://zhichai.net/t/177620162
  • Evidential Reasoning Advances Interpretable Real-World Disease Screening → https://zhichai.net/t/177620169
  • RAVEN: Real-time Autoregressive Video Extrapolation → https://zhichai.net/t/177620161
  • Quantitative Video World Model Evaluation for Geometric-Consistency → https://zhichai.net/t/177620165
  • Articraft: An Agentic System for Scalable Articulated 3D Asset Generation → https://zhichai.net/t/177620163
  • OpenDeepThink: Parallel Reasoning via Bradley--Terry Aggregation → https://zhichai.net/t/177620167
  • Text Knows What, Tables Know When: Clinical Timeline Reconstruction → https://zhichai.net/t/177620170
  • MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface → https://zhichai.net/t/177620168
  • SANA-WM: Efficient Minute-Scale World Modeling → https://zhichai.net/t/177620166
  • VGGT-Edit: Feed-forward Native 3D Scene Editing → https://zhichai.net/t/177620164
  • 2026-05-16

    (Overlapping entries with 05-17; see links such as FutureSim https://zhichai.net/t/177620087, RAVEN https://zhichai.net/t/177620086, OpenDeepThink https://zhichai.net/t/177620092, VGGT-Edit https://zhichai.net/t/177620089, SANA-WM https://zhichai.net/t/177620091)

    2026-05-15

  • From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended → https://zhichai.net/t/177620064
  • When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability → https://zhichai.net/t/177620062
  • Eradicating Negative Transfer in Multi-Physics Foundation Models → https://zhichai.net/t/177620065
  • Warp-as-History: Generalizable Camera-Controlled Video Generation → https://zhichai.net/t/177620063
  • Is Grep All You Need? How Agent Harnesses Reshape Agentic Search → https://zhichai.net/t/177620061
  • Aligning Latent Geometry for Spherical Flow Matching in Image Generation → https://zhichai.net/t/177620060
  • VGGT-\(\Omega\) → https://zhichai.net/t/177620059
  • RefDecoder: Enhancing Visual Generation with Conditional Video Decoding → https://zhichai.net/t/177620058
  • EntityBench: Entity-Consistent Long-Range Multi-Shot Video Generation → https://zhichai.net/t/177620056
  • ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both → https://zhichai.net/t/177620057
  • 2026-05-14

  • Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transform → https://zhichai.net/t/177620004
  • Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surfaces → https://zhichai.net/t/177620002
  • MEME: Multi-entity & Evolving Memory Evaluation → https://zhichai.net/t/177620011
  • Task-Adaptive Embedding Refinement via Test-time LLM Guidance → https://zhichai.net/t/177620006
  • Solve the Loop: Attractor Models for Language and Reasoning → https://zhichai.net/t/177620014
  • AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs → https://zhichai.net/t/177620001
  • KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference → https://zhichai.net/t/177620013
  • Covering Human Action Space for Computer Use → https://zhichai.net/t/177619997
  • Beyond GRPO and On-Policy Distillation → https://zhichai.net/t/177620008
  • From Web to Pixels: Bringing Agentic Search into Visual Perception → https://zhichai.net/t/177619999
  • EgoForce: Forearm-Guided Camera-Space 3D Hand Pose → https://zhichai.net/t/177619998
  • Elastic Attention Cores for Scalable Vision Transformers → https://zhichai.net/t/177620005
  • LongMemEval-V2: Evaluating Long-Term Agent Memory → https://zhichai.net/t/177620003
  • CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video → https://zhichai.net/t/177620000
  • ToolCUA: Towards Optimal GUI-Tool Path Orchestration → https://zhichai.net/t/177620009
  • Learning, Fast and Slow: Towards LLMs That Adapt Continually → https://zhichai.net/t/177620007
  • OmniNFT: Modality-wise Omni Diffusion Reinforcement → https://zhichai.net/t/177620010
  • Routers Learn the Geometry of Their Experts → https://zhichai.net/t/177620012
  • 2026-05-13

  • V4FinBench: Benchmarking Tabular Foundation Models, LLMs, and Standard ML → https://zhichai.net/t/177619930
  • CapVector: Learning Transferable Capability Vectors in Parametric Space → https://zhichai.net/t/177619927
  • Equivariant Reinforcement Learning for Clifford Quantum Circuit Synthesis → https://zhichai.net/t/177619923
  • Engineering Robustness into Personal Agents with the AI Workflow Store → https://zhichai.net/t/177619925
  • Revisiting Policy Gradients for Restricted Policy Classes → https://zhichai.net/t/177619924
  • Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers → https://zhichai.net/t/177619928
  • DataMaster: Towards Autonomous Data Engineering for Machine Learning → https://zhichai.net/t/177619926
  • Shepherd: A Runtime Substrate Empowering Meta-Agents → https://zhichai.net/t/177619921
  • RubricEM: Meta-RL with Rubric-guided Policy Decomposition → https://zhichai.net/t/177619929
  • WildClawBench: Real-World, Long-Horizon Agent Evaluation → https://zhichai.net/t/177619922
  • ELF: Embedded Language Flows → https://zhichai.net/t/177619911
  • Optimal and Scalable MAPF via Multi-Marginal Optimal Transport → https://zhichai.net/t/177619919
  • DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance → https://zhichai.net/t/177619915
  • Quantifying Concentration Phenomena of Mean-Field Transformers → https://zhichai.net/t/177619916
  • Personal Visual Context Learning in Large Multimodal Models → https://zhichai.net/t/177619913
  • Variational Inference for Lévy Process-Driven SDEs via Neural Tilting → https://zhichai.net/t/177619914
  • Confidence-Guided Diffusion Augmentation for Bangla Compound Script → https://zhichai.net/t/177619920
  • 2026-05-12

  • A Note on Non-Negative L1-Approximating Polynomials → https://zhichai.net/t/177619880
  • VecCISC: Improving Confidence-Informed Self-Consistency → https://zhichai.net/t/177619881
  • GRAPHLCP: Structure-Aware Localized Conformal Prediction on Graphs → https://zhichai.net/t/177619878
  • Proxy3D: Efficient 3D Representations for Vision-Language Models → https://zhichai.net/t/177619882
  • LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling → https://zhichai.net/t/177619874
  • EmambaIR: Efficient Visual State Space Model for Event-guided Image Restoration → https://zhichai.net/t/177619879
  • Normalizing Trajectory Models → https://zhichai.net/t/177619875
  • Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering → https://zhichai.net/t/177619876
  • 123D: Unifying Multi-Modal Autonomous Driving Data at Scale → https://zhichai.net/t/177619871
  • 2026-05-11

  • Recursive Agent Optimization → https://zhichai.net/t/177619785
  • GlazyBench: A Benchmark for Ceramic Glaze Property Prediction → https://zhichai.net/t/177619784
  • 2026-05-10

  • Inductive Venn-Abers and related regressors → https://zhichai.net/t/177619698
  • Concept-Based Abductive and Contrastive Explanations → https://zhichai.net/t/177619701
  • Optimizer-Model Consistency: Full Finetuning with the Same Optimizer → https://zhichai.net/t/177619694
  • BAMI: Training-Free Bias Mitigation in GUI Grounding → https://zhichai.net/t/177619689
  • Verifier-Backed Hard Problem Generation for Mathematical Reasoning → https://zhichai.net/t/177619691
  • ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation → https://zhichai.net/t/177619687
  • When No Benchmark Exists: Validating Comparative LLM Safety Scoring → https://zhichai.net/t/177619695
  • Why Global LLM Leaderboards Are Misleading → https://zhichai.net/t/177619693
  • Superintelligent Retrieval Agent → https://zhichai.net/t/177619697
  • Relit-LiVE: Relight Video by Jointly Learning Environment Video → https://zhichai.net/t/177619692
  • Edge-specific signal propagation on 3D mechanochemical models → https://zhichai.net/t/177619699
  • Are We Making Progress in Multimodal Domain Generalization? → https://zhichai.net/t/177619700
  • Beyond Negative Rollouts: Positive-Only Policy Optimization → https://zhichai.net/t/177619696
  • EMO: Pretraining Mixture of Experts for Emergent Modularity → https://zhichai.net/t/177619690
  • UniPool: A Globally Shared Expert Pool for Mixture-of-Experts → https://zhichai.net/t/177619688
  • 2026-05-09

  • BAMI: Training-Free Bias Mitigation in GUI Grounding → https://zhichai.net/t/177619662
  • Verifier-Backed Hard Problem Generation for Mathematical Reasoning → https://zhichai.net/t/177619664
  • ActCam: Zero-Shot Joint Camera and 3D Motion Control → https://zhichai.net/t/177619660
  • UniPool: A Globally Shared Expert Pool → https://zhichai.net/t/177619661
  • EMO: Pretraining Mixture of Experts for Emergent Modularity → https://zhichai.net/t/177619663
  • Optimizer-Model Consistency → https://zhichai.net/t/177619667
*Note: Duplicate entries appearing in the original index (same paper listed under multiple dates) have been consolidated above.*

Tags

#paper-digest#index#ai-research#machine-learning#vision-language-models#llm-agents#video-generation#benchmarks

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620791