AVA-Encoder: Towards Agent-Native Video Representation Learning
Field: Computer Vision Authors: Chuyue Li, Jinpeng Yu, Haozhe Wang arXiv: 2508.03420
Overview
Creative agents still lack an effective way to learn from high-quality human films, which limits their ability to produce cinematic-level videos. A key challenge is the absence of a structured video representation that is both faithful to cinematic content and directly usable for agent reasoning and manipulation.
To address this, the paper proposes the Agentic Video Auto-Encoder (AVA-Encoder), a framework that learns agent-native video representations via agentic auto-encoding.
Method
- AVA-Encoder converts a video into a knowledge graph (KG) representation, then reconstructs the video from it.
- The hierarchy and state nodes store structured text, while a linked asset layer holds generated images, audio, and video.
- Typed edges preserve the relationships between these text descriptions and assets in a form agents can easily understand, query, and edit.
- Video reconstruction differences drive a text-gradient optimization framework: evaluation feedback is expressed as natural-language update directions, used for:
- Outer loop: data-independent encoding strategy pseudo-training
- Inner loop: test-time refinement of the data-dependent KG representation
- AVA-Encoder outperforms the strongest external baseline by 20.7 percentage points.
- In a controlled policy-only setting, its pseudo-trained shot-level agent video encoder policy outperforms carefully hand-tuned policies while using 74.3% fewer system prompt tokens.
- The complete AVA-Encoder framework
- A reliable agent video reconstruction benchmark
- The first high-quality cinematic KG representation dataset
Results
Released Resources
The authors release:
*Collected automatically on 2026-08-14.*