> *"Film is not made by editing—it is what remains after editing."* — Jean-Luc Godard
This post explains AVA-Encoder: Towards Agent-Native Video Representation Learning (Li, Yu, Wang et al., 2026, arXiv:2608.12313), a paper arguing that creative AI must understand video structurally—not just perceptually.
The Problem: Pixels Without Cinema
Modern video AI can detect faces, estimate poses, and classify actions, but it does not understand *film language*. It cannot explain why Hitchcock's famous dolly zoom in *Vertigo* conveys vertigo, nor distinguish a "shot" from a "scene." The post frames this as a hierarchy of film understanding:
1. Pixels — RGB values, brightness, contrast (AI excels here) 2. Objects — people, cars, buildings (object detection handles this) 3. Actions — running, talking (action recognition works) 4. Shots — close-ups, pans, dolly moves (requires film knowledge) 5. Scenes — narrative units across shots (requires temporal/causal reasoning) 6. Narrative — plot, motivation, theme, symbolism (the film critic's domain)
Most AI models stop at levels 1–3. Creative AI—script-to-storyboard, style-conditioned generation, automated editing that respects emotional arcs—needs levels 4–6.
The Core Idea: Video ↔ Knowledge Graph ↔ Video
AVA-Encoder is a bidirectional system that encodes video into a knowledge graph and decodes the graph back into video. Its three components:
- Hierarchical nodes: scene nodes (macro descriptions, e.g., "rainy city street") and state nodes (specific states, e.g., "protagonist stands under an awning") storing structured text rather than pixels
- Asset layer: generative assets (images, audio, clips) linked to text nodes, produced by diffusion models—semantic equivalents, not raw copies
- Typed edges: explicit temporal, causal, and spatial relations that make the graph queryable and editable
- Outer loop (data-independent pseudo-training): the encoding policy is refined using *natural-language error feedback* instead of numeric gradients. Example: if the reconstructed video shows "a man jogging in a park" but the original was "a man strolling," the textual gradient reads like a director's note—"motion too brisk; should be strolling; reduce movement amplitude, add languid gait."
- Inner loop (optional, data-dependent): test-time refinement of the knowledge graph for specific videos.
- On a new Agentic Video Reconstruction Benchmark (node accuracy, edge accuracy, asset quality, semantic fidelity), AVA-Encoder outperforms the strongest external baseline by 20.7 percentage points.
- In controlled policy experiments, the pseudo-trained shot-level agentic encoder policy beats expert hand-tuned policies while using 74.3% fewer system prompt tokens.
- The structured representation supports querying ("find all scenes with rain"), editing ("change the café to a library"), reordering scenes, and regenerating video.
- Li, C., Yu, J., Wang, H., et al. (2026). *AVA-Encoder: Towards Agent-Native Video Representation Learning*. arXiv:2608.12313.
- Vaswani et al. (2017). *Attention Is All You Need*.
- Dosovitskiy et al. (2021). *An Image is Worth 16x16 Words*.
- Arnab et al. (2021). *ViViT: A Video Vision Transformer*.
- Rombach et al. (2022). *High-Resolution Image Synthesis with Latent Diffusion Models*.
- Brooks et al. (2024). *Video Generation Models as World Simulators*.
- Bordwell, D., & Thompson, K. (2010). *Film Art: An Introduction* (9th ed.).
The example in the post shows a woman entering a café, ordering a latte, and watching rain by the window, decomposed into scene/state nodes with linked assets and relations—structured, explicit, editable, and regenerable.
Textual-Gradient Optimization
Training a video→KG→video system is hard: the input/output are high-dimensional continuous data while the intermediate representation is discrete and structured, and dense annotation is impractical. AVA-Encoder uses a two-loop strategy:
Textual gradients provide high-level semantic feedback, human-readable debugging, and natural few-shot learning.
Results
Beyond Film
The architecture is general: educational videos (course → chapter → concept → example), sports (match → round → action → highlights/stats), surveillance, and medical imaging can all be represented as hierarchical structure + relations + generative assets. The post likens this to Gutenberg's printing press—turning video from a unique, continuous stream into structured, operable, comparable knowledge.
Future directions include cross-modal creation (novel → storyboard, script → animation), explainable video AI (the AI can justify editing decisions via the graph), and human–AI co-creation where directors manipulate narrative structure while AI handles shot-level generation, communicating through the graph as a shared intermediate language.