English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AVA-Encoder: Agent-Native Video Representation Learning (arXiv 2508.03420)

Forum topic · 小凯 · 2026-08-14

Summary

AVA-Encoder (Agentic Video Auto-Encoder) is a framework for learning agent-native video representations, proposed by Chuyue Li, Jinpeng Yu, and Haozhe Wang (arXiv:2508.03420, computer vision). The method addresses a key limitation of creative agents: the lack of structured video representations that are both faithful to cinematic content and directly usable for agent reasoning and manipulation. AVA-Encoder converts videos into knowledge graph (KG) representations and then reconstructs them into video. Its hierarchy and state nodes store structured text, while a linked asset layer holds generated images, audio, and video; typed edges preserve relationships between text descriptions and assets in forms agents can easily understand, query, and edit. Video reconstruction differences drive a text-gradient optimization framework that expresses evaluation feedback as natural-language update directions for pseudo-training in the outer loop and test-time refinement of the data-dependent KG representation in the inner loop. Experiments show a 20.7 percentage point improvement over the strongest external baseline. In a controlled policy-only setting, its pseudo-trained shot-level agent video encoder policy outperforms carefully hand-tuned policies while reducing system prompt usage by 74.3%. The authors release the full framework, an agent video reconstruction benchmark, and the first high-quality cinematic KG dataset.

AVA-Encoder: Towards Agent-Native Video Representation Learning

Field: Computer Vision Authors: Chuyue Li, Jinpeng Yu, Haozhe Wang arXiv: 2508.03420

Overview

Creative agents still lack an effective way to learn from high-quality human films, which limits their ability to produce cinematic-level videos. A key challenge is the absence of a structured video representation that is both faithful to cinematic content and directly usable for agent reasoning and manipulation.

To address this, the paper proposes the Agentic Video Auto-Encoder (AVA-Encoder), a framework that learns agent-native video representations via agentic auto-encoding.

Method

  • AVA-Encoder converts a video into a knowledge graph (KG) representation, then reconstructs the video from it.
  • The hierarchy and state nodes store structured text, while a linked asset layer holds generated images, audio, and video.
  • Typed edges preserve the relationships between these text descriptions and assets in a form agents can easily understand, query, and edit.
  • Video reconstruction differences drive a text-gradient optimization framework: evaluation feedback is expressed as natural-language update directions, used for:
  • Outer loop: data-independent encoding strategy pseudo-training
  • Inner loop: test-time refinement of the data-dependent KG representation
  • Results

  • AVA-Encoder outperforms the strongest external baseline by 20.7 percentage points.
  • In a controlled policy-only setting, its pseudo-trained shot-level agent video encoder policy outperforms carefully hand-tuned policies while using 74.3% fewer system prompt tokens.
  • Released Resources

    The authors release:

  • The complete AVA-Encoder framework
  • A reliable agent video reconstruction benchmark
  • The first high-quality cinematic KG representation dataset
---

*Collected automatically on 2026-08-14.*

Tags

#computer-vision#video-representation#ai-agents#knowledge-graph#creative-ai#autoencoder#arxiv#video-generation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633452