On August 6 (Beijing time), arXiv published a notably counter-trend paper: Activity Frames: Compiling Deterministic Pipelines for Agent Memory from Screen Activity, by independent researcher Nossa Iyamu. The paper also trended on HuggingFace Daily Papers.
The problem
The mainstream approach to long-term agent context is "have an LLM write a summary" — feed what happened on screen today into an LLM and stuff the resulting text into the prompt. This has clear drawbacks: summaries are not cacheable (rerunning the same window on the same day may produce different text), not auditable (to know what the agent actually saw, you must replay raw logs), and not reproducible.
The approach
Iyamu's answer is a zero-model, fully deterministic pipeline:
- Every screen transition (window focus change, URL navigation, scroll event) is compiled into an "event frame" by a set of rules.
- Three constant parameters (dwell 90s / session gap 300s / flicker merge 20s) merge event frames into "activity frames".
- An R-tree index stores all frames from a 24-hour window as a byte-consistent binary block.
- 86x compression: On the author's own 51 active days of data (128,756 frames), a full day of raw capture is compressed into a prompt block 86 times smaller. The paper itself notes this ratio mostly comes from "removing duplicate pixels + merging time windows", so it isn't directly comparable to LLM summarization — LLM summaries don't store raw pixels anyway.
- 68ms compilation: Running the full pipeline over a day's 128,756 frames takes 68ms. This is the most direct benefit of determinism + zero models: hot caching, offline recomputation, and CI integration all become possible.
- 98.4% QA accuracy: Read this carefully. The experiment is a self-designed QA task (100 questions over the 51-day screen log), where the agent answers using either compiled activity frames or raw logs against ground-truth answers. The baseline — GPT-4o-generated daily summaries — scores 66-80%. So on the author's own QA dataset, deterministic compression decisively beats LLM summarization. The conclusion's scope is narrow.
- arXiv abstract: https://arxiv.org/abs/2608.05784
- arXiv HTML full text: https://arxiv.org/html/2608.05784
- GitHub repo (implementation + sample data): https://github.com/nossa-iyamu/activity-frames
- PyPI package: https://pypi.org/project/activity-frames/
- Author's write-up (passions.com): https://www.passions.com/@nossa/55-days-of-screen-memory
- HuggingFace Daily Papers: https://huggingface.co/papers/2608.05784
- The evaluation dataset is author-constructed; 100 questions cover limited application scenarios
- No comparison against current SOTA RAG / long-context approaches (e.g., MemGPT, Letta, Anthropic Context Retrieval)
- Single person, single device; multi-device sync and cross-device deduplication unaddressed
- Three constants (dwell / session gap / flicker merge) are hardcoded; cross-user/cross-app generalizability unknown
- Semantic understanding of screen content (what the screen said, what decisions were made) still requires an LLM; Activity Frames only solves the "activity stream" dimension, not the "content stream"
- Privacy: local-only is the default, but cloud-edge collaboration boundaries are undiscussed
- No horizontal comparison with Apple Intelligence's screen-aware privacy mechanisms or OpenAI Operator's safety boundaries
Key numbers
Positioning
Unlike OpenAI's Operator, Anthropic's Computer Use, or Apple's ScreenKit — all "LLM looks at the screen" approaches — Activity Frames never lets the LLM see the screen; the LLM only reads the compiled activity frames. This makes the LLM-side cost and latency nearly zero, but the pipeline's model of "what happened on screen" is hand-crafted, and the author admits limited coverage of video frames, audio, and cross-window drag interactions.
The paper also includes a formal proof of cacheability: as long as the day's event sequence is unchanged, the activity frame output is byte-identical. That sounds trivial, but it doesn't hold for LLM summaries. Cachability means the same frame set can be shared across agents, persisted, and differentially updated.
Availability
The implementation is open-sourced: activity-frames on GitHub (https://github.com/nossa-iyamu/activity-frames), a same-named Python package on PyPI (v0.1.0, released 2026-08-06), with code, CLI, and compile.py all included. In a companion write-up the author discloses the discrepancy of "55 days of personal use vs 51 in the paper" and explains it (final cleanup days excluded from the measurement window).
Sources:
Assessment
This is a sample of "de-AI-ification" for agent long-term context engineering. The core thesis is not "LLM summaries aren't good enough" but "some tasks don't need an LLM at all" — compiling screen activity into agent memory can be done entirely by a deterministic pipeline, saving substantial LLM cost and latency. The author acknowledges limited generalizability: narrow QA scope, n=1 subject.