Paper Overview
Field: Machine Learning Authors: Sean Augenstein, Li Ding, Jihwan Lee, Keith Rush, Andrey Zhmoginov Published: 2026-09-21 arXiv: 2609.24979
Summary
On-device large language models (LLMs), such as those running on mobile phones, are ripe for improvement via personalization. The limited compute resources of mobile devices impose limits on model scale and thus model quality, making any realizable quality gains highly impactful. At the same time, their personal nature (i.e., the close coupling to a particular user) means that a given on-device LLM tends to be used in similar, predictable patterns over time.
This paper presents a novel method for personalizing on-device LLMs: it trains a hypernetwork to map a user's context tokens to a low-rank adaptation (LoRA) well-suited to that user. Once the trained common artifacts are deployed to users' devices, each user uses the hypernetwork to synthesize a personalized LoRA entirely on device.
Key Advantages
The method combines the benefits of existing LLM customization approaches while avoiding their drawbacks:
- Like in-context learning (ICL) (unlike PEFT), the on-device stage is computationally feasible, requiring only a neural network forward pass.
- Like PEFT (unlike ICL), it modifies the target base LLM via weights (LoRA), avoiding negative effects of extending the input sequence, such as increased latency.
- Mobile-friendly: beyond the on-device compute and latency advantages, the architecture internally reuses parts of the target LLM's weights, so it requires almost no additional storage.
Evaluation
The authors validate the advantages of LoRA-generating hypernetworks on several representative personalization datasets, comparing against ICL, PEFT, and other baselines. Notably, the personalization experiments focus on the more challenging and less-studied task of long-text generation.
--- *Auto-collected on 2026-09-23*