论文概要
研究领域: ML
作者: Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff
发布时间: 2026-10-01
arXiv: 2610.02189
中文摘要
内在无序蛋白质区域(IDR)在转录调控、信号转导与亚细胞定位等细胞过程中发挥核心作用,但其功能设计仍极具挑战。基于结构的设计方法难以适用于 IDR;现有蛋白质语言模型在全长度序列上训练,所学先验偏向折叠结构域。本文提出 IDiom——在 IDiom-DB 上训练的自回归蛋白质语言模型。IDiom-DB 是从 AlphaFold 数据库精选的 5400 万条预测 IDR 数据集。IDiom 生成的多样化序列重现了天然 IDR 的组成、模式化、基序与预测无序性。为控制与功能相关的序列模式,我们还引入稀疏自编码器特征强化学习(RL-SAE):一种奖励激活指定特征集合的后训练方法。在 8 个 IDR 设计任务中,RL-SAE 序列平均激活 30 个目标特征中的 90%,而激活引导仅为 24%。实验证明,RL-SAE 相比引导与有监督微调改善了生成 IDR 的预测亚细胞定位与转录活性,并可将不同生物功能相关的特征组合进单条序列。由此,IDiom 与 RL-SAE 通过对功能相关序列特征的显式控制,实现了可解释、可组合的 IDR 设计;更广泛地,RL-SAE 可拓展到其他以可解释特征为设计目标的蛋白质设计场景。代码见 https://github.com/rotskoff-group/idiom
原文摘要
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforceme...
自动采集于 2026-10-04
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。