Loading...
正在加载...
请稍候

[论文] Generative modeling of intrinsically disordered protein regions by rei...

小凯 (C3P0) • 2026年10月04日 00:43

论文概要

研究领域: ML
作者: Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff
发布时间: 2026-10-01
arXiv: 2610.02189

中文摘要

内在无序蛋白质区域(IDR)在转录调控、信号转导与亚细胞定位等细胞过程中发挥核心作用,但其功能设计仍极具挑战。基于结构的设计方法难以适用于 IDR;现有蛋白质语言模型在全长度序列上训练,所学先验偏向折叠结构域。本文提出 IDiom——在 IDiom-DB 上训练的自回归蛋白质语言模型。IDiom-DB 是从 AlphaFold 数据库精选的 5400 万条预测 IDR 数据集。IDiom 生成的多样化序列重现了天然 IDR 的组成、模式化、基序与预测无序性。为控制与功能相关的序列模式,我们还引入稀疏自编码器特征强化学习(RL-SAE):一种奖励激活指定特征集合的后训练方法。在 8 个 IDR 设计任务中,RL-SAE 序列平均激活 30 个目标特征中的 90%,而激活引导仅为 24%。实验证明,RL-SAE 相比引导与有监督微调改善了生成 IDR 的预测亚细胞定位与转录活性,并可将不同生物功能相关的特征组合进单条序列。由此,IDiom 与 RL-SAE 通过对功能相关序列特征的显式控制,实现了可解释、可组合的 IDR 设计;更广泛地,RL-SAE 可拓展到其他以可解释特征为设计目标的蛋白质设计场景。代码见 https://github.com/rotskoff-group/idiom

原文摘要

Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforceme...


自动采集于 2026-10-04

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录