[论文] Generative modeling of intrinsically disordered protein regions by rei...

研究领域: ML 作者: Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant…

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff 发布时间: 2026-10-01 arXiv: 2610.02189

中文摘要

内在无序蛋白质区域(IDR)在转录调控、信号转导与亚细胞定位等细胞过程中发挥核心作用,但其功能设计仍极具挑战。基于结构的设计方法难以适用于 IDR;现有蛋白质语言模型在全长度序列上训练,所学先验偏向折叠结构域。本文提出 IDiom——在 IDiom-DB 上训练的自回归蛋白质语言模型。IDiom-DB 是从 AlphaFold 数据库精选的 5400 万条预测 IDR 数据集。IDiom 生成的多样化序列重现了天然 IDR 的组成、模式化、基序与预测无序性。为控制与功能相关的序列模式,我们还引入稀疏自编码器特征强化学习(RL-SAE):一种奖励激活指定特征集合的后训练方法。在 8 个 IDR 设计任务中,RL-SAE 序列平均激活 30 个目标特征中的 90%,而激活引导仅为 24%。实验证明,RL-SAE 相比引导与有监督微调改善了生成 IDR 的预测亚细胞定位与转录活性,并可将不同生物功能相关的特征组合进单条序列。由此,IDiom 与 RL-SAE 通过对功能相关序列特征的显式控制,实现了可解释、可组合的 IDR 设计;更广泛地,RL-SAE 可拓展到其他以可解释特征为设计目标的蛋白质设计场景。代码见 https://github.com/rotskoff-group/idiom

原文摘要

Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforceme...


*自动采集于 2026-10-04*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens