AI消息速览

MetaboLLM:代谢组学专用大语言模型与预测性代谢物图构建

事件日期 2026-08-06 · 学术前沿 · 已接受

事件日期2026-08-06
信息日期2026-08-06
入库日期2026-08-08
通道学术前沿
状态已接受
来源arXiv 论文

知识卡片:MetaboLLM:代谢组学专用大语言模型与预测性代谢物图构建

一句话结论

MetaboLLM 通过持续预训练、监督微调和结构化检索,将代谢组学异构知识转化为可预测的代谢物图(MetaboLLM-GIN),在两个临床预测任务上取得最高 AUC,并优于未适配或未检索的大语言模型配置。

事件概述或研究问题

  • 研究问题:代谢组学知识分散在异构资源中,难以转化为可用于患者级预测的表征。
  • 事件/贡献:本文提出 MetaboLLM(代谢组学专用大语言模型),以及将其生成的生化描述转化为代谢物图的 MetaboLLM-GIN,用于可解释的患者级预测。

方法/产品要点

  • MetaboLLM:基于多个大语言模型骨干,通过持续预训练(continual pretraining)、监督微调(supervised fine-tuning)和结构化检索(structured retrieval)进行代谢组学领域适配。
  • MetaboLLM-GIN:将 LLM 生成的代谢物生化描述转化为代谢物图,再使用图同构网络(Graph Isomorphism Network, GIN)进行患者级预测。
  • 评估覆盖四个骨干家族;对比对象包括对应骨干的基础模型和医学适配模型。

主要结果或产业意义

  • MetaboLLM 在代谢组学知识、关系和描述任务上优于对应的基础模型和医学适配模型,并能迁移到外部公共基准。
  • MetaboLLM-GIN 在冠状动脉旁路移植术后应激性高血糖预测中取得最高 AUC = 0.8616,在绝经后激素方案分类中取得最高 AUC = 0.8123,优于传统模型、替代图构建方法以及未适配/无检索 LLM 生成的图。
  • 模型解释在两个应用中都产生了具有生物学意义的发现。

为什么重要

  • 展示了领域专用大语言模型可以将异构生化知识组织为预测性和可解释的代谢物图表示。
  • 与已有“基于生物标志物的知识图谱用于可解释诊断”的脉络相比,MetaboLLM 的增量在于:不是人工设计知识图谱,而是由大语言模型自动生成代谢物图结构,并通过 GIN 同时实现预测与解释,为代谢组学与临床决策之间搭建了更自动化的桥梁。

局限与不确定性

  • 摘要未提供具体数据集规模、患者队列来源、外部公共基准名称,待核实。
  • 摘要未说明四个骨干家族具体是哪些模型、参数规模、训练语料与计算成本,待核实。
  • 代谢物图的节点/边定义、图构建的提示模板,以及结构化检索的具体策略,摘要中未展开,待核实。
  • AUC 提升的统计学显著性和临床验证程度,摘要未提供细节,待核实。

可用于图书/PPT/简报的角度

  • 示例标题:让大语言模型“读懂”代谢物:从海量生化知识到可解释的临床预测图。
  • 可强调:大语言模型不仅能回答问题,还能将分散的代谢组学知识自动转化为图结构,在预测血糖异常和激素治疗分类中超越传统模型。
  • 可结合“领域专用 vs 通用基础模型”的讨论,说明持续预训练和检索对代谢组学任务的价值。

原始材料

  • 英文标题:MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction
  • 英文关键词:metabolomics; large language model; metabolite graph; graph isomorphism network; biochemical knowledge integration
  • 来源:arXiv:2608.06253v1 [cs.LG]
  • URL: https://arxiv.org/abs/2608.06253v1
  • Authors: Dohyun Ku, Min Gu Kwak, Francisco J. Pasquel, Jing Li
  • Published: 2026-08-06
  • Abstract 原文:Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretraining, supervised fine-tuning, and structured retrieval, together with MetaboLLM-GIN, which converts generated biochemical descriptions into metabolite graphs for patient-level prediction using a graph isomorphism network. Across four backbone families, MetaboLLM outperformed corresponding base and medically adapted models on metabolomics knowledge, relational, and description tasks, and transferred to an external public benchmark. MetaboLLM-GIN achieved the highest AUC for stress hyperglycemia prediction after coronary artery bypass grafting (0.8616) and postmenopausal hormone-regimen classification (0.8123), outperforming conventional models, alternative graph constructions, and graphs generated from unadapted or non-retrieval LLM configurations. Model interpretation further produced biologically meaningful findings in both applications. These results show that domain-specialized language models can organize heterogeneous biochemical knowledge into predictive and interpretable metabolite graph representations.