<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Zhixiang Ren · Recent Papers</title>
    <link>https://zhixiang-ren.github.io/</link>
    <description>AI scientist working across AI for Science, drug discovery, bioinformatics, multimodal foundation models, and AI agents.</description>
    <language>en</language>
    <atom:link href="https://zhixiang-ren.github.io/rss.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>The limits of bio-molecular modeling with large language models: a cross-scale evaluation</title>
      <link>https://doi.org/10.1093/bioinformatics/btag550</link>
      <guid isPermaLink="false">https://doi.org/10.1093/bioinformatics/btag550</guid>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
      <description>Yaxin Xu, Yue Zhou, Tianyu Zhao, Zhengyu Ma, Fengwei An, Zhixiang Ren

Bioinformatics · Published · 2026

Abstract Motivation The modeling of bio-molecular system across molecular scales remains a central challenge in scientific research. Large language models (LLMs) are increasingly applied to bio-molecular discovery, yet systematic evaluation across multi-scale biological problems and rigorous assessment of their tool-augmented capabilities remain limited. Results We reveal a systematic gap between LLM performance and mechanistic understanding through the proposed cross-scale bio-molecular benchmark: BioMol-LLM-Bench, a unified framework comprising 26 downstream tasks that covers 4 distinct difficulty levels, and computational tools are integrated for a more comprehensive evaluation. Evaluation on 13 representative models reveals 4 benchmark-specific observations: chain-of-thought-style training does not consistently improve performance on the evaluated biological tasks; the evaluated hybrid mamba–attention model shows strong performance on long bio-molecular sequence tasks; supervised fine-tuned models show task-specific specialization with reduced performance in some general settings; and current LLMs perform better on classification tasks than on challenging regression tasks under this benchmark setting. Availability Source code is available at https://github.com/AI-HPC-Research-Team/BioMol-LLM-Bench</description>
    </item>
    <item>
      <title>Phenotype-driven de novo molecular design from gene expression signatures</title>
      <link>https://doi.org/10.64898/2026.07.21.739736</link>
      <guid isPermaLink="false">https://doi.org/10.64898/2026.07.21.739736</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <description>Yaxin Xu, Taojie Kuang, Shuang Ge, Haomin Wu, Mingqing Wang, Huan Xu, Fengwei An, Zhengyu Ma, Qiang Cheng, Zhixiang Ren

bioRxiv · Preprint · 2026

A bstract Target-based and structure-guided drug design remain central to modern drug discovery, but complementary strategies are needed when predefined targets or binding pockets do not fully capture disease biology. Gene-expression signatures provide scalable system-level readouts of disease and perturbation states, making them attractive inputs for phenotype-guided molecular design. However, preserving phenotypic information during molecular generation remains challenging, and chemically plausible molecules may lose connection to the intended biological response. Here, we present Tx2Mol, a transcriptome-guided framework that translates gene-expression signatures into candidate molecules while maintaining biological guidance throughout generation. We evaluated Tx2Mol across three biological settings: bulk gene perturbation, single-cell perturbation, and patient-derived disease signatures; and three validation dimensions: chemical plausibility, structural compatibility, and phenotypic preservation. Across 10 cancer-relevant bulk gene-perturbation benchmarks, Tx2Mol outperformed 9 transcriptome-guided baselines, improving maximum Tanimoto similarity to known ligands by 24.10% on average and by 50.67% on HDAC1. Structure-based analyses further supported structurally novel candidates with favorable predicted target binding. Tx2Mol also generalized to noisy single-cell perturbation profiles and preserved drug-induced transcriptional responses through in silico drug-perturbation validation. Patient-derived disease signatures further guided molecular generation toward approved-drug chemical space. Together, these results support gene-expression phenotypes as actionable guidance signals for phenotype-directed molecular design and candidate prioritization.</description>
    </item>
    <item>
      <title>Improving Variant Effect Prediction by Steering Sparse Mechanistic Features in Protein Language Models</title>
      <link>https://doi.org/10.64898/2026.05.12.724472</link>
      <guid isPermaLink="false">https://doi.org/10.64898/2026.05.12.724472</guid>
      <pubDate>Fri, 15 May 2026 00:00:00 GMT</pubDate>
      <description>Mingqing Wang, Meng Yuan, Athanasios V. Vasilakos, Yonghong He, Zhixiang Ren

bioRxiv · Preprint · 2026

Abstract Protein language models (PLMs) like the ESM series encapsulate immense evolutionary knowledge within their high-dimensional continuous embeddings. However, these latent representations are densely entangled, obscuring the fine-grained biophysical constraints necessary for precise functional resolution. To unlock the full expressive power of these embeddings, we propose PLM-SAE, a mechanistic framework that employs Sparse Autoencoders (SAEs) to disentangle PLM representations into discrete, biologically interpretable activations. By isolating and directly intervening on critical functional features, we fundamentally enhance the structural and mutational awareness of the underlying embeddings. We rigorously validate this embedding enhancement on variant effect prediction (VEP). In the unsupervised zero-shot setting, our sparse modulation elevates the state-of-the-art ESM-3 model, yielding performance improvements across 114 deep mutational scanning datasets and delivering an 80.8% relative improvement on challenging targets like the human E3 ubiquitin ligase HECD1. Furthermore, our target-specific differentiable gating mechanism achieves consistent performance gains in over 80% of evaluated datasets with an average Spearman ρ increase of +0.138. Finally, extending this approach to a cross-fitness multitask architecture establishes new state-of-the-art results on 17 VenusMutHub datasets, highlighted by a 169.0% performance surge in small-molecule binding predictions. Our work demonstrates that refining the highly entangled latent manifold via sparse modulation provides a robust and generalizable foundation for enhancing downstream PLM capabilities.</description>
    </item>
    <item>
      <title>Learning Protein Structure-Function Relationships through Knowledge-guided Representation Decomposition</title>
      <link>https://arxiv.org/abs/2605.23960</link>
      <guid isPermaLink="false">https://arxiv.org/abs/2605.23960</guid>
      <pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate>
      <description>Mingqing Wang, Zhiwei Nie, Athanasios V. Vasilakos, Yonghong He, Zhixiang Ren

arXiv · Preprint · 2026

Proteins encode diverse functions within complex three-dimensional structures, yet most deep learning representations remain highly entangled, obscuring the biophysical signals that underlie function. Here we introduce ProtDiS, a knowledge-guided framework that decomposes pretrained protein micro-environment embeddings into biologically grounded and task-relevant dimensions. Inspired by the information bottleneck principle, ProtDiS learns representations that balance informativeness and compression, yielding structural features that are more specific, independent, and information-efficient, and achieving consistent improvements across twelve downstream tasks, with the largest gains under structure-based splits. Protein- and residue-level analyses further show that ProtDiS differentiates proteins with similar folds but divergent functions and captures fine-grained biophysical signals critical. These findings suggest that knowledge-guided decomposition provides a general and interpretable approach for structuring latent spaces in protein structural modeling. The source code and implementation details are publicly available at https://github.com/AI-HPC-Research-Team/ProtDiS.</description>
    </item>
    <item>
      <title>Pseudodata-Guided Invariant Representation Learning Boosts the Out-of-Distribution Generalization in Enzymatic Kinetic Parameter Prediction</title>
      <link>https://doi.org/10.1021/acs.jcim.5c03204</link>
      <guid isPermaLink="false">https://doi.org/10.1021/acs.jcim.5c03204</guid>
      <pubDate>Mon, 11 May 2026 00:00:00 GMT</pubDate>
      <description>Haomin Wu, Zhiwei Nie, Hongyu Zhang, Zhixiang Ren

Journal of Chemical Information and Modeling · Published · 2026</description>
    </item>
    <item>
      <title>Enhanced Drug-drug Interaction Prediction Using Adaptive Knowledge Integration</title>
      <link>https://arxiv.org/abs/2603.12885</link>
      <guid isPermaLink="false">https://arxiv.org/abs/2603.12885</guid>
      <pubDate>Fri, 13 Mar 2026 00:00:00 GMT</pubDate>
      <description>Pengfei Liu, Jun Tao, Zhixiang Ren

arXiv · Preprint · 2026

Drug-drug interaction event (DDIE) prediction is crucial for preventing adverse reactions and ensuring optimal therapeutic outcomes. However, existing methods often face challenges with imbalanced datasets, complex interaction mechanisms, and poor generalization to unknown drug combinations. To address these challenges, we propose a knowledge augmentation framework that adaptively infuses prior drug knowledge into a large language model (LLM). This framework utilizes reinforcement learning techniques to facilitate adaptive knowledge extraction and synthesis, thereby efficiently optimizing the strategy space to enhance the accuracy of LLMs for DDIE predictions. As a result of few-shot learning, we achieved a notable improvement compared to the baseline. This approach establishes an effective framework for scientific knowledge learning for DDIE predictions.</description>
    </item>
    <item>
      <title>A Multi-task Large Reasoning Model for Molecular Science</title>
      <link>https://arxiv.org/abs/2603.12808</link>
      <guid isPermaLink="false">https://arxiv.org/abs/2603.12808</guid>
      <pubDate>Fri, 13 Mar 2026 00:00:00 GMT</pubDate>
      <description>Pengfei Liu, Shuang Ge, Jun Tao, Zhixiang Ren

arXiv · Preprint · 2026

Advancements in artificial intelligence for molecular science are necessitating a paradigm shift from purely data-driven predictions to knowledge-guided computational reasoning. Existing molecular models are predominantly proprietary, lacking general molecular intelligence and generalizability. This underscores the necessity for computational methods that can effectively integrate scientific logic with deep learning architectures. Here we introduce a multi-task large reasoning model designed to emulate the cognitive processes of molecular scientists through structured reasoning and reflection. Our approach incorporates multi-specialist modules to provide versatile molecular expertise and a chain-of-thought (CoT) framework enhanced by reinforcement learning infused with molecular knowledge, enabling structured and reflective reasoning. Systematic evaluations across 10 molecular tasks and 47 metrics demonstrate that our model achieves an average 50.3% improvement over the base architecture, outperforming over 20 state-of-the-art baselines, including ultra-large-parameter foundation models, despite using significantly fewer training data and computational resources. This validates that embedding explicit reasoning mechanisms enables high-efficiency learning, allowing smaller-scale models to surpass massive counterparts in both efficacy and interpretability. The practical utility of this computational framework was validated through a case study on the design of central nervous system (CNS) drug candidates, illustrating its capacity to bridge data-driven and knowledge-integrated approaches for intelligent molecular design.</description>
    </item>
    <item>
      <title>Prototype-based continual cell-type annotation reveals cellular state transitions in expanding single-cell atlases</title>
      <link>https://doi.org/10.64898/2026.03.05.709973</link>
      <guid isPermaLink="false">https://doi.org/10.64898/2026.03.05.709973</guid>
      <pubDate>Sun, 08 Mar 2026 00:00:00 GMT</pubDate>
      <description>Shuang Ge, Qiming He, Yiming Ren, Yaxin Xu, Mingqing Wang, Zhiwei Nie, Huan Xu, Qiang Cheng, Shuqing Sun, Zhixiang Ren

bioRxiv · Preprint · 2026

ABSTRACT Large-scale single-cell atlases provide an increasingly comprehensive view of cellular diversity, but their continued expansion across studies poses a fundamental challenge: preserving consistent cell identities while capturing biological variation in cellular states. Most existing annotation frameworks are built around static references, making it difficult to incorporate newly generated datasets into established cellular representations without retraining on historical data. When updated sequentially, these methods are further constrained by catastrophic forgetting and batch-specific biases, limiting scalability and the continuity of knowledge integration. Here we introduce scEvolver, a continual learning framework for single-cell annotation that incrementally accumulates knowledge through memory-guided refinement of cell-type prototypes without revisiting historical data. Across sequencing platforms, tissue contexts and molecular modalities, scEvolver supports robust annotation and external query mapping with substantially fewer labelled reference cells. By preserving consistent cell-type semantics across datasets while capturing biologically meaningful within-class heterogeneity, scEvolver enables the identification of epithelial cell-state transitions in inflammatory gut disease. External mapping to the healthy Human Lung Cell Atlas further reveals shared cell-state deviations across multiple diseases, including an FCGR3A + inflammatory monocyte programme in sarcoidosis, chronic obstructive pulmonary disease and idiopathic pulmonary fibrosis, highlighting scEvolver’s potential to characterize context-specific cellular dynamics in complex disease settings.</description>
    </item>
    <item>
      <title>A self-feedback knowledge elicitation approach for chemical reaction predictions</title>
      <link>https://doi.org/10.1016/j.engappai.2025.111112</link>
      <guid isPermaLink="false">https://doi.org/10.1016/j.engappai.2025.111112</guid>
      <pubDate>Mon, 01 Sep 2025 00:00:00 GMT</pubDate>
      <description>Pengfei Liu, Jun Tao, Zhixiang Ren

Engineering Applications of Artificial Intelligence · Published · 2025</description>
    </item>
    <item>
      <title>Deep learning methods for protein representation and function prediction: A comprehensive overview</title>
      <link>https://doi.org/10.1016/j.engappai.2025.110977</link>
      <guid isPermaLink="false">https://doi.org/10.1016/j.engappai.2025.110977</guid>
      <pubDate>Mon, 01 Sep 2025 00:00:00 GMT</pubDate>
      <description>Mingqing Wang, Zhiwei Nie, Yonghong He, Athanasios V. Vasilakos, Qiang (Shawn) Cheng, Zhixiang Ren

Engineering Applications of Artificial Intelligence · Published · 2025</description>
    </item>
    <item>
      <title>A Unified Peptide Generative Framework via a Weakly Order-Dependent Autoregressive Language Model and Lifelong Learning</title>
      <link>https://doi.org/10.1021/acs.jcim.5c00623</link>
      <guid isPermaLink="false">https://doi.org/10.1021/acs.jcim.5c00623</guid>
      <pubDate>Mon, 11 Aug 2025 00:00:00 GMT</pubDate>
      <description>Zhiwei Nie, Daixi Li, Yutian Liu, Fan Xu, Hongyu Zhang, Xiansong Huang, Xudong Liu, Zhennan Wang, Yiming Ma, Yuxin Ye, Feng Yin, Wen-Bin Zhang, Zhixiang Ren, Zhihong Liu, Zigang Li, Jie Chen

Journal of Chemical Information and Modeling · Published · 2025</description>
    </item>
    <item>
      <title>A multi-modal genomic knowledge distillation framework for drug response prediction</title>
      <link>https://doi.org/10.1007/s10489-025-06768-9</link>
      <guid isPermaLink="false">https://doi.org/10.1007/s10489-025-06768-9</guid>
      <pubDate>Fri, 01 Aug 2025 00:00:00 GMT</pubDate>
      <description>Shuang Ge, Shuqing Sun, Huan Xu, Qiang Cheng, Zhixiang Ren

Applied Intelligence · Published · 2025</description>
    </item>
    <item>
      <title>Predicting protein stability changes upon mutations with dual-view ensemble learning from single sequence</title>
      <link>https://doi.org/10.1093/bib/bbaf319</link>
      <guid isPermaLink="false">https://doi.org/10.1093/bib/bbaf319</guid>
      <pubDate>Wed, 02 Jul 2025 00:00:00 GMT</pubDate>
      <description>Zhiwei Nie, Yiming Ma, Yutian Liu, Xiansong Huang, Zhihong Liu, Peng Yang, Fan Xu, Feng Yin, Zigang Li, Jie Fu, Zhixiang Ren, Wen-Bin Zhang, Jie Chen

Briefings in Bioinformatics · Published · 2025

Abstract Predicting the protein stability changes upon mutations is one of the effective ways to improve the efficiency of protein engineering. Here, we propose a dual-view ensemble learning-based framework, DVE-stability, for mutation-induced protein stability change prediction from single sequence. DVE-stability integrates the global and local dependencies of mutations to capture the intramolecular interactions from two views through ensemble learning, in which a structural microenvironment simulation module is designed to indirectly introduce the information of structural microenvironment at the sequence level. DVE-stability achieved state-of-the-art prediction performance on seven single-point mutation benchmark datasets, and comprehensively surpassed other methods on five of them. Furthermore, DVE-stability outperformed other methods comprehensively through zero-shot inference on multiple-point mutation prediction task, demonstrating superior model generalizability to capture the epistasis of multiple-point mutations. More importantly, DVE-stability exhibited superior generalization performance in predicting rare beneficial mutations that are crucial for practical protein directed evolution scenarios. In addition, DVE-stability identified important intramolecular interactions via attention scores, demonstrating interpretable. Overall, DVE-stability provides a flexible and efficient tool for mutation-induced protein stability change prediction in an interpretable ensemble learning manner.</description>
    </item>
  </channel>
</rss>
