The NIAID Office of Data Science and Emerging Technologies (ODSET) highlights publications that feature innovative uses of data science and bioinformatics in infectious, immune-mediated and allergic disease research.
Explore NIAID data science publications on PubMed:
- NIAID-funded publications that use data science or computational biology (since 2023).
- Publications funded or co-funded by ODSET.
- Publications from the Harnessing Big Data to Halt HIV initiative.
If you would like to feature a publication on this page, please contact Data Science. Publications should feature research related to infectious, immunologic, and allergic diseases; include data science or a related discipline; and cite NIAID funding in the manuscript. Please include in your email:
- The title of your published article.
- A link to the article.
- A 50-60 word description of the article.
183 Results
Risk and Protective Factors Associated With HIV-Related Mortality in Children Receiving Antiretroviral Therapy: A Model-Based Meta-Regression
September 3, 2026 JAMA Pediatrics
Causes of higher rates of mortality among children getting antiretroviral therapy (ART) in resource-limited settings are not fully understood. Here, the authors perform a meta-analysis using regression and machine learning techniques to identify modifiable factors associated with HIV mortality while receiving ART.
Integrating multi-omics data with network modeling to characterize Pseudomonas aeruginosa persister cell metabolism
September 3, 2026 Journal of Bacteriology
To better characterize the metabolism of Pseudomonas aeruginosa persister cells, which are transient variants that can tolerate antimicrobial treatment and are associated with chronic infections, the authors integrated transcriptomics, metabolomics, and genome-wide metabolic modeling. They identify metabolic pathways critical for persister survival.
SeqBoard: a genomics-based data dashboard for comprehensive wastewater virome monitoring
September 3, 2026 Journal of the Americal Medical Informatics Association
The authors present a public health dashboard to display viral genomic sequencing data from wastewater. By using agnostic genomic sequencing, as opposed to traditional PCR methods, the dashboard workflow is able to detect and visualize trends of any virus genomes present.
Central T Cell Tolerance from Sparse Peptide Sampling
August 19, 2026 Science Advances
Negative selection of T cells critical for preventing autoimmune responses, but how general tolerance to self-peptides is achieved from the relatively small number of self-peptides tested in the thymus is not well understood. The authors here present a model of negative selection that incorporates estimates from experimental methods to show how generalized tolerance can be achieved through sparse random sampling.
Design of an immunogen containing multidimensionally conserved and immunogenic parts of the HIV proteome
August 14, 2026 PLoS Computational Biology
To aid efforts to develop HIV vaccines, the authors developed a computational model to identify regions of HIV proteins that are conserved and difficult to be rescued by compensatory mutations.
Pocket restraints guided by B-cell epitope prediction improve Chai-1 antibody-antigen structure modeling
August 3, 2026
Accurately predicting antibody-antigen (AbAg) interactions is critical for designing therapeutics and diagnostics, but current deep learning methods often fail to predict the correct structure. Clifford and Nielsen present two open source tools that predict sequence binding restraints for use in AbAg structure predictions that they show improve the accuracy of predicted AbAg complexes.
Treeline Provides a Unified Strategy for Optimising Phylogenetic Trees Under Alternative Criteria
August 3, 2026
Phylogenetic trees can be built to optimize different objectives and these different trees may be more or less accurate. To facilitate comparing phylogenetic trees constructed using different optimization strategies, Wright presents Treeline as a part of the DECIPHER R package.
Predicting functions of uncharacterized gene products from microbial communities
August 3, 2026
Although critical for a wide range of biological and environmental functions, the majority of microbial proteins are uncharacterized. To address this, the authors developed a computational model to predict protein function based on metatranscriptomic co-expression patterns and other microbial community-wide data.
Seqwin: ultrafast identification of signature sequences in microbial genomes
August 3, 2026
Signature sequences are regions of a genome that are specific to a target microbial taxonomic group and therefore can be used to develop sensitive and specific diagnostics for that microbial group. The authors present Seqwin as an efficient computational method to discover such sequences from large-scale genomic data.
Uncovering Household Tuberculosis Infection Testing and Care Patterns Using a Novel Bioinformatics Linkage Strategy
August 3, 2026
This study utilized a bioinformatics linkage strategy in a large clinical database to identify and track tuberculosis (TB) infection and treatment among household members of patients with known prior TB infection. The authors identify household patterns of TB infection that they hope can inform approaches to TB infection testing and treatment
Same-Slide Spatial Multiomics Integration with IN-DEPTH Reveals Tumor Virus-Linked Spatial Reorganization of the Tumor Microenvironment
August 3, 2026 Cancer Discovery
To enable spatial transcriptomics and proteomics to be performed on the same tissue sample, the authors developed a novel spatial multiomics technique and corresponding graph representation method for quantifying the spatially linked pathways identified. The authors make the graph representation method publicly available as an R package.
ClusterApp to visualize, organize, and navigate metabolomics data
July 29, 2026 PLoS Computational Biology
To address the lack of easy-to-use tools for methodologically sound clustering analysis of metabolomics data, the authors developed ClusterApp to enable Principal Coordinate Analysis and visualization. The tool is freely available both through a web interface or a Docker image.
MicNet: integrating spatially resolved transcriptomes and pathology images by contrastive deep neural network
June 11, 2026 Genome Biology
Integrating gene expression data and pathology imaging remains a challenge despite recent advances. Here, the authors develop a deep learning-based method, MicNet, that projects spatial transcriptomic and imaging data onto a shared representation and enables improved analysis of gene expression in a spatial context.
High-resolution metabolomic analysis of stool reveals expanded biomarkers of C. difficile colitis and insights into pathophysiology
June 2, 2026 Microbiology Spectrum
Clostridioides difficile infection is facilitated by disturbances in the gut microbiome, and associated changes in the corresponding gut metabolome. To aid development of improved diagnostics, the authors compared the metabolic profiles of stool from infected and control participants. Numerous metabolic pathways were altered in infected samples, which the authors suggest may promote inflammation and facilitate infection.
Multi-omic analysis reveals nitric oxide dependent remodeling in classically activated macrophages and identifies negative regulation mediated by AKR1A1
June 1, 2026 Redox Biology
To investigate the regulatory effect of nitric oxide (NO), which is generated by macrophages in response to certain stimuli, the authors profiled the proteome and transcriptome of activated and not activated macrophages. They show NO signaling suppresses expression of electron transport chain components, while increasing expression of enzymes involved in redox defense, including AKR1A1.
Predictions from deep learning propose substantial protein-carbohydrate interplay
May 26, 2026 PNAS
Interactions between proteins and carbohydrates underlie numerous biological functions, but it is challenging to experimentally identify protein-carbohydrate interactions. To address this, the authors developed a neural network to predict whether a given protein interacts with carbohydrates that they show to be highly accurate.
Phyling: phylogenetic inference from annotated genomes
May 6, 2026 G3
The authors present a new tool, Phyling, for inferring phylogeny directly from genomic data that is designed to efficiently handle large-scale datasets while maintaining accuracy. Phyling is available as an open-source Python package.
BERTopic-driven term extraction from biomedical texts toward ontology population: evaluating vaccine ontology with Plotkin's vaccines corpus
May 3, 2026 Journal of Biomedical Semantics
To support the use of ontologies for vaccine data, the authors developed a natural language processing-based pipeline to semi-automatically extract ontology-relevant ideas from biomedical text. The code underlying the BERTopic tool is freely available.
Human Immunodeficiency Virus-Associated Proteomic Signature of Myocardial Fibrosis and Incident Heart Failure
April 29, 2026 Journal of Infectious Diseases
To better understand the higher risk of myocardial fibrosis and subsequent heart failure faced by people with HIV, the authors performed blood plasma proteomics and cardiac imaging on people with and without HIV. They identify proteins, enriched for processes such as T-cell activation and ephrin signaling, associated with myocardial fibrosis in people with HIV and time to heart failure.
Reliable detection of Host-Microbe Signatures in cancer using PRISM
April 13, 2026 Cancer Cell
The authors present a computational framework for improving decontamination and accurately identifying microbial sequences from low-biomass sequencing data, such as that collected from human tumors.
Classification of outcomes in antimalarial therapeutic efficacy studies with Aster
April 13, 2026 Antimicrobial Agents and Chemotherapy
Difficulties in distinguishing reoccurrences of the same infection from new infections pose a major challenge for accurate assessment of the efficacy of antimalarial therapies. The authors provide a novel statistical framework, available as an R package, that accounts for factors leading to misclassification and outperforms algorithms currently recommended methods.
TEpiNom: A computational framework integrating population data to prioritize Plasmodium falciparum T cell epitopes
April 2, 2026 Vaccine
Vaccine development against malaria is challenging due to variation in both Plasmodium falciparum (the malaria-causing parasite) and human genetics. This study presents a tool, TEpiNom, which integrates pathogen sequence diversity, predicted T cell binding, and population-specific HLA allele frequencies to identify high priority epitopes for further investigation.
Variation in Microbiome Composition and Faecal Metabolites Are Associated With Differential Susceptibility to DSS-Induced Colitis
April 1, 2026 Immunology
Using mouse models, the authors identify differences in microbiome composition that are associated with a higher likelihood of developing colitis. Additionally, metabolomics data revealed different metabolites associated with colitis susceptibility and severity.
Compressing the collective knowledge of ESM into a single protein language model
March 30, 2026 Nature Methods
Most current high-performing methods of predicting the effect of genomic variants use protein language models (PLMs) along with additional data such as protein structure or homology. Here, the authors show PLMs trained only on raw sequences are able to achieve state-of-the-art performance and accurately predict the severity of variant effects on clinical phenotypes.
MoleRate: comparing molecular relative evolutionary rates to detect convergent evolution
March 3, 2026 Evolution
To aid in studies of convergent evolution, the authors developed MoleRate, which uses likelihood models to identify changes in evolutionary rates among subsets of phylogenetic tree branches (i.e. genes or genomic regions) compared to the average of the overall genome tree. MoleRate is available as an open-source package that enables analysis and visualization.