ORCID Identifier(s)

0009-0009-7962-3713

Graduation Semester and Year

Summer 2026

Language

English

Document Type

Thesis

Degree Name

Master of Science in Computer Science

Department

Computer Science and Engineering

First Advisor

Chengkai Li

Second Advisor

Gautam Das

Third Advisor

Negin Fraidouni

Abstract

Identifying genes responsible for plant traits remains a fundamental challenge in plant biology. Although advances in genomics have enabled researchers to identify large sets of candidate genes associated with a trait, experimentally validating each candidate remains costly and time-consuming, creating a need for computational methods that can prioritize likely causal genes while providing interpretable biological evidence. We present GeneSieve, a system for candidate gene prioritization and gene discovery across diverse plant species. Given a natural-language trait description and a set of candidate genes from any angiosperm species, GeneSieve constructs a Trait-Gene Association Graph (TGAG) for each candidate by integrating four complementary evidences: trait similarity, trait-gene associations, gene co-expression, and sequence homology. The underlying biological data comes from four keystone plant species—rice, maize, soybean, and Arabidopsis—linking 1,055 traits to 127,611 genes through 952,630 curated gene-trait associations and more than 2.4 billion gene-gene co-expression relationships. Candidate genes are prioritized using a relational graph convolutional network (R-GCN) that learns from TGAGs' heterogeneous topology and relationships. The model was trained and evaluated using cross-validation on more than 3,000 TGAGs generated from a curated benchmark of experimentally validated causal gene-trait associations. GeneSieve consistently outperformed baseline methods, placing the validated causal gene among the top five candidates in approximately 70% of queries. GeneSieve also provides an interactive web platform for candidate gene prioritization and evidence visualization. The novelty of GeneSieve lies in addressing key limitations of existing candidate gene prioritization methods by constructing query-specific heterogeneous graphs that integrate LLM-derived trait similarity, curated trait-gene associations, RNA-seq-derived gene co-expression, and sequence homology into a unified graph framework. This enables evidence-driven R-GCN-based prioritization across diverse monocot and dicot plant species, rather than relying on static knowledge graphs, isolated evidence sources, or direct LLM prediction.

Keywords

Graph Neural Networks, Relational Graph Convolutional Networks, Large Language Models, Heterogeneous Graphs, Computational Biology, Bioinformatics, Candidate Gene Prioritization, Gene Discovery, Functional Genomics, Plant Genomics

Disciplines

Artificial Intelligence and Robotics | Bioinformatics | Databases and Information Systems

License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.

Available for download on Wednesday, August 11, 2027

Share

COinS