Graduation Semester and Year
Summer 2026
Language
English
Document Type
Thesis
Degree Name
Master of Science in Computer Science
Department
Computer Science and Engineering
First Advisor
Chengkai Li
Second Advisor
Gautam Das
Third Advisor
Negin Fraidouni
Abstract
Identifying genes responsible for plant traits remains a fundamental challenge in plant biology. Although advances in genomics have enabled researchers to identify large sets of candidate genes associated with a trait, experimentally validating each candidate remains costly and time-consuming, creating a need for computational methods that can prioritize likely causal genes while providing interpretable biological evidence. We present GeneSieve, a system for candidate gene prioritization and gene discovery across diverse plant species. Given a natural-language trait description and a set of candidate genes from any angiosperm species, GeneSieve constructs a Trait-Gene Association Graph (TGAG) for each candidate by integrating four complementary evidences: trait similarity, trait-gene associations, gene co-expression, and sequence homology. The underlying biological data comes from four keystone plant species—rice, maize, soybean, and Arabidopsis—linking 1,055 traits to 127,611 genes through 952,630 curated gene-trait associations and more than 2.4 billion gene-gene co-expression relationships. Candidate genes are prioritized using a relational graph convolutional network (R-GCN) that learns from TGAGs' heterogeneous topology and relationships. The model was trained and evaluated using cross-validation on more than 3,000 TGAGs generated from a curated benchmark of experimentally validated causal gene-trait associations. GeneSieve consistently outperformed baseline methods, placing the validated causal gene among the top five candidates in approximately 70% of queries. GeneSieve also provides an interactive web platform for candidate gene prioritization and evidence visualization. The novelty of GeneSieve lies in addressing key limitations of existing candidate gene prioritization methods by constructing query-specific heterogeneous graphs that integrate LLM-derived trait similarity, curated trait-gene associations, RNA-seq-derived gene co-expression, and sequence homology into a unified graph framework. This enables evidence-driven R-GCN-based prioritization across diverse monocot and dicot plant species, rather than relying on static knowledge graphs, isolated evidence sources, or direct LLM prediction.
Keywords
Graph Neural Networks, Relational Graph Convolutional Networks, Large Language Models, Heterogeneous Graphs, Computational Biology, Bioinformatics, Candidate Gene Prioritization, Gene Discovery, Functional Genomics, Plant Genomics
Disciplines
Artificial Intelligence and Robotics | Bioinformatics | Databases and Information Systems
License

This work is licensed under a Creative Commons Attribution 4.0 International License.
Recommended Citation
Umakant Pujar, Pranav, "GENESIEVE: A KNOWLEDGE NETWORK INTERFACE FOR GENE DISCOVERY ACROSS DIVERSE CROP SPECIES" (2026). Computer Science and Engineering Theses. 5.
https://mavmatrix.uta.edu/cse_theses2/5
Included in
Artificial Intelligence and Robotics Commons, Bioinformatics Commons, Databases and Information Systems Commons