Hero image

ARTICLE

How AI Is Pinpointing the Genetic Cause of Disease

ARTICLE

Julie Langelier

|

March 23, 2026

|

17 min read

Gladstone scientists who develop new technologies—both AI models on the computer and molecular tools in the lab—are coming together to solve one of the most challenging problems in science: what specific changes in the human genome cause disease?

This article is part of a series about the many ways our scientists are using—and developing—AI tools for biomedical research. Sign up for our newsletter to have these stories delivered right to your inbox.


Your genome is made up of all the DNA you inherit—about 3 billion letters of biological code. Today, you can have your genome sequenced for a few hundred dollars. That means you can know the exact order of all those DNA letters, a blueprint which is unique to you.

Although your DNA is about 99.9 percent the same as all other humans, the remaining tiny fraction contains millions of small differences (called genetic variants). These differences can influence how your body functions, how you respond to medications, and your chances of developing certain diseases.

“The ultimate goal would be to go see your doctor, have your genome sequenced, and have your doctor tell you what your specific genetic variants mean, whether they’re related to the symptoms you’re experiencing, and whether they might increase your risk for a disease,” says Seth Shipman, PhD, investigator at Gladstone Institutes.

But we’re not there just yet. Scientists and clinicians currently don’t know how to fully decipher the genome and figure out which of the 3 billion letters matter.

“We’re about to enter a world where everybody will have their DNA sequenced, but most of the information we get from a genome sequence is still not interpretable,” says Deepak Srivastava, MD, president of Gladstone. “Only a few of the millions of genetic changes are meaningful for disease, and we can’t yet pinpoint what those are.”

So, why is it so difficult to identify the exact genetic cause of disease? In part, it’s because such a huge number of DNA changes could be responsible that scientists can’t test all the possibilities. Without the ability to prioritize the countless options, progress has been slow, or even stalled.

This is where artificial intelligence, or AI, can completely shift the paradigm.

By using AI to predict outcomes before ever going into a lab, we can increase the speed of discovery a millionfold.”

Katie Pollard, PhD

“We can run thousands of experiments on the computer in one day that would take years in a traditional lab,” says Katie Pollard, PhD, director of the Gladstone Institute of Data Science and Biotechnology. “By using AI to predict outcomes before ever going into a lab, we can increase the speed of discovery a millionfold.”

Pollard and her colleagues are set on leveraging the power of AI to decode the entire genome and, for the first time, actually figure out what it means.

“We’re essentially trying to establish a lookup table,” says Gladstone Investigator Vijay Ramani, PhD. “If a patient comes in, we could then find their genetic variant and let them know how likely they are to develop a particular disease.”

Where AI Meets the Lab

At Gladstone, scientists who develop new technologies—both on the computer and in the lab—have joined forces to build a platform that combines advanced computational models with novel tools in the experimental lab to test their hypotheses about how DNA works.

The first component of this platform is an AI model, designed by Pollard’s team, that can look at millions of genome sequences and make predictions about which DNA changes are most likely causing disease.

These AI predictions won’t all be correct, so they are then tested in the Shipman Lab. This group invented a new genome editing method to edit the DNA of hundreds or thousands of human cells simultaneously in a single dish.

Then, the edited cells are analyzed using a new measurement technology developed by Ramani, which can confirm the edit that each cell received and, importantly, measure the effect of each edit on the function of the cell.

Finally, the scientists use all the data from the experiments to retrain the AI model, telling it which predictions were right and which were wrong, to make it better, faster, and more accurate—and start the whole process again.

“Nobody has tried this at this scale before,” Pollard says. “It’s been possible to make a single change, or a small number of changes, in DNA, but our team is going to interrogate the entire genome and try to understand how each letter of the genome works.”

This joint solution didn’t come about easily. Pollard, Shipman, and Ramani have had continuous discussions for over a year to get to this point.

“The fact that our labs can collaborate so seamlessly has been crucial,” Ramani says. “We kept going back and forth to figure out how we could improve our own technologies in a way that would overcome the bottlenecks in our fields. By working toward this common goal, we each pushed the limits of our labs and we can now achieve something unprecedented.”

A Quick Lesson in Genomics

Scientists used to assume disease happens when a gene is broken. So when they first sequenced the human genome in the early 2000s, they thought they would finally understand what causes disease. Instead, they found something unexpected: most disease-linked genetic differences aren’t actually in genes that make proteins—they’re in non-coding DNA.

Non-coding DNA, which makes up 98–99 percent of the genome, acts as a control panel, providing instructions for when, where, and how strongly genes are used. If you think of genes as lights in a skyscraper, non-coding DNA would be the switches and dimmers that control them. So, as it turns out, disease is most often not caused by broken bulbs, but faulty switches—perhaps causing a light to turn on when it shouldn’t or making it shine too dimly.

Studying non-coding DNA has been one of the hardest problems in genomics.

Its scale is overwhelming, and regions may only be active in one specific cell type, making it difficult to replicate in the lab. Still, the biggest challenge is the way DNA is neatly folded in 3D space to fit inside the nucleus of a cell. A stretch of non-coding DNA can influence a gene that’s hundreds of thousands of letters away. When DNA is folded, these two sections become physically close to one another. But with traditional sequencing techniques, which flatten DNA into a linear readout, those important spatial relationships become invisible.

These are exactly the types of problems AI is designed for—finding subtle patterns across enormous, complex datasets that humans would never spot.

From a Million Possibilities to a Thousand

Gladstone’s revolutionary platform begins in Pollard’s lab, where she and her team develop AI models using deep learning. Deep learning is a way for a computer to recognize patterns by learning on its own from a huge number of examples, and then use those patterns to make predictions about data it has never seen.

This AI approach is called “deep” because it uses many layers of processing. Each layer performs a simple operation, but when they’re combined, the system can learn very complex patterns—much like each step of a factory assembly line does a simple task, but at the end, the tasks can result in a very elaborate product.