Building the Data Foundation for AI-Driven ALS Drug Discovery
August 25, 2026 Manish Raisinghani
Artificial intelligence is changing how researchers approach drug discovery. It can help scientists analyze large amounts of information, identify patterns that might otherwise be missed, and generate hypotheses faster than traditional approaches alone. For a disease like ALS, where time is everything, that potential is especially compelling.
This year at the Target ALS Annual Meeting, leaders from Roche, Lilly Ventures, and insitro discussed how their companies are applying artificial intelligence and machine learning to drug discovery. Their optimism was accompanied by an important note of caution: while AI is already improving aspects of drug discovery, realizing its full potential in areas like clinical development will take time.
But there is another piece of the equation that receives less attention: AI is only as powerful as the data behind it.
AI and machine learning models identify patterns in the information they are given, which makes the quality and breadth of that data critically important.
If a dataset represents only a narrow segment of people living with ALS, the insights it generates may be similarly limited. If different types of clinical and biological information exist in disconnected systems, researchers may struggle to examine how they relate to one another. For AI to meaningfully contribute to the discovery of effective treatments and biomarkers for ALS, researchers need comprehensive datasets that capture the disease’s complexity. They also need that data to be organized, accessible, and connected to the resources required to test what they find.
Target ALS has spent more than a decade helping build that foundation. That long-term investment now puts us in a strong position to leverage advances in AI and machine learning.
Launched in 2024, the Target ALS Data Engine brings comprehensive multimodal datasets from these research efforts into a centralized resource for researchers worldwide. This includes clinical, demographic, and epidemiologic information, as well as multiple types of biological data, including cutting-edge long-read whole-genome sequencing, proteomics, RNA sequencing, and more.
The goal isn’t simply to collect more data; it’s to make that data impactful, and representation is part of what makes that possible.
ALS affects people of every race and ethnicity around the world, yet ALS research has historically lacked ethnic, genetic, and geographic diversity. Through our ALS Global Research Initiative (AGRI), Target ALS is expanding participation in clinical research efforts among diverse and historically underrepresented populations, helping build datasets that better reflect the ALS community. More representative datasets can provide researchers with a deeper understanding of disease biology and its progression, paving the path to identifying reliable biomarkers and effective treatments.
In other words, representation isn’t separate from data quality; it is part of data quality.
Identifying a promising pattern is only one step in the discovery process. Researchers still need to test whether those findings hold up experimentally.
That’s where the connection between data and biosamples becomes especially valuable. Not only can researchers access comprehensive datasets through our Data Engine, but matched biosamples, including blood, urine, cerebrospinal fluid, and postmortem brain and spinal cord tissue, are available for request through our Research Cores. This connection can help researchers move from identifying a potential signal in the data to testing what it might mean biologically.
But the value of data also depends on how readily researchers can access and use it.
This is why Target ALS makes the Data Engine freely available to academic and industry researchers worldwide, with access provided within 48 hours of signing a no-strings-attached data user agreement. This removes barriers of cost, logistics, and the time required to collect, generate, and curate comprehensive multimodal datasets. Researchers using the platform have seen the impact of that approach firsthand. Jamie Ifkovits, PhD, Scientific Director of Neurodegeneration at GSK, described open access and data harmonization as “major differentiators,” helping reduce barriers between academia and industry and enable faster, more efficient research.
Today, over 660 researchers across 35 countries are already using the platform. By lowering these barriers, Target ALS can also incentivize researchers and companies with expertise in AI, computational biology, and drug development to bring those capabilities to ALS research, whether or not they’ve worked on the disease before.
The result is simple: less time navigating barriers and more time pursuing the scientific questions that can accelerate the discovery process forward.
Target ALS began building this data and biosample framework as a reflection of our values. We deliberately disrupted the barriers to access, forged radical collaborations with academia, industry, and non-profits to build these tools, and created processes designed to get them into researchers’ hands quickly. That approach reflects the urgency of our mission: people living with ALS don’t have time to wait, and researchers need the data, samples, and infrastructure to move promising ideas forward faster.
As AI and machine learning continue to advance, their impact on ALS drug discovery will depend on the foundation beneath them. The algorithms will continue to evolve; however, our responsibility is to ensure ALS research has the right tools to make the most of them, so we can accelerate toward a future where we have effective treatments for all forms of ALS and everyone with ALS can live a long, high-quality life.
Related News
Brenda Bucklin's Story: Living with C9-ALS
August 10, 2026
Announcing the 2026 In Vivo Target Validation Projects
July 21, 2026