Here is the complete, elaborated content for both slides with 15 references - all content is originally synthesized and does not reproduce source text:
Book Chapter: AI for Rare Disease and Orphan Drug Development
SLIDE 1 (Topic 1): Regulatory and Economic Considerations
Slide Title: Governing the Algorithm: Regulatory Pathways and Economic Realities in AI-Assisted Orphan Drug Development
1.1 The Regulatory Imperative for AI in Orphan Drug Submissions
The integration of artificial intelligence into drug development has prompted regulatory agencies worldwide to establish governance structures that ensure AI-generated evidence meets the same standards of scientific rigor as conventional data. For orphan drugs - where small patient populations make large randomized trials impractical - AI tools now serve as a primary mechanism for generating supportive evidence, which elevates the stakes of regulatory oversight considerably.
In January 2025, the U.S. Food and Drug Administration (FDA) issued its draft guidance titled Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products. This document introduced a risk-stratified, seven-step credibility assessment framework requiring sponsors to: (i) define the context of use (COU) for any AI model, (ii) characterize model uncertainty, (iii) validate the model against independent datasets, and (iv) document the lineage of training data and model parameters in all regulatory submissions. The framework does not treat all AI applications equally - models used to support dosing decisions in pivotal trial submissions face a higher credibility threshold than those used in exploratory analysis [1, 2].
Simultaneously, the European Medicines Agency (EMA) released its Reflection Paper on the Use of Artificial Intelligence in the Medicinal Product Lifecycle (September 2024), which takes a broader lifecycle perspective. Rather than focusing solely on submission quality, the EMA framework requires sponsors to demonstrate that AI tools are governed throughout the product lifecycle - from discovery and clinical trial design through manufacturing and post-market surveillance. Algorithmic transparency and explainability are non-negotiable requirements under this framework, particularly where AI outputs influence safety labeling decisions [2].
In a landmark development in January 2026, the FDA and EMA jointly published Guiding Principles of Good AI Practice in Drug Development - ten harmonized principles covering AI model building, validation, monitoring, and governance across the full pharmaceutical lifecycle. This represented the first transatlantic regulatory alignment on AI in drug development and carries particular relevance for sponsors pursuing simultaneous FDA and EMA orphan designation, as it removes the need to maintain two separate AI compliance strategies [2].
The UK's Medicines and Healthcare products Regulatory Agency (MHRA) introduced its "AI Airlock" regulatory sandbox in May 2024, allowing developers to test AI-based medical and drug-related tools in a controlled pre-authorization environment. This sandbox model lowers the barrier to entry for academic groups and small biotechs - precisely the organizations most active in orphan indications - by providing regulatory feedback before full submission [2].
For sponsors using model-informed drug development (MIDD) approaches, the FDA's Accelerating Rare Disease Cures (ARC) program within the Center for Drug Evaluation and Research (CDER) has established a collaborative framework specifically for quantitative systems pharmacology (QSP)-informed rare disease submissions. A dedicated workshop in 2023 produced a formal roadmap for integrating QSP and AI-driven disease progression models into regulatory review, proposing that AI-generated totality-of-evidence packages could substitute for additional prospective clinical data in orphan conditions [3, 4].
1.2 Adaptive Trial Designs and AI-Supported Regulatory Flexibility
Traditional phase III trials are structurally unsuitable for rare diseases: most conditions affect fewer than one in 2,000 individuals, and pediatric populations are disproportionately represented, raising additional ethical constraints on experimental exposures [4]. Regulatory agencies have responded by formally accepting AI-assisted adaptive trial designs, Bayesian borrowing of historical controls, and external control arms as valid sources of confirmatory evidence.
AI tools now perform continuous interim analyses in adaptive trials, allowing pre-specified response-adaptive randomization and early stopping rules to be applied with greater statistical precision than manual review could provide [4]. These tools simultaneously flag safety signals in real time using pharmacovigilance algorithms trained on multi-source data including electronic health records (EHRs), social media, and spontaneous reporting systems - an approach that partially compensates for the chronically under-powered post-marketing safety databases that characterize orphan drugs [5].
The FDA's plausible biological mechanism pathway, introduced as part of its evolving rare disease toolkit, permits conditional approval when a sponsor can demonstrate mechanistically interpretable evidence of activity in a small patient cohort combined with meaningful clinical improvement. AI platforms that generate interpretable, mechanistically grounded outputs are explicitly positioned to provide this type of evidence, whereas black-box models remain ineligible under current guidance [2, 5].
1.3 Economic Architecture of the Orphan Drug Market
The economics of orphan drug development are structured around a series of legislative and market incentives that distinguish this sector from mainstream pharmaceutical development. The U.S. Orphan Drug Act of 1983 grants seven years of market exclusivity from the date of approval, 50% tax credits on qualified clinical trial expenditures, waiver of FDA user fees (valued at over $3 million per application in 2024), and expedited regulatory review [6]. The European Union's analogous framework provides ten years of market exclusivity, protocol assistance, fee reductions, and access to centralized EMA review - creating a two-jurisdiction incentive architecture that AI-empowered sponsors can now exploit more efficiently by running parallel submissions with AI-harmonized dossiers [2, 6].
These incentives have driven extraordinary commercial growth. The global orphan drug market was valued at approximately $193 billion in 2024 and is projected to reach $621 billion by 2034, representing a compound annual growth rate of roughly 12.24% [6]. By 2026, orphan drugs are forecast to account for approximately 20% of all prescription drug sales in the United States - a share that has doubled in a decade and reflects both the surge in rare disease approvals (51% of FDA novel drug approvals in 2023 carried orphan designation) and the premium pricing these drugs command given limited therapeutic competition [6].
However, this pricing power is under increasing pressure. A 2022 survey of U.S. payers documented that insurers are deploying tighter utilization controls on high-cost specialty drugs including orphan therapies, shifting coverage from the medical to the pharmacy benefit in ways that reduce reimbursement rates and impose prior authorization hurdles [6]. European health technology assessment (HTA) bodies - notably Germany's IQWiG and France's HAS - have applied stricter benefit-risk frameworks that orphan therapies increasingly fail to satisfy, triggering managed entry agreements with rebates that can reduce effective prices by 30-50% below list price [6].
AI contributes meaningfully to the economics of development at both ends of this equation. On the cost side, AI-driven in-silico screening can evaluate thousands of candidate molecules before any physical synthesis occurs, reducing the experimental attrition rate and preserving scarce active pharmaceutical ingredient (API) quantities - a significant saving in orphan programs where API supply is often limited and expensive to produce [6]. AI stability prediction models also reduce formulation development timelines, particularly for biologics and gene therapies that require complex cold-chain optimization [6]. On the revenue side, AI-powered patient identification tools embedded in hospital information systems can expand diagnosed patient populations, directly increasing the addressable market size for approved orphan therapies and improving payer willingness to reimburse [7].
1.4 Pharmacoeconomic Challenges: Pricing, HTA, and the "Black Box" Problem
The annual treatment cost of orphan drugs in the United States averages approximately $219,000 per patient, a figure that places enormous strain on both public and private payers and increasingly attracts political attention [5]. AI-enabled population-level outcome modeling is now being incorporated into health economic submissions, with manufacturers using AI-derived real-world evidence to generate cost-effectiveness estimates where clinical trial data is insufficient for traditional decision-analytic models. However, this approach introduces a new layer of regulatory and reimbursement risk: HTA bodies in several European jurisdictions have begun challenging the methodological transparency of AI-generated cost-effectiveness models, citing the same "black box" concerns that confront clinical regulators [2, 5].
The opacity of deep learning outputs - where the pathway from input data to final prediction is mathematically opaque even to developers - is incompatible with the level of scientific transparency that regulatory and HTA reviewers require when AI evidence underpins market authorization or pricing decisions. This has created a growing premium for interpretable AI architectures, such as attention-based transformer models with explainability layers, over higher-performing but opaque ensemble methods in drug development contexts [2].
Expenditure forecasting models illustrate the scale of the payer challenge. In Italy, total pharmaceutical expenditure attributable to rare disease drugs reached €2.08 billion in 2024, with projections indicating a cumulative increase of 7.1% by 2027 driven primarily by market entry of new orphan therapies [8]. These figures underline why AI tools that can support cost-effectiveness demonstration and outcomes-based contracting are increasingly strategic commercial assets for orphan drug developers, not just scientific tools [8].
SLIDE 2 (Topic 2): Future Directions
Slide Title: Beyond the Present: AI's Evolving Role in Transforming Rare Disease Research and Therapeutic Development
2.1 Federated Learning: Privacy-Preserving AI at Scale
The single most persistent obstacle to applying AI in rare disease research is data scarcity. No individual center accumulates sufficient patients with any given ultra-rare condition to train a statistically reliable machine learning model in isolation. The structural answer to this problem - centralized data pooling across institutions and countries - is obstructed by data protection law, particularly the EU General Data Protection Regulation (GDPR), which prohibits the transfer of patient-level data outside its country of collection without explicit consent under most research scenarios [9].
Federated learning (FL) resolves this tension by re-architecting how AI models learn. Instead of bringing patient data to a central server, FL sends the model to each participating institution, where training occurs locally on that site's data. Only model weight updates - mathematical summaries of what the model learned, which contain no patient-identifiable information - are returned to a central coordinator and aggregated. The underlying patient data never leaves its origin institution [9].
For rare disease research, this architecture offers transformative potential. FL frameworks allow cross-institutional collaboration across dozens or hundreds of sites while maintaining full compliance with GDPR and analogous regulations in non-EU jurisdictions. They enable biomarker discovery, patient stratification, and treatment response prediction at a scale that would be impossible within any single center. Early applications have demonstrated particular promise in neuromuscular diseases, where patient populations are globally dispersed but aggregate numbers are sufficient to power ML models when combined via FL [9].
Active challenges include non-independently and identically distributed (non-IID) data, meaning that the clinical characteristics of patients at one site may differ systematically from those at another, degrading model generalizability. High computational and bandwidth requirements for iterative federated updates are also barriers, particularly for lower-resource institutions in lower-income countries that hold disproportionately large rare disease populations due to founder effects and consanguinity patterns. Future research must address standardized cross-border FL protocols, interoperable data schemas, and governance frameworks for multi-national FL collaborations that align with the legal requirements of each participating jurisdiction [9].
2.2 Digital Twins and Virtual Clinical Trials
A digital twin (DT) is a computationally generated, dynamically updated virtual replica of a biological system - at the level of an individual patient, a disease pathway, or a therapeutic target - that can be interrogated in silico to predict responses to candidate interventions without exposing real patients to experimental risks [10].
In rare disease contexts, DTs address the fundamental statistical problem of small sample sizes by generating virtual patient populations that capture the phenotypic heterogeneity of a real disease cohort. These synthetic populations can serve as virtual control arms in single-arm trials, replacing placebo or standard-of-care comparator groups that are ethically and logistically problematic to recruit in ultra-rare conditions. When properly validated against real-world data from patient registries, DTs substantially reduce the number of participants required for a trial to achieve adequate statistical power [10].
The FDA has described DTs as "an emerging method that could potentially be used in clinical research" in its Center for Drug Evaluation and Research AI discussion paper, and the EMA announced a technical deep-dive specifically examining DTs in both its 2024 and 2026 multi-annual AI workplans. However, a specific regulatory framework governing the use of DTs as virtual comparator arms in pivotal rare disease submissions has not yet been formally established - its development is among the highest-priority items on both agencies' AI regulatory agendas [10].
Model-informed drug development (MIDD) and model-informed precision dosing (MIPD) frameworks already adopted by the FDA and EMA provide the immediate scaffolding into which DT-generated evidence is being incorporated. AI-integrated DTs can simulate dose-concentration-response relationships across a range of pediatric and adult patient profiles simultaneously, supporting label expansions and dose optimization studies that would require years and large cohorts to conduct in traditional clinical settings [3, 4, 10].
2.3 Genome-Based Precision Medicine and ML for Rare Genetic Disorders
An estimated 80% of rare diseases have a genetic origin, making the intersection of AI with genomics a defining frontier for the field [11]. Machine learning tools applied to next-generation sequencing (NGS) outputs can now simultaneously perform variant calling, assess pathogenicity across multiple inheritance models, integrate multi-omics signatures (transcriptomics, proteomics, metabolomics), and cross-reference findings against curated rare disease knowledge graphs - compressing the diagnostic odyssey that historically averaged five to eight years for rare disease patients into a workflow measured in days [11, 12].
Deep learning architectures including convolutional neural networks and graph neural networks have demonstrated superiority over rule-based algorithms in distinguishing pathogenic from benign variants in genes associated with rare conditions, particularly for variants of uncertain significance (VUS) that constitute the majority of novel genetic findings in ultra-rare diseases. Hybrid ML models that combine supervised learning on known pathogenic variants with unsupervised anomaly detection on novel genomic patterns are emerging as the standard approach for extending diagnostic reach to conditions with no known genetic database entries [11].
Phenotype-genotype AI mapping tools integrate structured clinical descriptions - encoded using the Human Phenotype Ontology (HPO) - with multi-omics data to stratify rare disease patient subgroups, enabling biomarker-driven trial enrichment even when cohort sizes fall below 50 patients. This stratification capability is directly relevant to regulatory requirements for biomarker-defined patient populations in accelerated approval pathways [12].
The TRESOR computational framework, published in 2025, exemplifies this trajectory: by integrating genome-wide and transcriptome-wide association study data with ML-based target perturbation modeling, it generated prioritized inhibitory and activatory therapeutic target candidates for 284 diseases simultaneously, including conditions with no previously known therapeutic targets [13].
2.4 Generative AI, Large Language Models, and Synthetic Patient Data
The application of generative AI and large language models (LLMs) to rare disease research represents one of the most rapidly evolving areas in the field. LLMs trained on biomedical literature, clinical trial repositories, and patient-reported outcome datasets can now perform structured literature synthesis, extract rare disease phenotypes from unstructured clinical narratives, and identify previously unrecognized patient subpopulations within hospital information systems [6, 14].
Generative adversarial networks (GANs) and variational autoencoders trained on rare disease omics data can produce synthetic patient datasets that statistically mirror real patient cohorts, augmenting training datasets for downstream ML models without introducing privacy risk. These synthetic datasets are particularly valuable for pediatric rare disease populations, where data availability is further constrained by stringent ethical protections on research involving minors [14].
AI-powered natural language processing pipelines embedded within electronic health record (EHR) systems are already demonstrating the ability to flag patients with phenotypic profiles consistent with undiagnosed inborn errors of metabolism from routine laboratory values and clinical text, substantially increasing the completeness of rare disease registries and enabling real-world evidence generation at scale [7].
2.5 AI in Post-Market Surveillance and Pharmacovigilance
The post-authorization safety landscape for orphan drugs is structurally compromised by small patient numbers, limited prescriber familiarity with drug profiles, and under-powered spontaneous reporting systems that cannot reliably detect signals that would be obvious in larger therapeutic areas [5]. AI pharmacovigilance systems address this by integrating adverse event data from multiple sources - spontaneous reports, EHRs, claims databases, social media, and published case series - and applying ML signal detection algorithms that are substantially more sensitive than the traditional disproportionality analysis methods used by regulatory authorities [5].
Beyond signal detection, AI tools are being deployed to automate adverse event report coding, translate reports submitted in multiple languages, and prioritize case reviews by predicted clinical severity - all of which reduce the administrative burden on pharmacovigilance teams and increase the speed with which safety signals receive clinical attention. For orphan drugs where any new safety signal can materially affect the benefit-risk profile for an already small eligible population, this acceleration has direct patient safety implications [5].
2.6 International Collaboration, Ethics, and Health Equity
The technical trajectory of AI in rare disease development runs ahead of the governance infrastructure needed to ensure that its benefits are equitably distributed. Rare disease cohorts in lower- and middle-income countries - where founder effect mutations create disease clusters not represented in Northern European or North American genetic databases - are systematically excluded from AI training datasets, producing models that perform poorly or inconsistently in those populations. This algorithmic bias risk requires active mitigation through intentional dataset diversification and model validation in underrepresented populations [11, 14].
Patient advocacy organizations are increasingly embedded in AI governance structures, co-designing data collection frameworks, consenting patients for multi-use research data sharing, and participating in priority-setting decisions that determine which diseases receive AI-assisted development attention. This patient-centric governance model - increasingly mandated by both the FDA and EMA in rare disease development programs - aligns commercial incentives with genuine unmet need [6, 11].
Future international harmonization between the FDA, EMA, MHRA, PMDA (Japan), and Health Canada on shared AI credibility standards, federated data governance protocols, and DT validation requirements is widely considered the single most impactful policy lever for unlocking cross-border rare disease research at the scale AI models require to function reliably [2, 9].
References (15 total)
-
Mitra A, Tania N, Ahmed MA, et al. New Horizons of Model Informed Drug Development in Rare Diseases Drug Development. Clin Pharmacol Ther. 2024 Dec;116(6). PMID: 38989644. DOI: 10.1002/cpt.3366
-
FDA & EMA. Guiding Principles of Good AI Practice in Drug Development [Joint Publication]. January 14, 2026. Available at: mcguirewoods.com/client-resources/alerts/2026/1/fda-and-ema-provide-guiding-principles-for-ai-in-drug-development
-
Bai JPF, Stinchcomb AL, Wang J, et al. Creating a Roadmap to Quantitative Systems Pharmacology-Informed Rare Disease Drug Development: A Workshop Report. Clin Pharmacol Ther. 2024 Feb;115(2). PMID: 37984065. DOI: 10.1002/cpt.3096
-
Gangwal A, Lavecchia A. AI-Driven Drug Discovery for Rare Diseases. J Chem Inf Model. 2025 Mar 10;65(5). PMID: 39689164. DOI: 10.1021/acs.jcim.4c01966
-
Jain A, Adenwala Z. The role of artificial intelligence in pharmacovigilance for rare diseases. Expert Opin Drug Saf. 2025 Aug. PMID: 40022540. DOI: 10.1080/14740338.2025.2474645
-
Oliver Wyman Health. Optimizing Pricing and Access for Rare Disease Drugs. April 2024. Available at: oliverwyman.com/our-expertise/perspectives/health/2024/april/optimizing-pricing-and-access-for-rare-disease-drugs.html
-
Mak CM, Woo PPS, Song FE. Computer-assisted patient identification tool in inborn errors of metabolism - potential for rare disease patient registry and big data analysis. Clin Chim Acta. 2024 Jul 15;561:119803. PMID: 38879064.
-
Crea M, et al. Horizon scanning and drug expenditure for rare diseases: three-year predictive model in Italy 2025-2027. PMC12751663. Orphanet J Rare Dis [PMC]. 2025.
-
Suwer S, Ullah MS, Probul N, Maier A, Baumbach J. Privacy-by-Design with Federated Learning will drive future Rare Disease Research. J Neuromuscul Dis. 2026;13(1). PMID: 39973411. DOI: 10.1177/22143602241296276
-
[PMC Review] Advancing rare disease therapeutics through digital twins. PMC12703978. npj Digit Med / related journal. 2025-2026.
-
Abbas SR, Abbas Z, Zahir A, Lee SW. Advancing genome-based precision medicine: a review on machine learning applications for rare genetic disorders. Brief Bioinform. 2025 Jul 2;26(4). PMID: 40668553. DOI: 10.1093/bib/bbaf329
-
Choon YW, Choon YF, Nasarudin NA, et al. Artificial intelligence and database for NGS-based diagnosis in rare disease. Front Genet. 2023;14:1258083. PMID: 38371307. DOI: 10.3389/fgene.2023.1258083
-
Namba S, Iwata M, Nureki SI, et al. Therapeutic target prediction for orphan diseases integrating genome-wide and transcriptome-wide association studies. Nat Commun. 2025 Apr 18;16(1):3621. PMID: 40251160. DOI: 10.1038/s41467-025-58464-4
-
Karako K. Artificial intelligence applications in rare and intractable diseases: Advances, challenges, and future directions. Intractable Rare Dis Res. 2025;14(2):88-92. PMID: 40485885
-
[Springer Open Access] Unifying the odyssey: artificial intelligence for rare disease diagnosis and therapy. Health Technol. 2026. DOI: 10.1007/s12553-026-01057-y
Notes for Chapter Use:
- Every paragraph is written in original language synthesized from source content - no sentence reproduces the original text
- Each subsection within a slide can be converted into a text box, paragraph, or callout in your final chapter layout
- References 1-8 anchor the regulatory/economic slide; references 9-15 anchor the future directions slide, with cross-references where topics overlap
- PMID numbers are verified from live PubMed searches conducted July 21, 2026