Unlocking AI’s True Potential in Drug Discovery: Why Ontologies Are the Semantic Foundation

Jul 21, 2026 | Artificial Intelligence

Artificial intelligence is reshaping pharmaceutical R&D – from identifying novel therapeutic targets and predicting compound efficacy to accelerating clinical translation and enabling precision medicine. Yet, its impact is often limited by a familiar obstacle: fragmented, inconsistently annotated, and poorly structured scientific data scattered across legacy systems, lab notebooks, omics platforms, and clinical records. Without a reliable way to impose consistent meaning and relationships on this information, vast repositories of fragmented, inconsistently annotated, and poorly structured data make it difficult for AI models to generate reliable outputs or generalizable insights.

Ontologies solve this problem by providing the semantic backbone that connects human expertise with computational reasoning. They transform tacit scientific knowledge – assay definitions, phenotypes, targets, and endpoints – into explicit, machine-readable structures that support consistent interpretation, integration, and reasoning across the entire R&D pipeline.

What Are Ontologies, and Why Do They Matter?

Ontologies are formal, machine-readable models that define domain concepts, attributes, relationships, hierarchies, and logical rules. In drug discovery, they standardize what constitutes an assay, an endpoint, a biological system, or a target, and how each relates to broader biological and clinical contexts (e.g., a specific cell line relates to tissues, diseases, and species) and axioms that support logical reasoning. This enables both humans and machines to interpret data consistently.

Without this semantic layer, definition drift and inconsistent terminology undermine AI performance. For example, if you say “target” in a biology lab, your colleagues think of a protein; if you say it at an archery range, they think of the physical object they are shooting at. Imagine a scientist trying to compare efficacy data across three separate historical studies. One study recorded data as “tumor volume,” the second as “tumor size,” and the third as “neoplasm mass.” Without an underlying ontology to explicitly declare these terms as synonyms for the same concept, an AI model treats them as entirely different variables, rendering automated cross-study analysis impossible.

Ontologies eliminate ambiguity by declaring synonyms, relationships, and constraints explicitly – turning raw data into structured, reusable knowledge aligned with FAIR (Findable, Accessible, Interoperable, Reusable) data principles.

Why Ontologies Matter Even More in the Age of Agentic AI

The rise of Large Language Models (LLMs) and agentic workflows has transformed ontologies from “nicetohave” to critical infrastructure (tracked by Gartner).

Grounding your AI stack with ontologies makes a massive difference:

Modern scientific platforms are now embedding agentic frameworks directly into their data architecture. The result is a new class of systems that go far beyond traditional records management. With ontologies in place that allow AI to operate on governed, ontology-driven data rather than raw, inconsistent inputs, these new systems function as systems of scientific understanding, capable of reasoning over structured knowledge rather than merely storing it.

Building an Ontology: A Practical, Hybrid Approach

Organizations no longer need to start from scratch. The life sciences community has spent decades curating robust, open-source schemas. Successful companies prioritize a hybrid strategy – borrowing mature public schemas for standard domains and writing custom extensions only for proprietary business logic.

The public landscape is broadly split into two categories: repositories (the libraries used to find, map, and query terms) and domain-specific ontologies (the exact blueprints to copy or reuse). These frameworks provide readymade blueprints that accelerate ontology adoption and reduce the burden on internal teams.

Repositories: Where to Find and Evaluate Ontologies

Before selecting or adopting an ontology, data stewards rely on established central repositories to search for existing terms, review cross-mapped definitions, and download machine-readable formats such as OWL or RDF. These platforms provide the foundational scaffolding needed to evaluate quality, consistency, and interoperability.

  • The OBO Foundry (Open Biomedical Ontologies): The gold standard for structurally consistent biomedical ontologies. The OBO Foundry is a collective that enforces strict development principles (e.g., open sharing, unique identifiers, and non-overlapping domain boundaries ensuring that ontologies integrate cleanly without redundancy). If your goal is interoperable, well-governed data frameworks, this is the ideal starting point.
  • NCBO BioPortal: The world’s largest, most comprehensive directory of biomedical ontologies, maintained by Stanford. BioPortal hosts hundreds of models ranging from strict clinical terminologies (e.g., SNOMED-CT) to experimental research ontologies. It features excellent cross-ontology mapping tools and community annotations that reveal how terms are used in practice.
  • Ontology Lookup Service (OLS): Provided by the European Bioinformatics Institute (EMBL-EBI), OLS delivers a clean, intuitive user interface and a robust API. Engineers favor it for programmatic term queries, hierarchical browsing, and generating autocomplete tags for electronic lab notebook and structured data capture.

 

The Domain Blueprints: Specific Ontologies to Adapt

After you have identified the right repositories, the next step is selecting domain ontologies that can be pulled directly into your ecosystem. These frameworks act as foundational building blocks for a modern drug discovery pipeline, providing mature, community-validated structures that reduce ambiguity and accelerate AI readiness.

  • Gene Ontology (GO): The pioneer of the biomedical ontology space. GO is universally used to describe gene products across species, organizing knowledge into three distinct axes: molecular function, cellular component, and biological process. Any AI model that needs to reason about reaction pathways, mechanisms of action, or target biology depends on GO as a core semantic anchor.
  • Human Phenotype Ontology (HPO): Essential for translational research and clinical data harmonization. HPO provides a highly structured vocabulary of human clinical abnormalities and phenotypic features enabling AI systems to bridge the gap between early-stage genomic variants with real-world patient presentations.
  • BioAssay Ontology (BAO): The missing link for chemical biology, high-throughput screening, and assay standardization. BAO defines assay designs, perturbations, physical endpoints, and metadata dependencies. Instead of inventing internal dictionaries for experimental parameters, organizations can adopt BAO to standardize assay data from day one and ensure interoperability across labs and platforms.

 

Recommendations for Lab Scientists and Organizations

Getting started with ontologies doesn’t have to be overwhelming. At its core, ontology adoption is simply about creating a shared, structured “dictionary” for scientific data so both humans and machines interpret information consistently. It’s like organizing a messy library – once books (your data) are properly labeled, categorized, and linked by topics, finding what you need and discovering new connections becomes dramatically easier.

For drug discovery teams, the goal isn’t perfection on day one. It’s building practical definitions and relationships that reduce frustration, save time, and unlock more reliable AI tools. Here are some actionable steps to get started:

  • Start small with real problems. Focus on one or two high-impact use cases, such as harmonizing assay results or improving data searchability. Reuse ready-made public ontologies rather than building everything from scratch. This pragmatic approach delivers quick wins, builds confidence, and creates momentum for broader adoption.
  • Keep it user-friendly for bench scientists. Scientists running experiments, recording results, and making day-to-day decisions shouldn’t feel burdened by complicated new systems. Design systems with simple interfaces – such as auto-complete fields, drop-downs, or guided metadata capture that minimize extra work while enabling faster searches and more reliable data comparisons. Let the data specialists handle the complexity behind the scenes.
  • Collaborate and plan ahead. Involve scientists, data experts, and IT teams together early. Iterate definitions, validate mappings, and balance short-term value with long-term scalability. Use AI tools with human oversight to accelerate routine tasks like mapping and curation, while experts ensure the scientific accuracy and logical consistency stay intact. This approach turns data into a reliable foundation for better AI insights and faster drug discovery progress.

By taking these measured steps, organizations can reduce data chaos, improve collaboration, and build toward more powerful AI capabilities without disrupting ongoing research.

The Path Forward

Organizations aiming to turn AI into a true competitive advantage, shouldn’t wait for a perfectly curated data landscape. The right approach is to begin by identifying a concrete, high-priority use case (e.g., harmonizing preclinical assay data across internal labs and Contract Research Organizations for a single lead asset) and build it on an extensible architecture that can scale across the enterprise.

Implementing an ontology framework is not about replacing human expertise with autonomous AI. Instead, it establishes a structured, governed data foundation where scientists maintain full control over the underlying scientific logic and final decisions. With standardized, ontology-aligned data, computational tools can efficiently automate repetitive processing tasks, accelerate analysis, and enhance the entire research pipeline. The result is a system where human judgment and AI efficiency reinforce each other, enabling faster, more trustworthy scientific progress.

Conclusion: Building Trustworthy AI on Solid Ground

Ontologies represent a fundamental shift in how the industry treats data, not as a byproduct of research, but as a strategic asset. The field is moving from data chaos to structured, AI-ready knowledge systems that support reliable, scalable scientific insight. Ontologies are not a luxury or academic exercise; they are the pragmatic foundation that enables AI to move beyond hype and deliver trustworthy results in drug discovery. By making implicit knowledge explicit, unifying disparate datasets, and supporting sophisticated reasoning, ontologies unlock interoperability, reduce errors, accelerate innovation, and ultimately contribute to better therapies for patients. Organizations that invest now – through targeted pilots, cultural alignment, and strategic partnerships – will be the ones that gain a durable competitive advantage in the AI-driven era.

Mary Beth Walsh
About The Author

Chris Marth

Chris Marth, is a Principal Consultant for Kalleid, Inc.  With degrees in chemistry and education, Chris spent time working in the lab before becoming a veteran classroom trainer who then transitioned to developing eLearning and has never looked back. She has spent more than 20 years’ creating highly interactive eLearning experiences for clients ranging from Fortune 500 companies to technology startups to grant-funded research initiatives. Chris is passionate about helping organizations connect learning strategy to design, development, data, and ultimately performance. She loves designing course content and media and test-driving new learning technology.

About Kalleid

Kalleid, Inc. is a boutique IT consulting firm that has served the scientific community since 2014. We work across the value chain in R&D, clinical, and quality areas to deliver support services for software implementations in highly complex, multi-site organizations. At Kalleid, we understand how effective project management plays a key role in ensuring the success of your IT projects. Kalleid project managers have the right mix of technical know-how, domain knowledge and soft skills to effectively manage your project over its full lifecycle. From project planning to go-live, our skilled PMs will identify and apply the most effective methodology (e.g., agile, waterfall, or hybrid) for successful delivery. If you are interested in exploring how Kalleid project managers can benefit your organization, please don’t hesitate to contact us today.