Scalable Extraction and Normalization of Biomedical Knowledge from Research Literature

Open Access
Article
Conference Proceedings
Authors: Supraja KrovvidiFinn VosLaurent HassonRamesh JakkaShloke Meresh
Abstract

The growth of biomedical literature poses significant challenges for researchers conducting systematic and scoping reviews. In fields such as the use of digital biomarkers for the treatment of heart failure and cardiovascular disease, manually screening thousands of papers is time-consuming and not scalable. To address this problem, we developed AI-assisted tools for large scale analysis and structured knowledge extraction from biomedical research articles. Our approach emphasizes cross-graph analytics, facilitating the exploration of relationships among key biomedical concepts, including digital biomarkers and digital health technologies. By automating the extraction and structuring of knowledge from unstructured text, our system aims to accelerate evidence of synthesis and support more comprehensive and up-to-date reviews in rapidly evolving biomedical domains.We developed a knowledge graph generation pipeline that extracts subject-predicate-object triplets representing scientific claims from research articles. To address redundancy caused by linguistic variation across documents, we implemented a multi-stage consolidation process focused on normalizing entities and relations. This process begins by validating and filtering extracted triplets, then applies lexical normalization to unify entity representations by removing stop words, resolving variants, and merging acronyms with their full forms. Entity types and relations are similarly standardized to ensure uniformity and clarity. The pipeline is designed to be modular and extensible, allowing for the integration of additional normalization strategies or domain-specific ontologies as needed.To further consolidate equivalent triplets, we leverage embedding-based semantic similarity, enabling the merging of semantically similar entities and relationships even when expressed differently across sources. Additionally, our pipeline utilizes biomedical ontologies such as RxNorm and MeSH to map entities to standardized concept identifiers. This ontology-based normalization ensures that references to the same biomedical concept are unified, regardless of linguistic or spelling differences. We evaluated our approach on a pilot corpus of 150 biomedical research articles, processing over 2,000 extracted triplets. The normalization pipeline reduced the number of unique entity variants by more than 50%, consolidating these into approximately 900 unique, semantically unified relationships. Manual review of a representative sample indicated entity normalization accuracy in the range of 90-95%. In conclusion, the integration of lexical, semantic, and ontology-based normalization offers a robust framework for reducing ambiguity and improving the interoperability of the resulting knowledge graphs. Moreover, this structured and unified representation of knowledge facilitates systematic reviews, meta-analyses, and data-driven decision-making in biomedical science, while enabling advanced querying, trend analysis, and the identification of novel associations between biomedical concepts

Keywords: knowledge graphs, AI, LLMs, normalization, entities, relations, biomedical, digital health technology, digital biomarker

DOI: 10.54941/ahfe1008186

Cite this paper
Downloads
0
Visits
4

More from this volume

Developing Command Capability Throughout the Career: A Competency-Based Model for Leadership and Communication in the NavyEvaluation of Hand Sewing Skills for the Design of a Clothing Repair Education Program: Pilot Study
View all articles in Human Systems Engineering and Design (IHSED2026): Future Trends and Applications