Quinex: Quantitative information extraction from text using open and lightweight LLMs

ARTICLE Open Access Download: PDF

Public summary

* Quinex, an efficient and domain-agnostic tool for extracting quantitative data from text, is presented.

* Quinex facilitates large-scale literature analyses and trend monitoring for a broad audience.

* Small, fine-tuned language models achieve state-of-the-art results for lightweight systems.

* To achieve this, significantly larger training datasets spanning diverse research fields were created.

* Qualitative improvements include consideration of implicit properties and detailed, critical qualifiers.


Abstract

Extracting quantitative data from the growing body of scientific literature is a challenge central to modern research across disciplines. While recent advances in large language models have significantly facilitated automation of this traditionally time-consuming task, their computational demands limit scalability and accessibility. Smaller specialized systems offer reduced computational requirements but sacrifice accuracy, domain generalization, or the extraction of contextual details. To address these gaps, we present Quinex, a domain-agnostic framework for quantitative information extraction based on comparably small language models. Quinex identifies quantities and their associated entities, properties, and qualifiers using a multi-turn question-answering approach. By addressing the data bottleneck, considering implicit properties, and optimizing extraction order and question templates, Quinex achieves state-of-the-art F1 scores of over 98% for quantities, 82% for entities, and 87% for properties across diverse scientific genres. It goes beyond identification by normalizing units, aligning them with a unit ontology, and detailing critical qualifiers such as references, spatiotemporal scopes, and determination methods. By reducing the effort required to extract structured quantitative data from texts, Quinex enables transformative applications, including automated literature reviews, quantitative search, and trend monitoring, and sets a new benchmark for scalable, accurate, and domain-agnostic quantitative information extraction.




Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(185) Cited by(0)

Relative Articles