Why Better Data Curation Can Shorten Biomedical R&D
In this Global Trial Accelerators episode, Jesús E. Moreno speaks with Pablo Yarza, founder of BIOINFILE, about a problem that sits underneath nearly every biomedical program: teams cannot move faster than the data they can actually trust.
That sounds obvious, but the conversation makes a sharper point. The issue is not just that data is messy. It is that useful scientific data often arrives late, incomplete, poorly structured for the task at hand, or already out of date by the time a team starts using it. For founders and R&D leaders, that turns data management from a background task into a real source of cost, delay, and technical risk.
Curation is not admin work. It is part of the product.
Yarza’s career started in academic bioinformatics, where he worked on reference sequence databases for microbial identification. His description of that work helps explain why curation matters so much. Primary repositories can grow quickly, but growth alone does not make them easy to use. Someone still has to decide which sequences best represent a species, check quality, and organize the data into something other scientists can rely on.
As he puts it, “Data curation, which is so valuable, is the bottleneck.” That line matters because it reframes curation. It is not clerical cleanup after the real science is done. It is the work that turns raw information into a usable reference.
Yarza also notes a frustrating reality from his academic experience: even a carefully built database could take about a year to release, and by then parts of it were already old because the underlying science had kept moving. In other words, the value of curation is high, but the maintenance burden is constant. A dataset is not finished just because it has been published once.
For companies, this matters when they assume publicly available scientific data is ready for immediate operational use. Yarza’s answer is more cautious. Academic data can be valuable, but it often needs another layer of work before it fits an industrial setting with quality expectations, traceability needs, and time-sensitive development goals.
Industry does not just need data. It needs current data.
One of the most practical parts of the episode is Yarza’s warning about data obsolescence. He describes a diagnostic lab using a commercial informatics platform whose reference databases were 10 years old. The brand and workflow were accepted, but the reference layer underneath had aged.
His broader argument is that secondary datasets age faster than many teams admit. Labels change. New references appear. Standards shift. New evidence forces reclassification. That means a dataset that once looked solid can quietly become a weak point in a pipeline.
Yarza says these reviews should happen at least annually. That is a useful benchmark for any team building tools on top of reference data, whether in genomics, transcriptomics, microbiome work, or diagnostic development. If the state of the art changes every year, the foundation has to change with it.
This is also where his distinction between primary and secondary data becomes useful. Primary data may remain intact as a record, but the curated datasets built from it need ongoing attention. Teams that ignore that difference risk treating a living knowledge layer as if it were a static file.
AI can help, but it cannot rescue bad inputs
The episode does not dismiss AI. Yarza says BIOINFILE uses AI and large language models in its own processes for tasks such as summarization and classification. But he is clear about where the limit is.
“What is more important here is to check the quality of the input and the quality of the output.” That standard applies whether a team is building a biomarker model, searching for external cohort data, or assembling a scientific reference package for development work.
This is a useful corrective to the common idea that better models will somehow compensate for weaker source material. Yarza’s view is simpler: AI sits downstream from data quality. If the inputs are outdated, poorly matched to the use case, or badly curated, the system may run faster, but it will not become more reliable.
That point is especially relevant in biotech, where validation still decides whether a result is useful. New tools may speed retrieval, structuring, or comparison. They do not remove the need to verify what goes in and what comes out.
The hidden cost is not only money. It is time from the wrong people.
When Moreno asks about economics, Yarza gives one of the clearest figures in the conversation. He estimates that the data curation part of a project can cost about 50,000 euros, largely in employee time spent searching for, cleaning, and structuring public scientific data.
That number is only part of the story. The larger cost may be who is doing the work. Scientists, data specialists, and R&D staff can end up spending months on data acquisition and harmonization instead of advancing the device, assay, or molecule itself. Yarza argues that this slows the path to market and delays the point at which a company can start generating its own private data.
His proposed answer is focused outsourcing: let specialized groups handle the public-data acquisition and curation layer so internal teams can concentrate on proprietary development. Whether a company buys that service from BIOINFILE or another provider, the principle is the same. The faster a team can get fit-for-purpose reference data, the faster it can move to the work that actually differentiates the business.
The conversation ends by widening the lens beyond one company. Yarza argues that academia still holds much of the expertise needed to maintain high-quality datasets. “The ones who can curate the data, maintain the quality of the data sets are the researchers, are the professors.” His thesis is that knowledge transfer from academia to industry is still too slow, and better data infrastructure should help close that gap.
Listen to the full episode
If your team works with omics data, external cohorts, or scientific reference datasets, this episode is worth your time. You can listen to the full conversation with Pablo Yarza on the episode page.
About Global Trial Accelerators™
Global Trial Accelerators™ is the podcast for MedTech, Biopharma and Radiopharma founders navigating first-in-human clinical trials. It is hosted by Jesús E. Moreno and produced by bioaccess®, a CRO purpose-built for first-in-human trials across the Americas.