Why Small Clinical Datasets Need Different AI
Most AI coverage in medicine starts with scale: larger models, larger datasets, larger computing budgets. In this Global Trial Accelerators episode, Jesús E. Moreno talks with Dr. Joseph Geraci, CSO/CTO & Co-Founder at NetraMark, about a less comfortable truth. Clinical trials often do not look like the datasets that made modern AI famous.
Geraci’s argument is not that AI has little value in drug development. It is that the dominant math behind deep learning tends to reward average patterns, while real trial populations are full of small, meaningful differences. That mismatch matters because a therapy can fail on average and still work very well for a distinct subgroup of patients.
Disease labels are useful, but they are too broad
Geraci’s starting point is simple: many disease categories are clinically necessary but biologically coarse. A label such as Parkinson’s disease, major depression, or non-small cell lung cancer gathers together patients who may share a diagnosis while differing in the mechanisms driving their illness.
That becomes a trial design problem the moment a sponsor tests one drug across the whole labeled population. As Geraci puts it, “It’s kind of like trying to hit five target boards with one dart.” If the drug’s mechanism only matches one or two of those boards, the average trial result may look weak even when a real treatment effect exists.
In the conversation, he pushes the industry to think below the label level. Genetic, epigenetic, microbiome, imaging, and clinical features can all contribute to meaningful separation within what looks like one diagnosis on paper. The practical implication is not that every trial must collect every possible layer of biology. Geraci explicitly acknowledges that doing so is expensive and often unrealistic. His point is narrower and more actionable: use the data you do have to define patient groups more precisely than the label alone allows.
Why average-seeking AI struggles in trials
The most interesting part of the discussion is Geraci’s explanation of why the usual success story for AI does not transfer neatly into clinical research. Deep neural networks do very well when they can learn from massive datasets, because large numbers help the model settle around stable, average patterns.
Clinical trials are different. They are smaller, noisier, and shaped by many interacting factors: comorbidities, prior medications, physician interactions, placebo effects, and even the wider environment around the patient. In that setting, Geraci argues, the data do not form one smooth landscape. They form clumps.
That is why he says, “I don't understand this person, but I understand these people,” as a better goal for trial analytics. In other words, a useful system for clinical development may not need to explain every individual equally well. It may need to identify subsets of patients who share a coherent signal and distinguish them from others who do not.
This is also where his critique of labels becomes mathematical rather than conceptual. If the labels used for learning are already too simple, then training a model on them can harden the wrong assumptions. Instead of discovering a subgroup with a strong response, the model may blur that signal into an underwhelming average.
A different geometry for small, noisy data
Geraci describes NetraAI as an alternative learning framework built for this small-data problem. Rather than forcing the data into one average effect, he says the system reorganizes patients into substructures that better reflect how they relate to one another.
His metaphors in the episode are memorable because they make an abstract point concrete. He compares the process to a dynamic landscape where certain regions act like traps that pull similar patients together. The key idea is not the metaphor itself, but the consequence: the system searches for combinations of variables that define a subgroup with a strong and explainable response pattern.
That explainability matters. Geraci spends time on the combinatorial difficulty of finding the right variable combinations in a trial dataset. His claim is that a useful system must do more than separate responders from non-responders. It must also show which characteristics are driving that separation in a form that statisticians, sponsors, and regulators can work with.
He also places this approach alongside, not against, large language models. In his framing, one system finds the subgroup-defining structure and another system can help interpret the broader context around it. The value comes from the pairing rather than from pretending one method can do every part of the job.
The practical use case is trial strategy, not magic
The most grounded part of the episode is Geraci’s discussion of what sponsors can actually do with this kind of output. He describes feedback from the FDA that points toward a pragmatic use case: analyze a prior study, identify a model-derived subgroup, and carry that insight into the next trial through the statistical analysis plan.
His advice to sponsors is direct: “Take your big shot on goal. Try to make your drug work for as many people as possible.” But do not stop there. If you can pre-specify a subgroup that has a clearer signal, you give yourself another legitimate way to understand the result.
That framing is important because it avoids overstating the technology. Geraci is not describing a shortcut around trial rigor. He is describing a method for learning which patients may actually be benefiting, then using that information to design the next study more intelligently. In his words, “This is where it starts. Real precision medicine.”
Listen to the full conversation
If you work in clinical development, the episode is worth hearing in full because Geraci connects abstract math to a practical regulatory path without treating either one as secondary. You can listen here: Joseph Geraci, CSO/CTO & Co-Founder at NetraMark.
About Global Trial Accelerators™
Global Trial Accelerators™ is the podcast for MedTech, Biopharma and Radiopharma founders navigating first-in-human clinical trials. It is hosted by Jesús E. Moreno and produced by bioaccess®, a CRO purpose-built for first-in-human trials across the Americas.