Clinical trials are one of the most expensive processes in any industry. Developing a single drug takes somewhere between ten and fifteen years and over five billion dollars, and only a small fraction of candidates ever make it to market. The part that rarely gets discussed outside pharma is how much of that cost is caused not by science failing, but by the wrong hospitals being picked to run the trial.
The numbers have not changed in decades. Around 37% of activated trial sites enroll fewer patients than they promised. Another 11% enroll nobody at all. The combined effect is that more than half of all trials run over their planned timeline, with one in six taking more than twice as long as expected. Each extra day of delay costs a drug sponsor around $500,000 in unrealized sales from the delayed medicine, plus direct daily running costs on top of that.
Pharma has known about this problem for a long time. The answer so far has been to buy scoring tools from specialist vendors or rely on the analytics platforms built by contract research organizations, the companies hired to manage trials on a sponsor's behalf. Those tools are built on aggregated industry data, meaning they score a site based on how similar sites have performed across the whole industry. That is a reasonable starting point but a poor predictor. A hospital that performed brilliantly on a cardiovascular trial may be completely wrong for a rare oncology study run under a different protocol.
Databricks has now released a free, open-source tool called the Site Feasibility Workbench that works differently. It trains its prediction models on the sponsoring company's own internal records: the actual enrollment rates from their previous trials, the screen failure patterns at specific sites, the qualification history, the amendment data. The more trials a company has run, the better the predictions get. Unlike a licensed vendor tool, there is no static version. Each new study adds to the training data.
The regulatory dimension is increasingly important here. The FDA, alongside international regulators, has been building out formal guidance on how AI models used in drug development should behave. The core expectation is that any AI-driven recommendation must be explainable. A regulator or data monitoring committee should be able to ask why a particular site was selected or excluded and receive a documented, traceable answer, not a vendor's black-box score. The Workbench stores the reasoning behind every site recommendation in a way that can be reviewed and audited like any other piece of trial documentation.
There is also a diversity angle that is becoming a compliance requirement rather than an aspiration. Under legislation passed in the United States in 2022, drug sponsors must now demonstrate upfront how they plan to achieve representative trial populations. The Workbench builds diversity as a scored dimension in site selection, not an afterthought. It can also be audited to check whether community hospitals and minority-serving institutions are being systematically underweighted by the model.
The broader market context matters here. The global market for AI in clinical trials is currently valued at around 2.7 billion dollars and is projected to grow at roughly 25% per year through 2030. But most of that investment is going into point solutions that still live outside the core data infrastructure. The pattern playing out across pharma is similar to what happened in financial services a decade ago: companies building many disconnected tools rather than changing the architecture underneath.
What Databricks is doing with this release is essentially arguing that the architecture is the product. The tool itself is free and open source. The commercial bet is that pharma organizations running on Databricks infrastructure will find it easier to use this than to integrate yet another external vendor tool, and will deepen their reliance on the platform as a result.
For operations leaders at pharma companies and biotech firms, the practical question is simpler than it sounds. If a company has years of trial history sitting in internal systems that never get used to inform the next site selection decision, that is institutional knowledge being wasted on every new study. Tools like this exist to close that gap. Whether this specific tool is the one that does it, or a competitor builds something equivalent, the direction is clear: site selection decisions are moving toward data-driven models trained on internal history, and companies that build that capability early will run faster, cheaper, and more compliant trials than those that do not.