Regulation2 min read

The Real Cost of Training AI on Untracked Data

June 2, 2026Synthesized from 1 source: AWS

As companies begin building their own AI models using internal data, a quiet compliance problem is growing: most of them cannot prove which data trained which model, and regulators in finance, insurance, and healthcare are increasingly asking exactly that question.

There is a gap forming inside most organisations that are building AI. The data team works in one system. The AI training team works in another. And somewhere in between, the record of what actually happened gets lost.

This matters now because the rules are catching up. The EU AI Act, which took full effect in 2024, can fine companies up to 7% of global annual turnover for deploying high-risk AI systems without proper documentation of training data. In US financial services, the Federal Reserve's longstanding model risk framework requires firms to document what data trained each model and who validated it. The same pressure applies to insurance, healthcare, and any industry where AI influences decisions about people or money.

The problem is structural. When a data scientist pulls data from a governed storage system like Databricks and sends it into Amazon's SageMaker training service to build a model, the governance record in Databricks does not automatically know what happened next. The model gets built. The data gets used. But the audit trail breaks at the boundary between the two platforms.

Databricks and AWS have published a technical blueprint that closes this gap. The approach keeps Databricks as the single source of truth for data permissions and history, even when the actual model training happens inside Amazon's infrastructure. When the training job finishes, it writes a lineage record back into Databricks, so the full chain is visible in one place: raw data, processed data, training job, model version.

Databricks calls this feature "bring your own data lineage," and it is currently in public preview. Over 700 companies already use Databricks Unity Catalog to manage governance across multiple tools and platforms, suggesting this kind of cross-platform record-keeping is already a priority for large enterprises.

The broader trend here is significant. For years, AI governance was treated as a technical problem for engineers. That is changing fast. Regulators increasingly expect the same standard of documentation for AI model decisions as they do for financial transactions. An auditor who asks how a credit decision or insurance risk score was generated now expects a traceable answer, not just a model file.

The companies at most risk are not the ones ignoring AI. They are the ones adopting it quickly, using whichever tools are available, and assuming the governance question can be sorted out later. It rarely can be sorted out retroactively without significant rework.

For business leaders, the practical question to ask your teams is simple: if a regulator asked us today which data trained our most important AI model, and who had authorised access to that data, could we answer that within 24 hours? If the answer is no or maybe, the gap is already there.

Stay informed

Get AI intelligence like this delivered to your inbox.


You May Also Find Valuable