← Back to insights
AI & Law
Auditing training data: settle provenance and licensing before launch
Training-data provenance and licensing is the thing AI products get questioned on most in fundraising and litigation.
In AI diligence or litigation, training-data provenance and licensing are often the first thing questioned. Scraped public data, licensed datasets, and user-uploaded content each have different license boundaries.
Without a clear provenance ledger, a challenge over infringement or non-compliance is hard to rebut and hard to size for exposure.
The pragmatic approach is a provenance-and-licensing ledger, filtering or replacing high-risk sources, and a data-to-model-version mapping.