Study Finds Widespread License Drift in Open-Source AI Models and Applications
A new arXiv preprint audits licenses for over 364,000 datasets and 1.6 million models on Hugging Face, along with 140,000 GitHub projects, uncovering systemic non-compliance: 35.5% of model-to-application transitions relicense under more permissive terms, removing restrictive clauses. The authors also introduce a rule engine that detects 86.4% of license conflicts in software applications, and release both their dataset and tool for further research.
Why it matters: The findings highlight significant legal and ethical risks in the open-source AI ecosystem, underscoring the need for automated license compliance tools.
Full story at: arXiv Software Engineering ↗