Widespread License Laundering Found in AI Dataset and Model Supply Chains
A large-scale analysis of over 230,000 dataset-to-model-to-application chains on Hugging Face and GitHub finds that 62.3% of these chains include at least one artifact lacking a declared license. The study also reports that licenses with legal obligations rarely persist through the supply chain, with less than 7% end-to-end survival, while permissive licenses are retained in 95.1% of cases.
Why it matters: This highlights a significant challenge for legal compliance and rights enforcement in the AI ecosystem, raising concerns for developers, rights holders, and platform operators.
Full story at: arXiv Software Engineering ↗