A monthly-refreshed compliance matrix has been published for all 4,893 Indic-language datasets on the Hugging Face Hub, revealing that 65.1% declare no license tag as of August 1, 2026.

  • 3,185 of the datasets have no license tag, accounting for 46.1% of all downloads.
  • 154 datasets use CC-BY-NC licenses, posing risks for commercial training pipelines.
  • 139 datasets carry ambiguous tags such as 'other', 'cc', or 'unknown'.
  • The full data is available in CSV and Parquet formats under CC0 license.

The matrix helps teams avoid licensing traps by identifying datasets that are effectively disqualified from commercial use under EU AI Act GPAI documentation duties.