Hugging Face has released The Stack v3, described as the largest open code dataset to date. The release provides two distinct access methods to accommodate different user needs.
- stack-v3-train: A near-deduplicated, quality-filtered, and PII-redacted version with inline contents, accessible via the Hugging Face datasets library.
- stack-v3-full: The entire 114 TB corpus stored as an HF Storage Bucket, preserving all duplicates with cluster IDs and stubs for excluded files.
This release offers researchers and developers access to a massive, raw code repository while also providing a pre-processed variant for immediate training use.