Koshur Pixel introduces a synthetic OCR dataset with 613,078 image-text pairs generated from the KS-PRET-5M corpus using SynthOCR-Gen. It includes over 25 augmentation strategies and spans diverse fonts and textual scales, from words to full-page documents, enabling scalable training for Kashmiri OCR systems.
Koshur Pixel: First Large-Scale Synthetic OCR Dataset for Kashmiri
vldmrbesk releases 51k-sheet Tsiolkovsky archive with measured recognition accuracy
A complete personal archive of Konstantin Tsiolkovsky, comprising 51,008 sheets and 2,019 archival files, has been released as a CC0 dataset. The collection spans fifty years of work by the self-taught rocket theorist and includes handwritten manuscripts alongside typed copies.
Hugging Face releases The Stack v3, the largest open code dataset
Hugging Face has released The Stack v3, described as the largest open code dataset to date. The release provides two distinct access methods to accommodate different user needs.
IBM open-sources CodeAlchemy, a massive synthetic dataset of high-quality code
IBM has open-sourced CodeAlchemy, a pipeline and synthetic dataset designed to improve AI model performance by pairing code with execution traces. The release includes nearly 1 trillion tokens across 15 programming languages, totaling at least 200 times the size of Wikipedia.
MultiSynt/MT releases 4.8T-token parallel corpus across 36 languages
Researchers introduce MultiSynt/MT, an open synthetic parallel corpus containing approximately 4.8 trillion target-language tokens across 36 European languages. The dataset is generated by translating 100 billion high-quality Nemotron-CC tokens using Tower+ and OPUS-MT/HPLT-MT systems.
Current AI launches Open Source AI Gap Map v0.1 with 421 indexed products
Current AI, a non-profit founded at the AI Action Summit in February 2025, has launched the Open Source AI Gap Map v0.1 to index the state of open source AI.