QuantVectors is seeking annotated document datasets in Indic languages from India, including Hindi, Marathi, Gujarati, Bengali, Punjabi, Tamil, Urdu, Telugu, Odia, Kannada, Malayalam, and Assamese. The datasets must include invoice, receipt, utility bill, payment advice, packing list, commercial invoice, and credit note types, with approximately 400 documents per language, human-verified annotations, and 99%+ accuracy. Datasets must be commercially licensable and can be open-source or commercial, with a request for HuggingFace datasets, research datasets, or vendors specializing in this space.
Seeking Indic Document Datasets for AI/OCR Training in India
vldmrbesk releases 51k-sheet Tsiolkovsky archive with measured recognition accuracy
A complete personal archive of Konstantin Tsiolkovsky, comprising 51,008 sheets and 2,019 archival files, has been released as a CC0 dataset. The collection spans fifty years of work by the self-taught rocket theorist and includes handwritten manuscripts alongside typed copies.
Hugging Face releases The Stack v3, the largest open code dataset
Hugging Face has released The Stack v3, described as the largest open code dataset to date. The release provides two distinct access methods to accommodate different user needs.
IBM open-sources CodeAlchemy, a massive synthetic dataset of high-quality code
IBM has open-sourced CodeAlchemy, a pipeline and synthetic dataset designed to improve AI model performance by pairing code with execution traces. The release includes nearly 1 trillion tokens across 15 programming languages, totaling at least 200 times the size of Wikipedia.
MultiSynt/MT releases 4.8T-token parallel corpus across 36 languages
Researchers introduce MultiSynt/MT, an open synthetic parallel corpus containing approximately 4.8 trillion target-language tokens across 36 European languages. The dataset is generated by translating 100 billion high-quality Nemotron-CC tokens using Tower+ and OPUS-MT/HPLT-MT systems.
Current AI launches Open Source AI Gap Map v0.1 with 421 indexed products
Current AI, a non-profit founded at the AI Action Summit in February 2025, has launched the Open Source AI Gap Map v0.1 to index the state of open source AI.