A complete personal archive of Konstantin Tsiolkovsky, comprising 51,008 sheets and 2,019 archival files, has been released as a CC0 dataset. The collection spans fifty years of work by the self-taught rocket theorist and includes handwritten manuscripts alongside typed copies.

  • The archive contains 36,446 struck-out passages marked for authorial revision and 299,939 explicit uncertainty marks from the transcription model.
  • Recognition accuracy is measured using 1,759 pairs of manuscript and typescript readings from 224 files, revealing a median word agreement of only 37%.
  • Against printed editions, character agreement reaches 98.1% on typescripts and 81.1% on handwriting, providing a concrete benchmark for handwritten Russian recognition.
  • The dataset serves as a testbed for uncertainty calibration and HTR model fine-tuning, addressing the lack of existing Russian handwriting resources.

The release provides a rare opportunity to study one hand aging over time and historical orthographic changes, while offering a rigorous evaluation metric for OCR systems on difficult historical scripts.