The group behind the Eunoia model, built on a Llama-3.1-Dolphin-3.0-8B base, has open-sourced "The Forge," a specialized architectural pipeline and algorithmic fine-tuning strategy under the CC0 1.0 Universal license.

  • The approach prioritizes structural crystallization over unbounded context windows to avoid memory saturation and informational noise.
  • The raw code includes optimizations for bare-metal CPU clusters, such as disabling gradient clipping and using native FP32 caching to prevent RAM fragmentation.
  • A custom dataset class enforces a strict linear sequence order to preserve the contextual structure of historical data during training.
  • The release includes the AGIO-CIMIENTO-dataset.txt, over 500,000 words of evaluation tracking logs, and raw GPQA Diamond evaluation results to verify general capability retention.

The authors publish the experiment from its start to encourage independent reproduction and testing on faster hardware, noting that the final mathematical results will only be known after several days of training.