AbstractPhil has released Beatrix V3, a 376 million parameter byte-level language model that utilizes "splat memory" in every attention block. The model features a vocabulary of 256 and a context window of 4,096 bytes, trained on 64.4 billion bytes over two RTX 5090 cards.

  • The architecture achieves 0.9861 bits per byte on held-out web text without loss spikes.
  • It includes detachable 13.7 million parameter adapters after every block to handle specific curriculum stages.
  • The model demonstrates strong text encoding capabilities, matching T5 performance in its middle blocks.
  • Splat memory shows limitations in long-distance recall compared to softmax controls, prompting plans for a hybrid approach.

The release includes the model weights and technical documentation, inviting feedback on architectural improvements such as hybrid attention mechanisms and forgetful memory modules.