Swiss AI has released Apertus 1.5, a family of 8B and 70B parameter language models that introduce multimodal capabilities and a new reasoning feature. The update involves continued pretraining on an additional multimodal mix of 4T tokens for the 8B model and 2T tokens for the 70B model.
- Supports multimodal inputs including images, audio, and text to generate text responses.
- Introduces a "thinking mode" that allows the models to reason on input before generating responses.
- Increases default context length to 262,144 tokens, a four-fold increase over Apertus 1.0.
- Enhances instruction-following and tool-use capabilities through an improved post-training recipe.
- Maintains fully open weights, data, and training details while supporting a wide range of languages.
The release aims to advance the state of multilingual, multimodal, and transparent AI by providing developers with versatile interaction options and improved performance on reasoning tasks.