Together has released a significant update to its inference platform, introducing "Dedicated Model Inference" designed to give users complete control over performance, cost, and quality for open-weight models. The update also announces a closed beta for custom training capabilities, including full-weight and LoRA reinforcement learning.

  • Deployments support canary, blue-green, and rolling updates with automatic rollback on threshold breaches.
  • A/B and shadow testing allow evaluation of new weights against real traffic without impacting users.
  • Model caching and distribution layers were rebuilt to deliver roughly 4× faster warm starts.
  • Autoscaling supports metrics such as inflight requests, GPU utilization, and time to first token.
  • Custom training beta includes supervised fine-tuning with checkpoints deployable directly to production.

The platform aims to simplify the transition from experimentation to production by handling deployment lifecycle management and observability, allowing teams to safely iterate on model versions weekly.