Together AI ने Kubernetes पर Together GPU Clusters के लिए preemptible compute की सार्वजनिक पूर्वावलोकन की घोषणा की है, जो interruption-tolerant workloads के लिए कम लागत वाला विकल्प प्रदान करता है। Preemptible nodes unused NVIDIA accelerated capacity का उपयोग करते हैं और on-demand दर का 50% फ्लैट शुल्क लिया जाता है।
- जब क्षमता कहीं और आवश्यक होती है तो नोड्स को वापस ले लिया जा सकता है, जिससे पाँच मिनट की drain sequence शुरू होती है जहाँ pods को SIGTERM प्राप्त होता है checkpoint बनाने और बाहर निकलने के लिए।
- बिलिंग sub-hourly है, उपयोग हर एक से दो मिनट में मापा जाता है।
- Preemptible nodes label के माध्यम से मौजूदा clusters में जुड़ते हैं, जिससे महत्वपूर्ण घटक standard nodes पर रह सकते हैं।
- यह सुविधा छोटे experiments, ablations और अस्थायी inference bursts का समर्थन करती है जो checkpoints से फिर से शुरू हो सकते हैं या interruption के बाद retry कर सकते हैं।
इससे टीमों को batch jobs और config sweeps जैसे non-critical workloads के लिए spare capacity का लाभ उठाने की अनुमति मिलती है, जबकि user-facing replicas guaranteed standard infrastructure पर रहते हैं।