FitCheck is a new tool that estimates peak VRAM requirements for LoRA, QLoRA, and full fine-tuning of large language models without requiring PyTorch, CUDA, or actual GPU hardware.
- It analyzes the model's Hugging Face config.json and parameter-count metadata to report memory breakdowns, usable capacity, headroom, and fit verdicts.
- The tool supports serving estimates based on model weights and KV cache, along with an advisor for exploring batch size, sequence length, and LoRA rank.
- Testing on a Tesla T4 showed a 2.4% mean absolute error across 57 calibration runs and produced 12/12 correct fit-boundary verdicts in a separate holdout test.
The tool helps users determine if their LLM configuration will fit in GPU memory before starting a job, aiming to reduce guessing and out-of-memory errors.