Turning weight count and context length into a real VRAM budget, instead of guessing and watching a container crash on startup.
Part of: Serving Open Models in Production