Ship on day one, not day thirty.
Declare the model in the UI. ContextOS provisions the runtime on dedicated GPU hardware. No inference framework to install, no CUDA to configure, no batching to tune. GPU utilisation is a first-class autoscaling signal.