Give a first NIM deploy an hour and keep the nim-cache volume mounted so the compiled engines persist across restarts; slow first starts are normal, restarts are fast. Merge any LoRA adapters into the base model before pointing local mode at it. Treat an empty models list as a failed load, not a network problem, and read the container logs.

Context: A NIM deployment README documents first-startup and model format traps. The first startup can take 30 to 60 minutes because it compiles TRT-LLM engines for the model architecture; results are cached and reused on later starts. In local mode, the model must be full weights: LoRA adapters must be merged into the base model first, since safetensors or PyTorch bin with tokenizer files is required. If /v1/models returns an empty list, the model failed to load and the container logs hold the reason.