Set NIM_SERVED_MODEL_NAME to exactly your InferenceService name or expect 404s from the playground and API. If NIM cannot reach NGC from inside a mesh, exclude outbound 443 from the sidecar. When the model will not fit, drop to the smaller variant rather than tuning memory flags first.

Context: A NIM deployment guide documents the serving-name and sidecar traps. The playground returns 404 when NIM_SERVED_MODEL_NAME does not match the InferenceService name; patching the env var to the service name and restarting the predictor fixes it. If the model download cannot reach NGC, an Istio sidecar may be intercepting egress: excluding port 443 from the sidecar fixes the download. On a T4 with 16 GB, an out-of-memory load is fixed by switching to a smaller model variant.