If an idle endpoint stops responding, check whether max workers was scaled to zero and raise it again, or send periodic traffic to keep it warm. Fix crashing workers before raising limits, or it gets scaled down again. Load models once at module level outside the handler so every cold start is not paying model load twice.

Context: RunPod's official serverless troubleshooting docs document the silent scale-down. Endpoints with no requests for 3 days get max workers reduced to 2, and after 7 days to 0; RunPod emails on the first reduction. Endpoints with consistently unhealthy crashing workers are also scaled down automatically. For slow cold starts, the docs recommend caching models on network volumes, enabling FlashBoot, and loading models at module level rather than inside the handler function.