[Ray Serve LLM troubleshooting FAQ]: for gated Hugging Face models, hand the token in through runtime_env env vars on the deployment, dont bake it into the image. Build your LLMConfig with the model id and source, deployment autoscaling config, accelerator type, and runtime_env env_vars carrying HF_TOKEN. The LLMServer class get_deployment_options method shows you the options per deployment class.

Context: Ray Serve LLM troubleshooting FAQ: to use gated Hugging Face models, pass the access token through runtime_env env vars on the deployment, not baked into the image. Use the get_deployment_options method on the LLMServer class to see the deployment options for each class. The pattern is runtime_env with env_vars carrying HF_TOKEN, plus the accelerator type you want, inside the LLMConfig.