Configure autoscaling around in-flight tokens, not concurrency, and keep at least one replica alive since scale-to-zero is unsupported. Match engine_config field names to the engine you actually run, vLLM names for vLLM and TRT-LLM names for TRT-LLM, or the deploy fails at startup. For cache-sensitive workloads, keep max_scale_down_rate low and scale_down_delay in the 300 to 600 second range so replicas drain gently and your time-to-first-token stays stable.

Context: A Baseten usage reference documents the autoscaling and config gotchas that cause failed or slow deploys. LLM serving autoscales on in-flight tokens, not request concurrency, so setting concurrency_target is rejected and only target_in_flight_tokens applies, with 50K to 150K a sane starting range. Scale-to-zero is not supported, so min_replica must be at least 1. engine_config field names are engine-native: vLLM uses max_num_seqs while TRT-LLM uses max_batch_size, and mixing the two fails at startup. Scale-down also erodes KV cache, causing TTFT spikes unless max_scale_down_rate stays low.