Policies watch queue depth, utilization, latency, or request rates and add or remove replicas. Startup time and model-loading cost require buffers or predictive scaling to avoid slow responses during bursts.
Autoscaling adjusts the number or size of serving resources in response to demand or operational signals.
Policies watch queue depth, utilization, latency, or request rates and add or remove replicas. Startup time and model-loading cost require buffers or predictive scaling to avoid slow responses during bursts.