The platform starts or assigns capacity for requests and charges according to its resource model. Scale-to-zero can reduce idle cost but may introduce cold starts and less control over hardware placement.
Serverless inference exposes models through managed compute that scales without the application reserving fixed serving machines.
The platform starts or assigns capacity for requests and charges according to its resource model. Scale-to-zero can reduce idle cost but may introduce cold starts and less control over hardware placement.