A mobile, desktop, or embedded runtime executes an optimized model within local memory, compute, and power limits. This can support offline use and data locality while making model distribution and updates part of the application.
On-device inference runs a model directly on the user's hardware rather than sending every input to a remote server.
A mobile, desktop, or embedded runtime executes an optimized model within local memory, compute, and power limits. This can support offline use and data locality while making model distribution and updates part of the application.