The application packages an optimized model and runtime within device limits for memory, power, and compute. Local execution can reduce network dependence and data transfer, but updating and monitoring many devices is harder.
Edge inference runs a model near the data source on a device such as a phone, sensor, vehicle, or local gateway.
The application packages an optimized model and runtime within device limits for memory, power, and compute. Local execution can reduce network dependence and data transfer, but updating and monitoring many devices is harder.