- Reduces precision to shrink memory use and speed up inference on hardware that supports the chosen format.
- The actual win depends on the runtime and GPU. A smaller checkpoint does not automatically mean lower latency.
- Test quality per component. STT may lose rare words, an LLM may lose tool reliability, and TTS may introduce audible artefacts even when aggregate benchmarks barely move.
Connections
- part_of Inference Optimization