• Reduces precision to shrink memory use and speed up inference on hardware that supports the chosen format.
  • The actual win depends on the runtime and GPU. A smaller checkpoint does not automatically mean lower latency.
  • Test quality per component. STT may lose rare words, an LLM may lose tool reliability, and TTS may introduce audible artefacts even when aggregate benchmarks barely move.

Connections