- Strong instruction following, tool use, sub-agent transfers.
- Better accuracy than OSS models, at higher cost per token.
- Speed depends heavily on host — specialized fast hosts and colocated inference win.
- Smaller variants trade cost/latency vs quality — pick size after measuring.
- Livekit Gemma4 is 200ms TTFT
Connections
- part_of Large Language Model
- tradeoff_vs GPT-OSS — Better accuracy vs cheaper tokens.
- tradeoff_vs Minimax — Reliable tools/transfers vs unstable tool loops despite fine TTFT.
- measured_by Time to First Token