This document provides an organized comparison of GPU architectures, deployment platforms, LLMs, and speech models (TTS/STT) relevant for deploying a voice agent. It also includes pricing and feature comparisons across popular GPU cloud providers.
GPU Architecture Comparison
| Architecture | Release Year | Example GPUs | Tensor Core Gen | Special Precision | Approx. Tensor TFLOPS | Other Notable Features |
|---|---|---|---|---|---|---|
| Turing | 2018 | T4, RTX 20 series | 1st Gen | FP16, INT8 | ~65 TFLOPS (T4, FP16) | First Tensor Cores, RT cores for RTX |
| Ampere | 2020 | A100, A10G, RTX 30 series | 3rd Gen | TF32, FP16, BF16, INT8 | ~312 TFLOPS (A100, FP16) | TF32 precision mode for training |
| Ada Lovelace | 2022/23 | L4, L40S, RTX 40 series | 4th Gen | FP16, BF16, INT8, sparsity | ~180 TFLOPS (L40S, FP16) | Improved efficiency, higher clocks |
| Hopper | 2022 | H100, H200 | Transformer Engine | FP8, FP16, BF16, INT8 | ~1000 TFLOPS (H100, FP8) | Dynamic mixed-precision (FP8) |
| Blackwell | 2024+ | B200 | Next Transformer Engine | FP4, FP8, FP16, INT8 | ~2000 TFLOPS (FP4/FP8 est.) | Lower precision for LLMs |
GPU Cloud Platforms Comparison
| Feature | Lightning AI | Modal | Koyeb | RunPod | Fal AI | Cerebrium AI |
|---|---|---|---|---|---|---|
| GPU Types | T4, L4, L40S, A100, H100, H200, B200 | T4, L4, A10G, A100 (40/80GB), L40S, H100, H200, B200 | RTX 4000, L4, A6000, L40S, A100, H100, Tenstorrent N300 | A4000, A5000, A6000, A100, H100, H200, B200, RTX 3090/4090 | A6000, A100, H100, H200, B200 | T4, L4, A10, A100, H100, H200, Trainium, Inferentia |
| Monthly Fees | 50 free GPU hours, 30) | $250 compute/month | $29/mo | — | — | — |
| Pricing (T4/A100/H100) | T4: 2.71/hr, H100: $5.52/hr | T4: 2.50/hr, H100: $3.95/hr | T4-like: 2/hr, H100: $3.3/hr | A100(80GB): 4.47/hr | A100(40GB): 1.89/hr, H200: $2.10/hr | T4: 2.50/hr, H100: $3.95/hr |
| Storage | Free: 50GB, Pro: 200GB | No direct charge | $0.50/GB/mo | $0.05–0.20/GB/mo | Ephemeral (7-day) | $0.05/GB/mo |
| Cold Start | 78–537s | 1–4s | <200ms | ~200ms–12s | ”Near-instant” (benchmarks: 222–537s) | ~2s (GPU) |
| Docker/Containerization | Custom Docker, autoscaling | Custom Docker, 1s spin-up | GitHub deploys | Full Docker, serverless | fal deploy or Docker | Custom Docker |
| VPC / Networking | Enterprise VPC | Internal mesh only | No external VPC | Internal VPC | No external VPC | Region isolation |
| Security & Compliance | SOC2, HIPAA | HIPAA (BAA) | Private mesh | SOC2/HIPAA (in progress) | Custom licenses | SOC2, HIPAA |
Large Language Models (LLMs)
| Model | Parameters | GPU | Other Info |
|---|---|---|---|
| Llama 3.1 8B Instruct | 8B | ~16GB | 128k context length, JSON tool calls supported |
Domain Models (Medical)
| Model | Parameters | GPU | Benchmark / Notes |
|---|---|---|---|
| OpenBioLLM | 70B, 7B–13B | A100 80GB | SOTA results, outperforms GPT-4/Gemini/Med-PaLM-2; Llama 3 License (commercial use under 700M MAU) |
| Meditron | 70B | — | 70.2% accuracy on USMLE-style Qs, better than GPT-3.5 |
| Me-LLaMA | 70B | — | Beats ChatGPT on 7/8 datasets, GPT-4 on 5/8 |
Text-to-Speech (TTS)
| Model | Parameters | GPU Memory | Quantization |
|---|---|---|---|
| Kokoro | 82M | 326MB | FP32 |
Speech-to-Text (STT)
| Model | Parameters | GPU Memory | Quantization |
|---|---|---|---|
| Whisper-Medium | 769M | ~5GB | FP16 |
Deployment Cost Estimation
| Provider | Model / Setup | Cost per Hour | 30-Day 24/7 Estimate |
|---|---|---|---|
| Lightning AI | H100 (80GB) – 26 CPU / 1 GPU, 1513 TPS | $2.70/hr | $1944 |
| Lightning AI | A100 (80GB) – 30 CPU / 312 TPS | $1.55/hr | — |
| Together AI | Llama | $0.88 per 1M tokens | — |
| Together AI | Whisper | $0.09/hr | — |
| Cluster (Kubernetes) | H100 (80GB, 208 CPU) | $3.19/hr | — |
| Dedicated Deployment | — | $27/hr | — |
| Modal | A100 (40GB) | $2.10/hr | — |
| Modal | H100 | 0.0473 CPU/hr | — |
🌐 Enterprise GPU Hosting Summary
| Provider | GPU Options | H100 $/hr | A100 $/hr | T4 $/hr | REST API | WebSocket | Docker | Autoscaling | Serverless GPU | Scale-to-Zero | VPC | Billing | Free Tier | Enterprise Features |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Lightning AI | H100, A100, L40S, T4, A10G | $0.42+ | $0.42+ | $0.42+ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | Per-second | 50 GPU-hrs | SOC2, HIPAA |
| Cerebrium AI | H100, H200, A100, L40S, L4 | $4.68 | $2.48 | $0.58 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | Per-second | $30 credit | SOC2, HIPAA |
| Modal | B200, H200, H100, A100, L40S | $3.95 | $2.50 | $0.58 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | Per-second | $30 credit | HIPAA, SSO |
| Koyeb | H100, A100, L40S, A6000 | $3.30 | $2.00 | — | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | Per-second | $10 credit | ISO27001 |
| RunPod | H100–T4 | 4.47 | 2.72 | 0.58 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | Per-second | Credits | Enterprise |
| Fal AI | H100, A100, A6000, B200 | $1.89 | $0.99 | — | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | Per-second | Credits | Enterprise |
| Replicate | H100, A100, L40S, T4 | $5.49 | $5.04 | $0.81 | ✅ | ❌ | ✅ | ✅ | ✅ | ✅ | ❌ | Per-second | Credits | Enterprise |
| Anyscale | H100, A100, V100, T4 | Contact | Contact | Contact | ✅ | ❌ | ✅ | ✅ | ✅ | ✅ | ✅ | Per-use | Contact | SOC2 |
| Together AI | H100, A100, L40S, T4 | Contact | Contact | Contact | ✅ | ❌ | ✅ | ✅ | ✅ | ✅ | ✅ | Per-token | Contact | SOC2 |
| Paperspace | H100, A100, V100, RTX | $5.95 | $3.09–3.18 | $0.56 | ✅ | ❌ | ✅ | ✅ | ✅ | ✅ | ✅ | Per-hour | Free | Enterprise |