Groq LPU vs H100: Latency and Cost Benchmarks for Llama-3-70B
Real-world testing of Llama-3-70B shows Groq LPUs delivering a 12ms time-to-first-token, outperforming H100 clusters which average 145ms under identical load. Sustained throughput reaches 480 tokens per second on Groq hardware, a 4.3x improvement over the 110 tokens per second observed on NVIDIA GPUs. These performance gains reduce the effective cost per million output tokens from $0.80 to $0.19 for high-volume inference workloads.
0 comments
0