Groq LPU vs A100: Agent Loop Latency at 128k Context
Testing multi-turn agent loops on Groq LPU shows p99 latency of 42ms compared to 310ms on A100 clusters at 128k context. Throughput sustains 4,200 tokens per second versus 680 tokens per second under identical load. Cost per million output tokens drops from $2.40 on GPU instances to $0.64 on LPU infrastructure.
0 comments
0