Why We Brought This Tool Into Our Lab
We brought GroqCloud into our lab for a narrow reason: user-facing agents spend too much time waiting for models to emit tokens.
An agent can have a fast vector store, a warm application container, and sub-20 ms internal service calls, yet still feel slow because generation takes several seconds. Faster decoding changes that interaction materially. It also shortens sequential agent loops where every tool result triggers another model call.
Groq attacks this bottleneck with its LPU inference architecture and an OpenAI-compatible API. We did not treat that architecture as proof of production performance. We separated four measurements that are too often collapsed into one headline:
- Client-observed time to first useful token: request start to the first non-empty content delta.
- End-to-end latency: request start to a completed stream.
- Client-side output-throughput proxy: returned completion tokens divided by elapsed time from the first non-empty content delta to stream end. We do not treat this as verified answer-token decoding speed when reasoning or other non-visible tokens may be included in completion usage.
- Workload cost: returned input and output tokens multiplied by the model-specific rates.
That distinction matters. Groq’s console records server-side latency, while an end user experiences network time plus server time. A dashboard can therefore look healthy while clients in another region see a worse tail. We also found that “first chunk” is not necessarily “first useful token”: role metadata or an empty content delta can arrive before visible text.
The historical LPU result that put Groq on engineering teams’ radar used approximately 100 input tokens and 200 output tokens. Groq’s Llama 2 Chat 70B endpoint reached 241 output tokens per second in that 2024 test, with an estimated 0.8-second response time for 100 output tokens. That is useful architectural evidence, but it is not a valid 2026 capacity number. The model, API traffic, and product lineup have all changed.
For an external comparison, we recorded the July 2026 snapshot summarized in this 2026 Groq benchmark. In our July 2026 snapshot comparison, we recorded median time to first token of 0.71 seconds for GPT-OSS 120B at high reasoning effort, 0.73 seconds at low reasoning effort, and 0.77 seconds for GPT-OSS 20B at high reasoning effort. Those figures include internal thinking before an answer token appears. We use them as a sanity check, not as substitutes for measurements from our deployment region.
The model lineup was our first warning against benchmark marketing. llama-3.1-8b-instant and llama-3.3-70b-versatile were scheduled to shut down on August 16, 2026. A benchmark built around either identifier is now historical evidence, not a basis for deployment. We based the runnable harness on openai/gpt-oss-20b and openai/gpt-oss-120b, while keeping the model configurable.
Our goal was not to prove that one accelerator always beats another. We wanted to answer a more operational question: can we route latency-sensitive agent turns to Groq without losing control of tail latency, billing, stream integrity, and fallback behavior?
Hands-On Walkthrough: Setup, Execution & Output
We used the chat-completions endpoint directly so the harness could inspect SSE events and response headers. The only secret required is a Groq API key:
python -m venv .venv
source .venv/bin/activate
pip install "httpx==0.27.0"
export GROQ_API_KEY="gsk_replace_me"
export GROQ_MODEL="openai/gpt-oss-20b"
export GROQ_RUNS="30"
We fixed the workload to one compact agent-planning turn:
- A system instruction requiring JSON.
- One user request describing three support-ticket actions.
- Temperature set to zero.
- A maximum of 256 output tokens.
- Streaming enabled.
- Sequential requests, so the result measures single-stream behavior rather than concurrent account capacity.
The following script calculates p50 and p95 for first useful token and total latency. It also records output throughput, token usage, routing headers, and cost. We intentionally fail a run if the stream ends without a completion signal or without authoritative usage; silently estimating billing tokens would make the cost result look more precise than it is.
# bench_groq.py
import json
import math
import os
import statistics
import time
import httpx
API_URL = "https://api.groq.com/openai/v1/chat/completions"
MODEL = os.getenv("GROQ_MODEL", "openai/gpt-oss-20b")
RUNS = int(os.getenv("GROQ_RUNS", "30"))
PRICES = {
"openai/gpt-oss-20b": {"input": 0.075, "output": 0.30},
"openai/gpt-oss-120b": {"input": 0.15, "output": 0.60},
}
MESSAGES = [
{
"role": "system",
"content": (
"You are a support operations agent. Return JSON only with keys "
"priority, actions, and escalation_reason. Use at most three actions."
),
},
{
"role": "user",
"content": (
"A paid customer cannot export invoices. Login works. The export "
"request returns HTTP 500, and retrying twice did not help. Plan the response."
),
},
]
def percentile(values, fraction):
ordered = sorted(values)
index = max(0, math.ceil(fraction * len(ordered)) - 1)
return ordered[index]
def run_once(client):
payload = {
"model": MODEL,
"messages": MESSAGES,
"temperature": 0,
"max_tokens": 256,
"stream": True,
"stream_options": {"include_usage": True},
"service_tier": "on_demand",
}
started = time.perf_counter()
first_content_at = None
finished_at = None
finish_reason = None
usage = None
text = []
with client.stream("POST", API_URL, json=payload) as response:
response.raise_for_status()
region = response.headers.get("x-groq-region")
ray = response.headers.get("cf-ray")
rate_headers = {
key: value
for key, value in response.headers.items()
if key.lower().startswith("x-ratelimit") or key.lower() == "retry-after"
}
for line in response.iter_lines():
if not line.startswith("data:"):
continue
data = line[5:].strip()
if data == "[DONE]":
finished_at = time.perf_counter()
break
event = json.loads(data)
if event.get("usage"):
usage = event["usage"]
choices = event.get("choices") or []
if not choices:
continue
choice = choices[0]
content = (choice.get("delta") or {}).get("content")
if content:
if first_content_at is None:
first_content_at = time.perf_counter()
text.append(content)
if choice.get("finish_reason") is not None:
finish_reason = choice["finish_reason"]
if first_content_at is None:
raise RuntimeError("Stream contained no non-empty content delta")
if finished_at is None or finish_reason is None:
raise RuntimeError("Stream ended without a verified completion signal")
if usage is None:
raise RuntimeError("No authoritative token usage was returned")
output_tokens = usage["completion_tokens"]
decode_seconds = max(finished_at - first_content_at, 1e-9)
return {
"ttft_ms": (first_content_at - started) * 1000,
"total_ms": (finished_at - started) * 1000,
"output_tps": output_tokens / decode_seconds,
"input_tokens": usage["prompt_tokens"],
"output_tokens": output_tokens,
"region": region,
"cf_ray": ray,
"rate_headers": rate_headers,
"response_chars": len("".join(text)),
}
def main():
api_key = os.environ["GROQ_API_KEY"]
headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
}
results = []
with httpx.Client(headers=headers, timeout=45.0) as client:
for index in range(RUNS):
result = run_once(client)
result["run"] = index + 1
results.append(result)
ttft = [row["ttft_ms"] for row in results]
total = [row["total_ms"] for row in results]
throughput = [row["output_tps"] for row in results]
input_tokens = sum(row["input_tokens"] for row in results)
output_tokens = sum(row["output_tokens"] for row in results)
price = PRICES[MODEL]
cost = (
input_tokens * price["input"] / 1_000_000
+ output_tokens * price["output"] / 1_000_000
)
summary = {
"model": MODEL,
"runs": RUNS,
"ttft_ms": {
"p50": percentile(ttft, 0.50),
"p95": percentile(ttft, 0.95),
},
"total_ms": {
"p50": percentile(total, 0.50),
"p95": percentile(total, 0.95),
},
"output_tps": {
"median": statistics.median(throughput),
"p05": percentile(throughput, 0.05),
},
"tokens": {"input": input_tokens, "output": output_tokens},
"estimated_cost_usd": cost,
"regions": sorted({row["region"] for row in results}),
}
print(json.dumps(summary, indent=2))
if __name__ == "__main__":
main()
We run it with:
python bench_groq.py > result.json
A successful execution emits this structure with values measured from the caller’s own network path:
{
"model": "openai/gpt-oss-20b",
"runs": 30,
"ttft_ms": {
"p50": "<measured by the script>",
"p95": "<measured by the script>"
},
"total_ms": {
"p50": "<measured by the script>",
"p95": "<measured by the script>"
},
"output_tps": {
"median": "<measured by the script>",
"p05": "<measured by the script>"
},
"tokens": {
"input": "<API-reported total>",
"output": "<API-reported total>"
},
"estimated_cost_usd": "<calculated from reported usage>",
"regions": ["<x-groq-region value>"]
}
What This Article Could Not Verify
Our retained evidence does not include a completed remote run, so we cannot report live latency percentiles from our own deployment. The harness is our reproducible artifact; the 2026 external figures provide a comparison. Anyone evaluating a production region should run it from the same network and runtime that will serve users.
Streaming Failures and Deployment Constraints
The most important failure was not a slow token. It was an apparently successful but incomplete stream.
We reproduced the stream contract locally with groq==0.9.0, httpx==0.27.0, and pydantic==2.8.2. Our HTTP transport emitted deterministic OpenAI-compatible SSE chunks, which let us test completion, a read failure after partial content, and a clean EOF without a finish chunk.
The complete fixture returned four SDK chunks, two non-empty content chunks, the text Hello from fixture., and finish_reason="stop".
When we injected a transport read failure after the first content delta, the SDK returned Hello and propagated an httpx.ReadError. It issued one request and did not replay automatically. That behavior is preferable to an invisible retry, but the application must distinguish partial UI output from a committed result.
The nastier case was clean EOF. The iterator returned normally after yielding Hello even though no finish reason had arrived. There was no exception. Code that treats normal loop termination as success can therefore persist a truncated answer, trigger an incomplete tool call, or bill a workflow as completed.
Our workaround is strict:
- Buffer or mark streamed content as provisional.
- Require a non-null finish reason before committing it.
- Treat EOF without a completion signal as an integrity failure.
- Never replay an agent turn blindly after partial output.
- Attach idempotency controls to downstream tools and business actions.
We also hit four broader deployment constraints.
Rate limits are account and model constraints, not accelerator throughput. A model may decode at hundreds of tokens per second while the account still rejects a burst. We log every x-ratelimit-* and retry-after header, cap concurrency with a semaphore, and retry 429 or transient 5xx responses with bounded exponential backoff and jitter. We do not retry malformed requests or replay partially executed tool turns.
Service tiers change failure semantics. We use on_demand for interactive traffic. Flex processing can provide up to ten times the current rate limits when capacity exists, but it may fail rapidly when capacity is constrained. Auto begins with on-demand limits and can fall back to flex. That makes flex useful for replayable enrichment jobs, not for an agent action that must execute exactly once.
Model identifiers are routing dependencies. Groq is not a generic deployment target where we upload arbitrary weights. We route only to its curated catalog, and scheduled model removal can break hard-coded applications. We now place model IDs behind configuration, run a startup capability check, and keep a quality-tested fallback at another provider.
Token counting must come from the response. Character counts, local tokenizers, and concatenated SSE fragments are useful diagnostics, but they are not billing records. Reasoning and tool-related tokens can make client estimates diverge. We calculate cost from authoritative prompt and completion usage, store the raw usage object, and reject a benchmark sample when usage is missing.
The same caution applies to latency. We separate network time from server time and treat input length as the primary TTFT driver in our measurement design (latency reference). We preserve input size, output cap, reasoning mode, region, service tier, and concurrency with every sample. Without those dimensions, p95 is just a number detached from a workload.
Scale, Latency & Cost vs. Alternatives
Our current comparison is not “LPU versus GPU” in the abstract. It is hosted Groq versus a conventional GPU API versus self-hosted vLLM for a specific traffic shape.
| Decision factor | GroqCloud | Hosted GPU inference API | Self-hosted vLLM or SGLang |
|---|---|---|---|
| Fast single-stream decoding | Core strength; published 2026 rates were 1,000 tok/s for GPT-OSS 20B and 500 tok/s for GPT-OSS 120B | Varies heavily by provider, batching, and model | Tunable, but low-load latency and saturated throughput require different configurations |
| Operational work | Low | Low | High: GPUs, autoscaling, model loading, observability, and upgrades |
| Model freedom | Curated Groq catalog | Usually broader | Highest, subject to hardware and license |
| Burst behavior | Controlled by account, model, and service-tier limits | Provider-specific quotas and queues | Controlled by owned capacity; excess traffic queues or fails |
| Cost model | Per input/output token | Usually per token | GPU-hours plus engineering and idle capacity |
| Streaming risk | SSE completion must be verified | Similar API-stream concerns | Full control, but full responsibility |
| Best fit | Interactive agents where latency has business value | Broad model access and managed fallback | Stable, high utilization or strict deployment control |
The July 2026 price points we validated were:
| Model | Published speed | Input per 1M tokens | Output per 1M tokens | Fixed workload cost |
|---|---|---|---|---|
| GPT-OSS 20B | 1,000 tok/s | $0.075 | $0.30 | $0.000150 |
| GPT-OSS 120B | 500 tok/s | $0.15 | $0.60 | $0.000300 |
The fixed workload cost assumes 1,000 input tokens and 250 output tokens:
GPT-OSS 20B:
(1,000 × $0.075 / 1,000,000) + (250 × $0.30 / 1,000,000)
= $0.000150 per turn
GPT-OSS 120B:
(1,000 × $0.15 / 1,000,000) + (250 × $0.60 / 1,000,000)
= $0.000300 per turn
At one million such turns, that is $150 for GPT-OSS 20B or $300 for GPT-OSS 120B before caching, retries, failed generations, tool-loop expansion, and taxes. The blended rates are $0.12 and $0.24 per million total tokens respectively, but that shorthand is valid only for this 4:1 input-to-output ratio. Output-heavy agents cost more because output tokens carry the higher rate.
Batch processing reduces synchronous pricing by 50% and does not consume standard API quotas, but its completion window makes it a different product. At the same workload, the nominal model cost becomes $75 or $150 per million turns. We would use that for evaluation runs, enrichment, and scheduled classification—not interactive chat.
For self-hosting, the useful break-even equation is:
required requests per hour =
GPU cluster cost per hour / hosted API cost per request
If a self-hosted cluster costs G dollars per hour, GPT-OSS 20B’s token-only break-even for this workload is G / 0.000150 requests per hour. For GPT-OSS 120B it is G / 0.000300. We then add engineering, failover capacity, observability, idle headroom, and the cost of meeting p95 during bursts. Comparing Groq’s per-token bill against a fully utilized GPU while ignoring idle time is not a serious production calculation.
We would also require identical models, prompts, output caps, concurrency, and regions before claiming that Groq is “N times faster” than a GPU service. The old 241 tok/s Llama 2 result illustrates the hardware’s decoding performance; it does not establish a current speed advantage over an unspecified GPU service.
Teams building a multi-provider router can review our broader tools collection. For workload-specific capacity and failover design, our AI infrastructure services cover the parts that raw tokens-per-second charts omit.
Our Final Verdict: When to Deploy, When to Skip
Groq works best as a low-latency inference option within a multi-provider system, not as the default provider for every workload.
Deploy it if:
- User-perceived generation speed directly affects conversion, retention, or operator productivity.
- Your agent makes several sequential model calls and faster decoding shortens the full loop.
- A supported GPT-OSS model meets your quality requirements.
- You can measure TTFT and completion latency from the production region.
- You calculate cost from returned usage rather than tokenizer estimates.
- You enforce completion signals before committing streamed output.
- You have bounded retries, concurrency controls, and idempotent tools.
- You can tolerate a curated model catalog and maintain a tested fallback.
Hold off or avoid it if:
- You need custom or fine-tuned weights that Groq does not host.
- Your compliance requirements have not been contractually verified.
- Your workload is offline and latency has little economic value.
- You depend on a deprecated model identifier.
- Your traffic exceeds available quotas and cannot be queued or shifted to batch.
- You need provider-independent behavior across every OpenAI-compatible extension.
- You cannot safely recover from a partial agent stream.
- You are choosing it solely because an old benchmark reported a large GPU multiple.
Our verdict is positive, with conditions. The published 2026 pricing is aggressive, the single-stream throughput targets are unusually high, and the API surface is straightforward. The operational risks sit around the accelerator: rate limits, service-tier semantics, model churn, network distance, token accounting, and ambiguous stream termination.
We would deploy Groq for latency-sensitive conversational turns behind our own gateway, with on-demand processing, explicit model configuration, raw usage logging, verified finish_reason, and a secondary provider. We would route replayable bulk work to batch or flex, compare median TTFT with the dated external figures only under comparable model and reasoning settings, and track p95 against our own workload-specific baseline and latency target.
We would not approve a rollout based on the 1,000 tok/s headline alone. Before routing production traffic to Groq, run the supplied harness from the actual service region, add a concurrency sweep that stays within the account’s quotas, and verify quality on the same agent traces used in production. If you want us to review that routing design, contact Effloow.
Top comments (0)