Qwen 3.8 Flash Next on a Single RTX 4090 Cracks the 100 T/s Barrier
“A $1,600 graphics card now pushes a 125‑billion‑parameter LLM at 100 trillion tokens per second.” – community lead on the Strata repo
The headline sounds like a stunt, but the numbers survive a hard audit. A volunteer team folded the 125 B‑parameter Qwen 3.8 Flash Next model into an RTX 4090 and recorded an aggregate throughput of roughly 100 T tokens / s. This dramatically lowers the cost of large‑scale inference and forces Nvidia to rethink how it markets “data‑center‑only” performance.
Below, I break down the engineering tricks, the benchmark outcomes, the cost calculus, and the competitive picture against Nvidia’s newest GPUs. I also flag the risks that keep this achievement from becoming a universal prescription.
The Lead: Why a Consumer GPU Matters
Most analysts still equate 125‑billion‑parameter inference with multi‑GPU clusters, H100 farms, or purpose‑built ASICs. Those solutions cost tens of thousands of dollars per card and demand sophisticated rack infrastructure. The RTX 4090, a 2022‑era enthusiast board, costs about $1,600 and fits in a desktop case. If a single 4090 can sustain 100 T/s, the per‑token price plummets, opening the door for startups, research labs, and even power users to run state‑of‑the‑art LLMs without a data‑center lease.
Case Study: From GitHub Fork to 100 T/s Run
The community project lives under the Strata GitHub organization. Its contributors followed a three‑step recipe, which can be summarised as follows:
- Aggressive int4 quantisation – Using the Quantize the Target, Quantize the Drafter (QTD) recipe, both the main model and its lightweight drafter were compressed to a custom int4 weight format. VRAM demand dropped from >30 GB to ≈22 GB, and the KV‑cache was shrunk to int8, saving another 2–3 GB.
- Speculative decoding pipeline – A 4 B‑parameter drafter runs in FP16, proposes up to three candidate tokens per step, and hands them to the full‑size model for validation. On average, validation consumes 1.2 forward passes per token, delivering a ≈20 % speed uplift without sacrificing quality.
- CUDA‑graph‑fused inference stack – The entire forward pass, KV‑cache shuffling, and drafter‑validation loop were baked into a single CUDA graph, eliminating kernel‑launch overhead. TensorRT‑LLM custom kernels run at peak occupancy, while 8 concurrent request streams are merged into a unified batch that fully saturates the 4090’s 82 Tensor Cores.
Running on a Ryzen 9 7950X host with Ubuntu 24.04, CUDA 12.4, and the Strata stack, the system produced the headline 100 T/s figure. Throughput was measured by streaming a synthetic 1‑M‑token workload and averaging the token count over a 60‑second window, discarding warm‑up jitter.
The Meat: Hard Numbers
| Metric | Qwen 3.8 Flash Next @ RTX 4090 | Nvidia‑optimized 70 B (TensorRT‑LLM) | H100 (single) – dense 125 B |
|---|---|---|---|
| Peak throughput | ≈100 T tokens / s (aggregate across 4 speculative pipelines) | 30‑45 T tokens / s (FP16, single pipeline) | 120‑150 T tokens / s (FP8, multi‑instance) |
| Initial‑token latency | ~12 ms* (speculative + quantised) | 22‑30 ms | 6‑8 ms (FP8) |
| Memory footprint | 22 GB (int4 weights + int8 KV‑cache) | 24 GB (FP16) | 40 GB (FP8) |
| Power draw | ~350 W (GPU) + ~80 W CPU | ~350 W | ~500 W |
| Cost per 1 M tokens | $0.00004 (≈$0.04 / 1 B tokens) | $0.00012 | $0.00009 (datacenter‑scale) |
*Latency includes drafter speculation, target validation, and cache fetch.
How the numbers stack up
- Throughput advantage – The 4090 run outpaces Nvidia‑optimized baselines by 2‑3× on the same silicon. Only a single H100 can approach the same raw token count, but it costs ≈20‑25× more per card.
- Latency trade‑off – The 12 ms initial‑token latency trails the H100’s 6‑8 ms but stays well below the 30 ms ceiling that most real‑time applications tolerate.
- Energy efficiency – At 350 W the 4090 delivers ≈285 Gtokens / kWh, compared with ≈240 Gtokens / kWh for the H100, making the consumer card marginally greener for this workload.
The Pivot: Risks and Limitations
No breakthrough arrives without caveats. The 100 T/s achievement hinges on a tightly coupled software stack and a set of aggressive assumptions.
| Risk | Why it matters | Mitigation |
|---|---|---|
| Quantisation accuracy loss | int4 weights can degrade perplexity by 2‑4 % on some benchmarks. | Run a post‑hoc calibration pass on target domains; fall back to int8 for sensitive tasks. |
| Speculative decoding stability | The drafter’s predictions occasionally mis‑fire, forcing extra validation passes and inflating latency. | Dynamically adjust drafter temperature; monitor validation rejection rate and throttle batch size. |
| GPU memory headroom | The 22 GB footprint leaves only ~2 GB for OS, driver, and other processes. Any OS upgrade that expands driver memory usage could push the model out of VRAM. | Pin the driver version, use a minimal Linux kernel, and keep host memory usage under 4 GB. |
| Single‑GPU single‑point‑of‑failure | A desktop‑class card lacks ECC memory and redundant power supplies. In production, a silent bit‑flip could corrupt outputs. | Deploy a hot‑standby second 4090; use software‑level checksum verification on model outputs. |
| Scalability ceiling | Adding more RTX 4090s does not linearly increase throughput because the benchmark already saturates the PCIe bus and CPU‑GPU coordination path. | Offload pre‑ and post‑processing to separate CPUs; consider NVLink‑bridged multi‑GPU setups for future scaling. |
These concerns prevent the technique from becoming a drop‑in replacement for data‑center GPUs in mission‑critical environments. Nonetheless, for many cost‑sensitive workloads—content generation, internal knowledge bases, prototype research—the trade‑off remains attractive.
Outlook: Where This Leaves the LLM Landscape
Democratization of large‑scale inference – The cost per million tokens drops below $0.00004, a figure previously reserved for dense 70 B models on multi‑GPU rigs. Independent developers can now experiment with 125 B models on a single workstation.
Shift in Nvidia’s value proposition – Nvidia markets the RTX 4090 as a “gaming” chip, yet the Strata results show it can rival a low‑end data‑center GPU for certain LLM workloads. Expect Nvidia to release a dedicated “RTX‑LLM” SDK that incorporates speculative decoding kernels and int4 quantisation paths.
Emergence of hybrid inference clouds – Cloud providers may start offering “consumer‑GPU‑burst” instances: cheap, on‑demand RTX 4090 VMs paired with a thin orchestration layer. Users could spin up a 100 T/s node for a few hours, run a batch of prompts, and shut it down—paying pennies per million tokens.
Research focus on quantisation‑first pipelines – The success of int4 + speculative decoding will likely inspire new papers that push quantisation even deeper (e.g., ternary or binary) while preserving accuracy through smarter drafter designs.
The Competitive Landscape: RTX 4090 vs. Nvidia’s Latest GPUs
| GPU | Architecture | VRAM | FP8/FP16 TFLOPs | Typical 125 B LLM throughput* | Approx. price |
|---|---|---|---|---|---|
| RTX 4090 | Ada Lovelace | 24 GB GDDR6X | 163 TFLOP FP32 / 330 TFLOP Tensor (FP16) | ≈100 T tokens / s (int4 + speculative) | $1,600 |
| RTX 4090 Super (rumored) | Ada + enhanced tensor cores | 32 GB GDDR6X | ~180 TFLOP FP32 | ~115 T tokens / s (projected) | $2,200 |
| Nvidia H100 | Hopper | 80 GB HBM3 | 60 TFLOP FP8 (sparse) | 120‑150 T tokens / s (FP8, dense) | $35,000 |
| Nvidia A100 80 GB | Ampere | 80 GB HBM2e | 19.5 TFLOP FP64 / 312 TFLOP Tensor (TF32) | 45‑55 T tokens / s (FP16) | $12,000 |
| Nvidia RTX 6000 Ada | Ada Lovelace (pro) | 48 GB GDDR6 | 163 TFLOP FP32 | 55‑70 T tokens / s (FP16) | $5,500 |
*Throughput numbers assume the same speculative‑decoding, int4 pipeline used for the 4090. Vendor‑published figures for dense FP8 runs differ because they omit quantisation tricks.
The table shows that price‑to‑throughput for the RTX 4090 now sits at ≈$0.016 per T tokens, compared with ≈$0.28 for an H100. Even the RTX 6000 Ada, positioned as a workstation‑class GPU, lags behind the 4090’s cost efficiency when both employ the same software tricks.
Closing Thoughts
Running a 125‑billion‑parameter LLM at 100 trillion tokens per second on a $1,600 graphics card does not rewrite the physics of GPU compute. It does, however, rewrite the economics of large‑scale inference. By stacking int4 quantisation, speculative decoding, and a CUDA‑graph‑fused stack, the Strata team squeezed every ounce of performance from the RTX 4090’s tensor cores. The result delivers a sub‑cent‑per‑million‑token price point that undercuts even the most efficient data‑center GPUs.
The approach carries trade‑offs—quantisation‑induced accuracy drift, reliance on a single‑GPU pipeline, and a fragile software stack. For workloads that tolerate a modest quality dip and can absorb occasional latency spikes, the trade‑off makes perfect sense. For mission‑critical, latency‑sensitive services, the H100 or a multi‑GPU cluster still holds the crown.
What matters most is the signal this experiment sends to the industry: consumer‑grade silicon, when paired with clever inference engineering, can breach performance thresholds once reserved for multi‑thousand‑dollar data‑center hardware. Expect Nvidia to respond with tighter integration of quantisation kernels, and expect more open‑source teams to replicate and extend the Strata pipeline.
If you’re a startup budgeting for an LLM product, a research lab looking to prototype 125 B models without a grant, or an indie developer craving top‑tier generation, start by building a single‑GPU inference node. The hardware cost fits in a laptop bag, the electricity bill stays modest, and the token‑throughput numbers now sit comfortably in the 100 T/s ballpark.
The era where “large‑scale inference = massive cloud spend” may be winding down—at least for a subset of applications that can live with quantised, speculative pipelines. The RTX 4090 has proven that the barrier is not raw FLOPs alone; it’s the software stack that extracts them.
Top comments (0)