On this page
- What Is Strata and How Does a 125B Model Fit in 12 GB of VRAM?
- How Fast Is Strata on a Real Gaming PC?
- Strata vs llama.cpp: What Do the Numbers Say on the Same $500 PC?
- What Does the Hype Hide? Agent-Written Code and a License Trap
- Should Your Business Run a 125B Model Locally?
- Frequently Asked Questions
- Can I run a 125B model on a 12GB GPU?
- Does Strata work on AMD cards?
- Is Qwen3.8-Flash-Next Apache 2.0 licensed?
- How much RAM do I need?
- Is Strata safe to run?
Last Updated: October 4, 2026
Key takeaways
- Strata runs the 125B Qwen3.8-Flash-Next mixture-of-experts model on a 12 GB GPU plus 64 GB of RAM by keeping hot experts in VRAM, pinning all 24,576 experts in system RAM, and computing cold-expert misses on the CPU.
- Official docs report 87 to 94 tok/s decode at short context on an RTX 5070 and 76.2 tok/s at 64k context on the Q2_0 quant.
- An independent RTX 4070 test measured 53.2 tok/s at 60k context, about twice a tuned llama.cpp baseline of 27.1 tok/s on the same PC.
- The model is not Apache 2.0: the Qwen Community 1.0 license allows commercial use but restricts hosted API businesses.
- The repository is 16 days old with 232 open issues, so build from source, pin a commit, and keep it off critical machines.
- In our own testing, a bigger 284B model runs at 41 tok/s on dual DGX Sparks: a multi-thousand-dollar tier for a different budget.
Yes, a 125-billion-parameter model now runs on a 12 GB gaming GPU, and the trick is not compression magic. Strata, an MIT-licensed inference runtime published on GitHub on September 24, 2026, serves Qwen3.8-Flash-Next at 87 to 94 tokens per second on an RTX 5070, according to the project's own benchmarks. An independent test by Toronto data analyst Kartikey Chauhan recorded 53.2 tok/s at 60k context on an RTX 4070, which he described as "About twice my best llama.cpp number." This post explains how the runtime works, what the benchmarks actually measure, and where the hype needs a haircut before you spend money.
What Is Strata and How Does a 125B Model Fit in 12 GB of VRAM?
Strata is a C++ inference server that runs Qwen3.8-Flash-Next, a 125B mixture-of-experts model, on consumer hardware. Instead of squeezing the whole model into VRAM, it keeps the hottest experts on the GPU, pins all 24,576 experts in system RAM, computes cache misses on the CPU, and reads a 28.8 GB n-gram table from the SSD a few rows at a time. Apps talk to it through OpenAI- and Anthropic-compatible APIs on 127.0.0.1:8080, with a browser chat, a monitor tab, and an MCP server for agents.
The model itself, released by Qwen on August 26, 2026, has 125B total parameters with only 6B active per token, a 51B-parameter n-gram embedding table, and a 4B multi-token prediction head. Its 48 layers stack Gated DeltaNet and Qwen Sparse Attention blocks, with 512 experts per layer of which 10 routed plus 1 shared fire per token, and it carries 262,144 tokens of native context extendable to 1M via YaRN. According to Artificial Analysis, it scores 56 on the Intelligence Index, 5th of 111 in its class, and Qwen itself labels it an under-trained preview ahead of Qwen4, per eesel.ai. The speed tricks are speculative decoding on top of that split: the MTP head drafts 3 tokens with 2.4 to 3.2 accepted per pass, prompt lookup drafts up to 5 on repetitive input so code edits finish 6 to 11 percent faster, and prompts ingest in 2,048-token chunks at more than 1,000 tok/s.

How Fast Is Strata on a Real Gaming PC?
On the project's reference box, an RTX 5070 12 GB with a Ryzen 5 7600 and 64 GB of DDR5, the Q2_0 quant decodes at 87.3 tok/s at 1k context and 76.2 tok/s at 64k, while prompts ingest at 2,171 tok/s at 32k, according to docs/DETAILS.md. Tighter 3-bit packs trade speed for quality: IQ3_XXS runs 61.9 and 57.2 tok/s at the same context points.
Engine 0.1.36 pushed decode from 89 to 93.5 tok/s at 4k context and from 64.5 to 76.4 at 128k, with prompt kernels 16 to 22 percent faster. More VRAM matters more than a faster GPU: the docs estimate roughly 140 tok/s short-context on a 24 GB RTX 3090, and each extra gigabyte holds about 700 more experts in cache. KV cache options include streaming from 64k when RAM allows, q4_0 at an 8 to 12 percent perplexity cost, and k8v4 which cuts KV memory 23 percent, worth 99 versus 85 tok/s on a 3090 Coder run at 198k. The variant menu matters for planning: Coder keeps 256 of 512 experts, fits in 32 GB of RAM, and holds 91.3 percent of the full model's SWE-bench Verified score, Swift 1.5 cuts thinking tokens 63 percent while passing the repo's own 8-question check 8 of 8 with 1,234 versus 2,682 output tokens, and Unsloth's UD-Q4_K_XL needs 111 GB and crawls at 7 to 8.5 tok/s from SSD on a 64 GB PC, proof that RAM rather than the GPU is the real gate.

Strata vs llama.cpp: What Do the Numbers Say on the Same $500 PC?
On the same $500-class PC, an i5-12600K with an RTX 4070 12 GB and 64 GB of DDR5, Chauhan measured 53.2 tok/s at 60k context for Strata against 27.06 tok/s for his best llama.cpp configuration after seven tuning steps. Strata also read his prompts at 2,013 tok/s.
| Metric | Strata | llama.cpp |
|---|---|---|
| Decode at 60k context | 53.2 tok/s | 27.1 tok/s |
| Setup effort | One-click installer, tuned packs | 7 manual tuning steps |
| Model coverage | Qwen3.8-Flash-Next GGUF packs | Thousands of GGUF models |
| Maturity | v0.1.38, 16 days old, MIT | Stable, multi-year project |
In his tiering, Strata is the fast default and llama.cpp is the stable fallback, which matches how we advise clients to think about young runtimes: adopt the speed when the work is low-stakes, keep the veteran engine one config switch away.
What Does the Hype Hide? Agent-Written Code and a License Trap
Two caveats deserve oxygen: nobody outside the project knows how much of Strata was written by AI coding agents, and Qwen3.8-Flash-Next ships under the Qwen Community 1.0 license, not the Apache 2.0 that several launch write-ups claimed.
Chauhan coined the term "napkin runtimes" for Strata and five peers (ninfer, DwarfStar, Splash, llamAmpere, gufo) and wrote plainly: "I don't know how much Strata was agent-written." DwarfStar's README admits "strong assistance from AI coding agents." His checkable prediction: by April 2027 at least 4 of the 6 engines will have gone 60 days without a commit. According to Startup Fortune's Julian Lim, the uniform speed of the praise itself became the story, because readers could not distinguish human enthusiasm from agent-generated hype. A community account, seekinganythingbutalpha, added the benchmark caveat: "we run 3-bit packs (IQ3_S, IQ3_XXS), so there is a gap to the official" numbers. On licensing, according to PacketNebula's breakdown, commercial use is allowed, but running the model as a managed API or "AI Work Assistant" business needs a separate license, deployments above 100 million monthly users or 20 million dollars in monthly revenue must display the model name, and the FP8 checkpoint alone weighs 172.8 GiB. Chauhan's trust guidance is blunt: build from source, pin a commit, and do not run it on a machine that holds anything you cannot afford to lose.
Should Your Business Run a 125B Model Locally?
If you need private coding or document inference and already own a 64 GB desktop, Strata is a credible $500 to $1,500 on-ramp to a 125B model. If you need million-token context on a larger model, you are in workstation-server territory, which is a different budget line entirely.
In our own Flowtivity testing, we run DeepSeek V4 Flash, a 284B FP8 model, across two NVIDIA DGX Sparks with tensor parallelism at 41 tok/s with 1M context, and a single Spark manages 12 to 15 tok/s at 131k. That is a multi-thousand-dollar path built for a different tier: bigger model, longer context, rack-grade memory bandwidth. Strata targets the opposite corner: one 12 GB card, one desktop, 64 GB of RAM, and a focused model. The durable part of either stack is not the engine, it is the interface. Chauhan's parts-bin observation is that napkin runtimes assemble general engines from GGML and GGUF parts, and what survives the churn is the OpenAI- and Anthropic-compatible HTTP API plus model routers such as llama-swap. That is exactly why we tell clients to standardize on the local API layer and treat the engine as swappable. A practical pilot: start with the Coder variant on 32 GB of RAM, keep llama.cpp as the fallback route, and use Qwen Cloud pricing of $0.16 input and $0.47 output per million tokens as the break-even yardstick for self-hosting.

Frequently Asked Questions
Can I run a 125B model on a 12GB GPU?
Yes, if the model is a mixture-of-experts with a small active footprint. Strata keeps hot experts in the 12 GB card, pins all experts in 64 GB of RAM, and computes misses on the CPU. Official numbers: 87 to 94 tok/s decode on an RTX 5070.
Does Strata work on AMD cards?
The project lists supported AMD cards alongside NVIDIA RTX 20/30/40/50 generations. You still need 12 GB of VRAM minimum, 64 GB of RAM recommended, an AVX2 CPU, and driver 580 or newer.
Is Qwen3.8-Flash-Next Apache 2.0 licensed?
No. It uses the Qwen Community 1.0 license: commercial use is allowed, hosted API businesses need a separate license, and very large deployments must display the model name.
How much RAM do I need?
Strata recommends 64 GB and pins every expert there. The Coder variant fits 32 GB at 91.3 percent of the full model's SWE-bench Verified score. Unsloth's smallest quant needs 75 GB.
Is Strata safe to run?
Treat it as young software: a 16-day-old repository with 232 open issues. Build from source, pin a commit, and keep it off machines that hold anything you cannot afford to lose.
One email a month, no noise
Practical AI notes for Australian businesses. Unsubscribe anytime.