Niko1221/Strata promises that a 125-billion-parameter model — Qwen3.8-Flash-Next — runs on your own gaming PC, that “a graphics card with 12 GB or more” is enough, and that it writes 94 tokens per second while doing it. It was created on 24 September 2026 and pushed to on 4 October, it reached the front page of Hacker News with 595 points and 279 comments, and it has 11,048 stars and 969 forks in 11 days.
The headline is the kind that usually collapses on contact. This one does not. I downloaded the repository, queried the GitHub API for the release, contributor and issue history, read the three documents that describe the engine, pulled the documentation’s own numbers apart, and did the memory-bandwidth arithmetic myself. The speed claim survives — but it survives for a reason the marketing sentence hides, and the same repository contradicts itself by a factor of 5.3 on prompt speed.
What the repository actually is
This is not a thin wrapper around someone else’s binary. It is an inference engine: 235 C++ files, 56 CUDA files, 167 headers, its own CUDA and SYCL backends, a vendored ggml fork in third_party/, CMake, a Dockerfile, 172 files under bench/ and 30 tests. It is MIT-licensed.
| Measured | |
|---|---|
| Created / last push | 24 Sep 2026 / 4 Oct 2026 |
| Stars / forks / watchers | 11,048 / 969 / 71 |
| Stars per day | 1,004 |
| Releases | 40 (v0.1.0 → v0.1.39 in 11 days, 3.6 per day) |
| Release asset downloads | 65,702 |
| Contributors | 68 (top: 520 commits, second: 21) |
| Open issues / PRs on page 1 | 98 / 67 |
| Files / directories / bytes | 852 / 113 / 27,100,222 |
| Licence | MIT |
Two details from that table deserve to be read together. The first contributor has 520 contributions against 21 for the second — 24.8 times as many — which is the footprint of a single-author project with a large drive-by community; several Hacker News commenters looked at the same shape and concluded the codebase is largely machine-written (“Main contributor: Claude”, wrote one). I did not audit authorship, and I would not repeat that as fact, but the concentration is real and worth knowing when you weigh how much independent review the engine has had.
The second detail is in the download counts. The most-downloaded release asset is strata-windows-x64.zip from v0.1.38 at 14,179 downloads. The second is the 49-second demo video, Pagoda.mp4, at 12,493 — 19% of every asset download in the project’s history, more than every build except one. For a 13.8 MB trailer to out-download 38 of the 39 alternative builds tells you how much of this project’s adoption is still people watching, not running.
The 12 GB is the least important number
Here is the requirement sentence most coverage repeated: a NVIDIA or AMD card with 12 GB of VRAM or more. Here is the rest of the same table: 32 GB of RAM minimum and about 80 GB of disk, with a model download of roughly 70 GB and a first start that loads “35-55 GB into your RAM” and can freeze the machine for one to three minutes.
Then read the sizing table in docs/MODELS.md, which is the honest version:
| Size | RAM + VRAM needed | Expert weights held in RAM |
|---|---|---|
| Q2_0 (fastest) | 37.6 GB | 34 GB |
| IQ2_XS (recommended) | 39.2 GB | 35.5 GB |
| IQ3_XXS | 47.0 GB | 43 GB |
| IQ3_S (best quality) | 54.8 GB | 50 GB |
| Coder (IQ1_M) | 29.6 GB shard 1 | 23 GB |
Q2_0 puts 34 GB of expert weights in system memory — 2.8 times the size of the entire 12 GB card. The graphics card holds the dense layers, the attention and DeltaNet mixers, the router, the output head and a cache of the most-used experts. The model itself lives in the RAM you already have.
The project’s own guidance follows from that, and it is stricter than the headline: with 32 GB of RAM the only model the installer recommends is the Coder, which keeps 256 of each layer’s 512 experts and is, in the documentation’s own words, “weaker outside code and in languages other than English”. With 48 GB you get Q2_0 or IQ2_XS. You need 64 GB before every size fits. A bigger graphics card makes it faster; the docs say explicitly that it does not lower the RAM you need.
So the purchase decision is not a graphics card. It is 64 GB of DDR5, which on most motherboards means four DIMMs or two 32 GB modules, and which is where the performance actually comes from.
Where 94 tokens per second actually comes from
The engine is a mixture-of-experts model: 48 layers × 512 experts = 24,576 experts, of which the router picks 10 per layer per token — 480 expert instances for one token. That sparsity is the whole trick, and it also lets you check the published numbers instead of trusting them.
Q2_0 keeps 34 GB of expert weights for 24,576 experts, so one expert is about 1.38 MB. One token reads 480 of them: 664 MB per token if nothing is cached. At the published 94 tokens/s that is 62.4 GB/s of sustained memory traffic — against a theoretical peak of 83.2 GB/s for the dual-channel DDR5-5200 the README specifies, and a realistic sustained figure closer to 50–60 GB/s. That is 75% of theoretical peak, before the operating system, the desktop and the browser take their share.
Now do the same arithmetic for the AMD row. The RX 9070 XT machine has a Ryzen 9 3900X and 47 GB of RAM, which means DDR4-3200: a 51.2 GB/s ceiling. The published figure there is 60 tokens/s, which needs 39.8 GB/s — 77.8% of that machine’s ceiling. The published speeds sit at almost exactly the same fraction of their respective memory-bandwidth peaks (75.0% and 77.8%), and their ratio — 60/94 = 0.638 — tracks the bandwidth ratio 51.2/83.2 = 0.615.
That is not a coincidence, and it is the most useful thing in this article. The 94 tokens/s figure is a memory-bandwidth number, and the hardware that decides it is the RAM, not the GPU. The GPU is a cache: the docs say every extra gigabyte of VRAM holds “about 700 additional experts”, and the measured numbers in DETAILS.md put VRAM residency at 1,589 experts normally and 3,872 with KV streaming enabled — between 6% and 16% of the 24,576. The other 84–94% are computed by your CPU from RAM, in parallel with the GPU, over the PCIe bus.
There is one more term that makes the published table close, and the project is upfront about it: speculative decoding. A small multi-token-prediction head drafts several tokens and the model verifies them in one pass; the docs claim 1.6–1.8× and say an average of 2.4 to 3.2 tokens are accepted per forward pass. At 94 tokens/s that means only 29–39 weight-reading passes per second, dropping the required RAM bandwidth to 19.5–26.0 GB/s. So the numbers do close — but only if the speculative decoder really accepts two to three tokens per pass on your prompt. On code and structured output it does; on unusual text the docs admit acceptance falls off and “a different answer to the same prompt moves it by several percent”.
This is why every reader-reported number above 100 tokens/s in the Hacker News thread comes from a machine with either a 24 GB card or 128 GB of RAM, and why the 4-bit Unsloth build that does not fit in RAM — 111 GB download, 77 GB of experts — collapses to 7-8.5 tokens/s on a 64 GB PC with a 12 GB card. When the experts have to come from the SSD per token, the RAM-bandwidth argument turns into an SSD one and the machine stops being usable for agentic work.
The documentation contradicts itself by 5.3×
The README tells you: “Strata reads the first message of a chat in full, about 1 minute per 30,000 tokens.” That is roughly 500 tokens/s of prompt processing.
docs/MODELS.md says: “A 32K prompt takes about 15 seconds with Q2_0”, in a table that lists 2,650 tokens/s for Q2_0 and 2,180 for the Coder. That is 2,180–2,650 tokens/s — 5.3 times faster than the README’s own sentence.
Both cannot be right, and the repository’s changelog says which one is stale. Prefill throughput was capped and slow until engine 0.1.13, which introduced automatic chunking up to 8,192 tokens; DETAILS.md then records engine 0.1.36 moving 32K prefill from 2,170 to 2,653 tokens/s with fused int8 tensor-core kernels. The “1 minute per 30,000 tokens” line describes an engine several dozen releases old and was never removed, while the 15-second line is from 0.1.26 and 0.1.36. It is a documentation bug rather than a hardware one — but it is the sentence a new user meets first, and prompt speed is the difference between a usable coding agent and one that spends a minute swallowing a file.
Two smaller versions of the same problem:
- The headline table mixes engine versions. The Q2_0 row was measured with engine 0.1.36; every other row in the same table is engine 0.1.26, and
DETAILS.mdsays so in a footnote (“The tables below are 0.1.26’s”). - The fastest AMD number is a configuration the installer refuses.
MODELS.mdannotates the Q2_0 AMD row as being above setup’s own RAM estimate for a 47 GB machine, installed with the override--model Q2_0 --yes. The best AMD number in the marketing table is flagged as out of spec by the installer that ships with it.
In the project’s defence: it does label estimates as estimates. The 24 GB RTX 3090 figure — “should write about 100-140 tokens per second” — sits under a heading called “Other GPUs (estimated)” with the note “Not measured - estimated from the runs above… Treat as…”, and the repository ships a file literally named experimental-speed-projection. That is more discipline than most projects of this size show. The problem is what happens downstream: the Hacker News submission carried the title “Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s”, a card the repository never benchmarks and a number that came from a user’s comment, not the maintainer. Retold once, an estimate became a measurement on hardware nobody tested.
What independent measurements say, and where the tradeoff lands
The Hacker News thread is unusually useful because the commenters did the work. Community reproductions broadly confirm the speed table rather than exceeding it: about 60 tokens/s for IQ3_S on a 3090, 65 tokens/s for IQ2_XS on an RX 9070 XT with a 5900X, “over 110 tokens/s” with 3-token MTP on a 4090 with 128 GB of DDR4. For scale, the same model’s IQ3_XXS under llama.cpp on a Ryzen 8845HS with 96 GB of RAM and no graphics card at all runs at 7 tokens/s decode and 60 tokens/s prefill — so Strata’s contribution is real, roughly an order of magnitude on decode at that tier.
The quality side is where the honest objections live, and the project does not answer them:
- The documentation publishes no accuracy numbers.
MODELS.mdrates its own sizes qualitatively — “good” (Q2_0), “better” (IQ2_XS), “great” (IQ3_XXS), “best: matches the full model on the published tests” (IQ3_S). The only number in the repo is a 4-bit KV cache perplexity loss of 8–12%, which is a warning about the cheaper cache, not a measure of the 2-bit weights. One commenter’s framing — “publishing benchmarks with quantized models should become standard practice” — is the right standard and this project does not meet it. - The one quoted quality figure is third-party and borrowed. The Coder’s “91.3% of the full model’s SWE-bench Verified score and 98.7% of LiveCodeBench v6” comes from ISTA-DASLab, the model’s authors, for an expert-pruned variant that keeps 256 of 512 experts per layer. It is a legitimate number about a different model than the one in the speed table’s top row.
- A quantified vision regression. A commenter ran a 50-image coordinate-regression benchmark at temperature 0 with the exact same GGUF and vision adapter weights under Strata and under llama.cpp: median error 154.8 px against 46.5 px, mean 168.8 against 81.4, a gap he described as “as large as the jump from a 9B model to a 35B model”. That matters because the vision path is optional and platform-limited — on AMD cards images are processed on Linux through the CPU and “on Windows they can’t yet” — and because
DETAILS.mditself concedes of the quantized encoder that “we have not measured its accuracy against BF16 yet”. A 0.9 GB encoder plus ~1.4 GB of reserved VRAM also costs text speed: 2–8% slower generation when images are enabled. - Size versus gain. The 125B model’s files are roughly six times larger than a 27B model at the same quantisation for, on the model card’s own benchmarks, under 10% of measured gain — the tradeoff a 64 GB machine is being asked to accept.
- Concurrency is a weak point. Strata serves one request at a time by default;
"parallel": 2is opt-in, and on a 12 GB card “this makes each answer slower” — because two conversations activate different experts and the VRAM expert cache, holding 6–16% of the table, thrashes.
Shard 2: the 28.8 GB file nobody mentions in the headline
One structural detail is more interesting than any of the benchmarks. Strata’s model is split into two shards, and shard 2 is a 28.8 GB “lookup table” that stays on your SSD. It is the model’s n-gram embedding layer — a PLE (Parallel Lookup Embedding) table of roughly 20 million bigram and trigram entries at layer 2, embedding dimension 2,560. During generation the engine fetches “a few rows per token” from your disk through unbuffered direct I/O (--ple-io direct).
This is the mechanism that lets the model carry parameters without paying for them in RAM — parameter scaling that the CPU and GPU never see. It is also a second per-token dependency on hardware the marketing copy never names: your SSD. If you are choosing between an NVMe drive and a SATA one, or running this on an external disk, this table is why the documentation insists on an SSD for first start, and why a machine that technically fits can still behave badly under one request per second of random reads.
The related engineering is real and worth crediting. KV streaming above 64K context keeps the cache in RAM and only the attended part in VRAM (--kv-resident 32768), at about 13.7 KB of RAM per context token — 1.7 GB at 128K — and takes Q2_0 at 262K context from 50.9 to 62.6 tokens/s by raising VRAM residency from 1,589 to 3,872 experts. A hybrid K8V4 cache cuts KV memory 23% and took a 3090 from 85 to 99 tokens/s on a 198K context. The open pull requests read like a systems project rather than a launch-week artefact: “Refactor PageCache to fix a 45× read amplification”, “hip: RDNA3 (gfx1100) WMMA kernel for the QSA block-score select/scorer”, “perf: opt-in resident expert exchange buffer rotation”. Forty releases in eleven days is not always a good sign, but this engine is being measured, profiled and fixed in public.
So should you run it?
If you have 64 GB of RAM, a 12 GB or larger NVIDIA card and an NVMe SSD, and you want a local model good enough for well-scoped work, this is one of the most capable things you can install today. Take IQ2_XS (the project’s own recommendation) rather than the fastest row in the table: Q2_0’s speed is real, but the documentation has no accuracy number to support its “good” label, and the gap between “fastest” and “recommended” is one quantisation step.
If you have 32 GB of RAM, you are choosing between the Coder — which the docs say is “weaker outside code and in languages other than English”, with Chinese answers that “came out wrong or looping” — and nothing. At that point a good 27B model at 4-bit on the card you already own is the better use of the machine: four bits of a smaller dense model against two bits of a model you can only partially evaluate is not a comparison that favours the bigger number.
If your priority is images, test before you trust. The encoder’s accuracy is unmeasured by the project, the AMD path is limited to CPU inference on Linux, and the only public, quantified comparison against llama.cpp on identical weights shows a 3.3× median-error penalty.
And ignore the “12 GB” framing when you plan. Decide from RAM: 32 GB → Coder only; 48 GB → Q2_0 or IQ2_XS; 64 GB → everything except the 4-bit Unsloth builds at 111 GB; 96 GB+ → IQ3_S or UD-IQ4_XS. Budget 80 GB of disk, expect the first start to lock the machine for one to three minutes, and expect one request at a time unless you deliberately pay for concurrency.
FAQ
Is the 94 tokens/s claim true? It is consistent with the hardware, and the arithmetic is checkable: 480 of 24,576 experts per token at 1.38 MB each is 664 MB of traffic per token, which at 94 tokens/s is 62.4 GB/s — 75% of the theoretical peak of the DDR5-5200 the README recommends. That is achievable only with a working VRAM expert cache and speculative decoding accepting 2–3 tokens per pass. The figure is not fabricated; it is a best-case number on a machine built around it.
Why does a 12 GB card run a 125B model at all? Because only 10 of 512 experts per layer are used per token, and because the model is split: dense weights and a hot-expert cache in VRAM, all 24,576 experts in system RAM, and a 28.8 GB n-gram lookup table on the SSD. A 12 GB card is enough to hold the small hot part. It is a cache, not the model’s home.
Does a bigger graphics card replace the RAM I need? No. The documentation says a bigger card makes it faster but “doesn’t lower the RAM needed” — with the exception of low-RAM mode, which maps experts from disk and, on a small card, means “most experts then come from the SSD and it is much slower”.
Is the 2-bit quantisation a real quality loss? The project does not publish the numbers that would answer this, which is itself the answer. Its own labels run “good” to “best” with no accuracy table for the base sizes, the only loss figure in the repo is an 8–12% perplexity hit for the 4-bit KV cache, and community claims run from “IQ3_XXS is within a point of the fully unquantized model” to an estimate that a 2-bit quant retains about 80% of the full model’s recall. Test on your own workload.
How slow is the 4-bit version if it does not fit in RAM? Unsloth’s UD-Q4_K_XL is a 111 GB download with 77 GB of experts and writes 7-8.5 tokens/s on a 64 GB PC with a 12 GB GPU, with slow long prompts. The 94 GB UD-IQ4_XS is better behaved: it keeps your RAM minus 24 GB of experts resident and streams the rest from an NVMe SSD, and wants 48 GB of RAM or more.
What is the strongest criticism of the project? Not the speed — the speed is real. It is that the fast, headline-grabbing configuration is the least verified one: 2-bit weights with no published accuracy evaluation, a vision path the project admits it has not measured and which one independent test found 3.3× behind llama.cpp on identical weights, and documentation that still tells you a 32K prompt takes a minute when its own sibling file says fifteen seconds.
Conclusion
Strata is the rare project whose extraordinary claim mostly survives. A 125B mixture-of-experts model does run on a machine with a 12 GB card, and the published speed table is internally consistent with the memory bandwidth of the hardware it names — so consistent that the two published rows land at 75% and 78% of their respective RAM ceilings.
But the sentence “12 GB is enough” is doing the work of a different sentence: “64 GB of RAM is enough.” The card is a cache for 6–16% of the experts; the model lives in RAM, a 28.8 GB table lives on your SSD, and the speed you feel is your memory controller, not your GPU. Measure the claim that way and it is true, checkable, and a genuinely impressive piece of local-inference engineering — one whose weakest point is not performance but the accuracy evidence it has not published, and the documentation it has not updated.