# Strata runs 125B Qwen3.8 Flash Next locally with 1.6 to 1.8x speedup

From The Forward Pass daily issue, October 5, 2026 (https://theforwardpass.net/archive/daily/2026-10-05). Source: https://theforwardpass.net/archive/daily/2026-10-05/strata-runs-125b-qwen3-8-flash-next-locally-with-1-6-to-1-8x-speedup

> This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

**Top Repo** · 2,124 Stars

Strata is a free, MIT-licensed local runtime and app that runs the 125B Qwen3.8 Flash Next on a Windows or Linux PC with a 12 GB NVIDIA or AMD GPU.

Here's what makes it work:

- The MoE model has 24,576 experts with 10 active per word. Strata keeps hot experts on the GPU, all experts in system RAM, and uses the CPU for the rest.
- Speculative decoding claims a 1.6 to 1.8x speedup with the same answer.
- Measured speeds hit 53 to 94 tokens/s on an RTX 5070 and 52 to 60 tokens/s on an RX 9070 XT, depending on quantized size.
- It exposes OpenAI-compatible and Anthropic-style local endpoints, plus chat, coding and optional image input (AMD image reading is Linux-only).
- A Coder variant drops half the experts and keeps 91% of the full model's SWE-bench Verified score.

One catch: you need about 80 GB of disk and 32 GB RAM is recommended. The Coder variant is weaker outside code, including CJK text.

Try it: clone https://github.com/Niko1221/Strata, run START-HERE.bat on Windows or ./setup.sh on Linux, then point clients at http://127.0.0.1:8080/v1.

Sources: [GitHub](https://github.com/Niko1221/Strata) · [GitHub](https://github.com/PrimeIntellect-ai/prime-rl/releases/tag/v0.9.1.dev179)
