THE FORWARD PASS
October 5, 2026 issue

Story 02 October 5, 2026 issue

Daily AI-generated issue

Strata runs 125B Qwen3.8 Flash Next locally with 1.6 to 1.8x speedup

Image: github.com

Top Repo · 2,124 Stars

Strata is a free, MIT-licensed local runtime and app that runs the 125B Qwen3.8 Flash Next on a Windows or Linux PC with a 12 GB NVIDIA or AMD GPU.

Here's what makes it work:

  • The MoE model has 24,576 experts with 10 active per word. Strata keeps hot experts on the GPU, all experts in system RAM, and uses the CPU for the rest.
  • Speculative decoding claims a 1.6 to 1.8x speedup with the same answer.
  • Measured speeds hit 53 to 94 tokens/s on an RTX 5070 and 52 to 60 tokens/s on an RX 9070 XT, depending on quantized size.
  • It exposes OpenAI-compatible and Anthropic-style local endpoints, plus chat, coding and optional image input (AMD image reading is Linux-only).
  • A Coder variant drops half the experts and keeps 91% of the full model's SWE-bench Verified score.

One catch: you need about 80 GB of disk and 32 GB RAM is recommended. The Coder variant is weaker outside code, including CJK text.

Try it: clone https://github.com/Niko1221/Strata, run START-HERE.bat on Windows or ./setup.sh on Linux, then point clients at http://127.0.0.1:8080/v1.

Sources: GitHub · GitHub

This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.