Editorial

Deep dive The Forward Pass research desk

Strata’s 125B Local Runtime Is a Hardware-Sharing Story

Strata reports consumer-GPU inference for Qwen3.8 Flash Next by splitting expert execution across GPU and CPU, with speculative decoding adding a claimed 1.6 to 1.8× speedup.

AI-generated article

This article is researched and written by AI models from the sources it cites, and every claim is checked against them. No human edits it before it is published.

Kind
Deep dive
By
The Forward Pass research desk
Published
October 11, 2026
Read time
7 min

What ships

Strata is a free, MIT-licensed local runtime and app for the 125B Qwen3.8 Flash Next model. The project describes support for Windows and Linux PCs with supported NVIDIA or AMD graphics cards carrying at least 12 GB of VRAM.

The release story is the runtime, not a new model generation. Strata packages a way to run an existing large model locally, exposes interfaces for applications and coding agents, and includes chat, coding and optional image input.

Its hardware requirements extend beyond the graphics card. The repository lists 32 GB or more of system RAM and about 80 GB of free disk space. It says 64 GB of RAM runs every listed model size, and recommends an SSD because the first start is much faster.

The project is not describing a 125B model resident entirely in 12 GB of VRAM. Its stated design distributes model storage and execution across the GPU, system RAM, CPU and SSD.

Strata also makes an explicit locality claim: “Nothing leaves your PC.” That is the project’s description of its software, not an independent assessment supplied with the performance results.

For engineers, the package has two parts worth evaluating separately: the execution strategy that makes the model fit, and the local API surface that lets existing tools use it.

How the model is spread across the machine

According to Strata, Qwen3.8 Flash Next contains 24,576 experts, with 10 used per word. The runtime keeps the few thousand most frequently used experts on the graphics card, holds all experts in system RAM, and uses the CPU to work on the remaining experts at the same time.

The repository also says the SSD holds a large lookup table. It does not give a detailed allocation breakdown for that table or the GPU-resident expert set in the supplied description.

This is the mechanism behind the consumer-PC claim. The GPU does not hold the complete expert collection. System RAM supplies the larger storage pool, while the CPU participates in execution rather than serving only as a host for the GPU.

The supported hardware list is specific:

  • NVIDIA GeForce RTX 20, 30, 40 and 50 series.
  • AMD Radeon RX 7900 XT and XTX, RX 7800 XT and 7700 XT, RX 9060 XT, RX 9070 and 9070 XT, Radeon AI PRO R9700, and RX 6800 and 6900 series.
  • At least 12 GB of VRAM, plus a current NVIDIA or AMD graphics driver.
  • Windows 10 or 11, or Linux.

Strata says two or three graphics cards can share the model through multi-GPU support.

The installer configures the engine. That makes the project’s listed system requirements the starting point for deployment, rather than the GPU memory figure alone.

What the speed numbers establish

Strata reports generation speeds of 53 to 94 tokens/s on an RTX 5070 with 12 GB of VRAM, across the listed quantizations.

For the RX 9070 XT, Strata reports generation speeds of 52 to 60 tokens/s, again depending on quantized size. These are project-reported measurements, not an independently reproduced comparison between the two GPU families.

The NVIDIA table has an important qualification. Its Q2_0 row uses engine version 0.1.36, while the other rows, including Coder, use 0.1.26. The repository labels those measurements as using 4K answers and 32K prompts.

That makes the table useful as a set of reported operating points, but not a controlled comparison of quantization choices alone. Engine version changes alongside the model variant in part of the table. The supplied facts do not isolate the contribution of each change.

The separate 1.6 to 1.8× claim concerns speculative decoding. Strata describes a small helper guessing upcoming words, followed by the large model checking them in a batch. The project says this produces “the same answer, 1.6-1.8x sooner.”

That wording should remain attached to its source. The supplied material does not specify the helper model, acceptance rates, or a per-workload breakdown for the claimed gain.

For an evaluation, keep three reported quantities distinct: prompt-processing throughput, generation throughput, and the speculative-decoding speedup. They describe different measurements. The repository provides figures for all three, but the supplied facts do not provide one unified benchmark that explains how each component contributes to an end-to-end application run.

Choosing between Coder and the full model

Strata’s Coder variant removes half of the experts. According to its authors, it fits in 32 GB of RAM and reaches 91% of the full model’s SWE-bench Verified score.

The benchmark claim is relative to the full model. It is not a statement that Coder resolves 91% of SWE-bench Verified tasks. The supplied facts do not include the absolute score for either version or the evaluation setup behind that comparison.

The project also states a concrete tradeoff: Coder is weaker outside coding, including Chinese and other CJK text. That limitation belongs in model selection, not just in a footnote to the smaller memory requirement.

For a coding-focused evaluation on a 32 GB machine, Coder is the variant the repository explicitly identifies as fitting that RAM budget. For mixed coding and noncoding use, its stated quality loss is relevant even if the application mostly reaches the model through a coding agent.

The repository says available system RAM determines which model size fits, and that 64 GB runs every size. It does not provide a complete size-by-size memory table in the supplied facts, so the 32 GB baseline should not be read as a promise that every variant fits.

The remaining resource requirements still apply: a supported GPU with at least 12 GB of VRAM and about 80 GB of free disk space. Strata is free software, but the source does not provide a hardware purchase cost or an operating-cost estimate. The concrete cost information here is the machine capacity the project asks users to supply.

Connecting applications and enabling images

Strata exposes an OpenAI-compatible interface at http://127.0.0.1:8080/v1. The repository says applications and coding agents can add it as an OpenAI-compatible provider, and that any API key and any model name work for that endpoint.

It also documents an Anthropic-compatible route at http://127.0.0.1:8080/v1/messages. For Claude Code, the project identifies http://127.0.0.1:8080 as the base URL.

For Codex CLI and other applications using the OpenAI Responses API, Strata documents /v1/responses. These routes give engineers several integration surfaces to test without treating the bundled chat interface as the only way to use the runtime.

The supplied facts establish that these routes exist. They do not enumerate protocol coverage or certify feature-by-feature parity with the hosted APIs. An application’s required behavior still needs to be checked against the local endpoint.

Image input is optional during setup. Strata says users can enable it through the “Images?” option, then attach pictures in an application or use the Picture control in chat.

AMD support has a platform-specific boundary: image reading works on Linux through the CPU, but is not yet supported on Windows. That restriction does not describe AMD text inference, which the project supports on both listed operating systems.

For a multimodal application, the relevant support matrix therefore includes both GPU vendor and operating system. The text-inference hardware list alone does not establish that image input is available.

Separate local inference from training evidence

A separate Prime RL release adds native trainer support for Qwen3.8 Flash Next under the model names qwen4_exp and qwen4_exp_text. That is relevant model-support evidence, but it is not validation of Strata’s local runtime.

Prime RL describes text training and Hugging Face checkpoint conversion. Its implementation includes hyper-connection residual streams, Gated DeltaNet, indexed sparse attention, PLE N-gram embeddings and sigmoid-gated MoE.

The reported validation used the official checkpoint for a 20-step math RL run, with 16 B300 trainer GPUs and TP8 inference. The run used a 4096-token context and a 3072-token completion limit, with thinking disabled.

Prime RL reports median warmed-up steps of 56.1 seconds and median broadcasts of 41.5 seconds. Across the 20 steps, all mean mismatch KL values were below 0.015, with a reported maximum of 0.001570820 and mean of 0.000701912.

Those are run-specific engineering measurements. They are neither consumer-PC inference results nor comparisons against named competitor models.

The release also qualifies its validation history. Tests were not rerun after the merge, and the reported results came from earlier commits or jobs.

The resulting evidence boundary is clear. Strata supplies the local deployment design, hardware requirements and inference claims. Prime RL supplies a separate text-training implementation and a bounded validation run. Engineers evaluating Strata should use the former for deployment expectations and should not treat the latter as independent confirmation of its speed or answer-equivalence claims.

Try it

Start at https://github.com/Niko1221/Strata. Download the project and follow its Windows or Linux setup instructions, then open the local interface at http://127.0.0.1:8080. No executable steps are listed because the supplied commands list is empty.

Sources