Strata

7 October 2026

The big news in local LLMs this week is a project called Strata, a new inference engine for Linux and Windows tailored to run Qwen3.8-Flash-Next on reasonably available hardware. I’ve been trying it out the last couple of days and I can confirm: this is the real deal. This is the good stuff.

Strata monitoring screen

Let’s compare this with what I was running a few days ago:

Enginellama.cppStrata
ModelQwen3.8-27BQwen3.8-Flash-Next
Parameters27B125B + 51B n-gram
Prefill (tok/s)1500–10003000–1500
Output (tok/s)50–30140–100
QuantisationQ4_K_MIQ3_XXS
Context (my setup)131072 @ Q8131072 @ Q8
VRAM (my setup)23 GB23 GB
System RAM0 GB39 GB

It scales as far down as 32 GB system RAM and 12 GB VRAM; however, you take a bit of a hit on quantisation. I got lucky—I recently upgraded my new PC to 64 GB RAM, so between that and its laptop-edition 5090 I’m able to comfortably fit a quantisation that doesn’t have a major drop in quality. Now I have over double the token rate and a much stronger model. This is even much better tok/s than you get today as an OpenAI subscriber running GPT-6.1 Sol. It’s absolutely wild.

There are a few things going on here that make such a jump possible. Strata is fundamentally using more of my computer than Qwen3.8-27B did. I’m using a bunch of extra RAM outside my GPU and CPU is involved too. The model architecture is different; apart from being a MoE design which means faster computation, many additional weights live on SSD in the n-gram table for fast lookup. The result is an optimisation of speed and intelligence that’s sympathetic to the spread of resources available on a fast PC, rather than assuming you’re racking up a bunch of NVIDIA Blackwell cards.

Now obviously, to get “Opus at home” I’m bringing some pretty serious equipment to bear. This is not your average laptop. At the same time, it’s “just” a high-end gaming laptop. That’s awesome for anyone who has access to that kind of hardware, and above all it shows that the future is bright for local inference. Models are only going to get more intelligent and more efficient from here.

I suppose things will get tougher for OpenAI and Anthropic in the years ahead. If you can get reasonably fast tokens on reasonably priced equipment, many people won’t want to pay big bucks to outsource the problem to the cloud. To their credit, I think companies like OpenAI understand that and it makes sense why they’re investing in more wrapping up AI in convenient packages like dots and their new intelligent UI system. Text chat isn’t going to be the dominant UI forever. Open source developers and the local inference crowd will probably always be playing catch-up. Fortunately, we have the tools to do so if we wish.

So, just how good is this model? I haven’t tried to run any objective benchmarks but I’ve used it for a number of tasks. It’s added some new features to my static site generator, fixed some bugs in a status page for me to see if anyone’s using my game servers, and performed a bit of security research. It successfully found and exploited an issue, nothing earthshattering, but not a bad effort at all for an hour or two of analysis running on my GPU. It’s fast and effective.

There are some benchmarks available: this analysis shows that IQ3_XXS only falls slightly short of baseline BF16 performance. The trick here is that it can be smaller than a flat Q4 quant by tuning weights adaptively—more important weights get more bits while less important weights get fewer. If Artificial Analysis is to be believed, this model lands somewhere around Sol 5.6 medium-high, and considerably ahead of Opus 4.6. Not shabby at all.

I was starting to think we’d have to stay around the 27B mark if we wanted both quality and speed. I’m very pleased to see that’s not true. What will be next?


Serious Computer Business Blog by Thomas Karpiniec
Posts RSS, Atom