Quick Answer: To run Qwen 125B on consumer hardware at up to 100 tokens per second, use the Strata inference engine. Strata utilizes aggressive IQ quantizations (like IQ2_XS) and offloads the model across your GPU's VRAM and system RAM. An RTX 3090 or RTX 4090 paired with 64GB of system RAM provides the ideal setup for local execution.

Running a 125-billion-parameter model locally used to require an enterprise server rack or a massive cloud budget. If you tried loading a model of that scale on a standard gaming rig, your system would instantly crash with an out-of-memory error. But things have changed. With the release of the Strata inference engine, you can now run Qwen 125B on consumer hardware like an RTX 4090 or RTX 3090 at speeds exceeding 100 tokens per second. Here is how to set it up and what to expect.

The Architecture: How Strata Fits a 125B Model into a Gaming PC

To understand how we can run Qwen 125B on consumer hardware, we have to look at how modern Mixture of Experts (MoE) architectures function. Qwen3.8-Flash-Next is not a dense 125B model where every single parameter is active for every token. Instead, it routes tokens to specific "experts" dynamically. This means only a fraction of the 125 billion parameters are computed per token, which drastically lowers the active compute requirement.

However, the entire model must still reside in memory. This is where the Strata inference engine comes in. Strata uses advanced Importance Quantization (IQ) formats, which compress the model weights down to 2 or 3 bits per parameter.

Here is a counter-intuitive finding that contradicts popular local LLM advice: most guides tell you to never go below 4-bit quantization because of severe perplexity degradation. While that holds true for smaller 8B or 14B models, it does not apply to massive 100B+ models. A 125B model quantized to 2 bits (like IQ2_XS) retains shocking coherence and easily outperforms a fully unquantized 14B model on complex reasoning tasks. The sheer scale of the base model acts as a buffer against quantization noise.

Strata splits these quantized weights between your GPU's fast VRAM and your system's slower DDR4/DDR5 RAM. By utilizing a highly optimized unified memory mapping system built on top of the GGML framework, Strata pulls off a feat that standard runtimes struggle to achieve without massive latency penalties.

That said, there's a real catch here when it comes to your physical hardware configuration.

Hardware Requirements: VRAM, RAM, and the SSD Bottleneck

To achieve acceptable generation speeds, your hardware must meet specific thresholds. You cannot simply throw any old components together and expect 100 tokens per second. The primary bottleneck in local LLM execution is memory bandwidth, not raw compute power.

Your GPU's VRAM is the fastest memory pool available. An RTX 4090 offers 24GB of GDDR6X VRAM with over 1 TB/s of bandwidth. Your system RAM (DDR5) is significantly slower, offering around 60-80 GB/s. When Strata splits the model, it keeps the most frequently accessed layers (and the active KV cache) in VRAM, while offloading the rest to system RAM.

According to benchmarks from the Strata GitHub repository, here is what you need to run the model:

  • Graphics Card: NVIDIA RTX 3090 or RTX 4090 (24GB VRAM) is highly recommended for 100+ T/s. However, cards with 12GB VRAM (like the RTX 5070 or RTX 4070) can still run the model at 50-60 T/s using tighter quantizations.
  • System RAM: 64GB is the sweet spot. While 32GB can run the specialized "Coder" variant, you need 64GB to load the full 125B model with room left for your operating system.
  • Storage: An NVMe SSD with at least 80GB of free space. Running this setup on a mechanical hard drive will result in agonizingly slow load times.

If your system RAM is insufficient, your operating system will begin paging memory to your SSD. This creates a massive bottleneck. When memory pages spill over to the SSD, token generation speeds drop from a blazing 100 T/s to a painful 7 T/s.

This next part trips people up every time, so let's walk through the exact installation process to avoid these performance traps.

Step-by-Step Guide to Run Qwen 125B on Consumer Hardware

Setting up Strata is surprisingly straightforward thanks to its one-click installer, but you must configure it correctly to maximize performance.

Step 1: Clone and Prepare the Repository

First, download the Strata package or clone the repository directly from GitHub. Ensure your graphics drivers are updated to the latest version before proceeding.

git clone https://github.com/Niko1221/Strata.git
cd Strata

Step 2: Run the Installer

  • On Windows: Double-click START-HERE.bat.
  • On Linux: Run ./setup.sh in your terminal.

The installer will automatically scan your system's GPU, CPU, and available RAM to recommend the optimal quantization size.

Step 3: Select Your Model and Context Size

When prompted, select the model variant that matches your system RAM. If you have 64GB of RAM, select IQ2_XS or IQ3_XXS. For context size, start with 32K. While the model supports up to 128K context, larger context windows require significantly more memory for the KV cache, which can push your system into swap space.

Step 4: The First-Run Allocation

Once you confirm your settings, Strata will download the model weights (approximately 70GB). When the model starts for the first time, your PC may become slow or unresponsive for 1 to 3 minutes.

Do not close the window. Strata is allocating and page-locking up to 55GB of system RAM to prevent the operating system from swapping it to disk. Once this process completes, your browser will automatically open the local dashboard at http://127.0.0.1:8080.

Step 5: Connect Your Tools

Strata hosts an OpenAI compatible local API on port 8080. You can point your favorite development tools, such as Cursor, Claude Code, or any local chat interface, directly to this endpoint:

  • Base URL: http://127.0.0.1:8080/v1
  • API Key: Any dummy string (e.g., local)

But which specific quantization should you actually choose for your daily workflow?

Quantization Comparison: Q2_0 vs. IQ3_S vs. Coder

Choosing the right quantization is a direct trade-off between speed and intelligence. If you are using the model primarily for software development, the specialized "Coder" variant is often your best bet.

Model VariantRAM RequiredSpeed (RTX 4090)Best Use Case
Q2_048 GB100-140 T/sMaximum speed, general drafting
IQ2_XS48 GB80-100 T/sBalanced reasoning, daily use
IQ3_S64 GB55-70 T/sComplex logic, deep analysis
Coder32 GB60-80 T/sCoding agents, SWE-bench tasks

The Coder variant is particularly interesting. The developers of Strata removed half of the non-essential experts from the base model to create this version. It retains 91% of the full model's SWE-bench Verified score while fitting comfortably on systems with only 32GB of RAM. However, it performs poorly on non-coding tasks and multilingual inputs.

While these numbers look great on paper, real-world execution introduces some unique edge cases that you must monitor.

Real-World Performance and the Offloading Edge Case

When you run a model of this scale, you will eventually encounter the limits of consumer hardware. The most common failure mode is memory thrashing during long-context conversations.

When you start a chat, the memory footprint is relatively small. But as the conversation grows toward 32K or 64K tokens, the KV cache (the memory used to store past conversational context) expands exponentially. If your combined VRAM and system RAM are filled to 99% capacity, the addition of a few thousand context tokens will force the system to offload active KV cache pages to your SSD swap file.

When this happens, your generation speed will instantly crater from 100 T/s to less than 2 T/s. The CPU becomes entirely blocked waiting for disk I/O operations. To prevent this, always leave a 10% memory buffer when configuring your context window in Strata, or limit your maximum context to 32K if you only have 48GB of system RAM.

Additionally, keep an eye on your GPU thermals. Sustained generation at 100 T/s keeps the GPU memory controller under constant load, which can spike VRAM junction temperatures on some RTX 3090/4090 cards. Ensure your PC has adequate ventilation.

Frequently Asked Questions

What is the Strata inference engine?

Strata is a free, open-source local inference engine designed to run massive LLMs like Qwen 3.8 Flash Next on consumer-grade hardware. It achieves high speeds by combining aggressive quantization with optimized CPU/GPU memory offloading.

How to run 125B model on RTX 4090 without running out of memory?

To run Qwen 125B on consumer hardware like an RTX 4090, you must use Strata's IQ2_XS or Q2_0 quantization. This compresses the model weights so they can be split between your 24GB of VRAM and 64GB of system RAM.

Why does offloading slow down LLM generation?

Offloading slows down generation because system RAM (DDR5) and PCIe bus bandwidth are significantly slower than GPU VRAM (GDDR6X). When the engine has to wait for the CPU to process layers stored in system RAM, it creates a processing bottleneck.

Can I run Strata on an AMD graphics card?

Yes. Strata natively supports AMD Radeon cards with 12GB of VRAM or more, such as the RX 7900 XT or RX 9070 XT, using the same one-click installer script.

Summary and Next Steps

Running a 125B parameter model locally is no longer restricted to enterprise datacenters. By deploying the Strata inference engine on an RTX 4090 or RTX 3090, you can achieve production-grade speeds of up to 100 tokens per second entirely offline. This setup gives you complete data privacy and eliminates recurring API costs.

Try setting up Strata on your local machine this week and monitor your memory allocation during long chats. If you want to optimize your setup further, read our breakdown of local LLM quantization formats next.