TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Strata GitHub project says its software can run the 125-billion-parameter Qwen3.8 Flash Next model locally on supported consumer PCs, with at least 12 GB of GPU memory and 32 GB of system RAM. Its published RTX 5070 test reached 94 tokens per second for answer generation with one compressed model setting; the supplied results do not establish 100 trillion tokens per second or a benchmark on an RTX 4090.
The Strata open-source project says its software can run the 125-billion-parameter Qwen3.8 Flash Next model on a consumer PC, using supported NVIDIA or AMD graphics cards with at least 12 GB of video memory. Its published tests report up to 94 tokens per second generating answers on an RTX 5070, but the supplied report does not document an RTX 4090 test or support the claim of 100 trillion tokens per second.
Strata’s published performance table covers two gaming PCs: one with an RTX 5070, 12 GB of VRAM, a Ryzen 5 7600 and 64 GB of system RAM; the other with an RX 9070 XT, 16 GB of VRAM, a Ryzen 9 3900X and 47 GB of RAM. On the NVIDIA system, answer-generation results ranged from 53 to 94 tokens per second across the listed model settings. The AMD results ranged from 44 to 60 tokens per second. These are project-reported measurements, not an independently verified benchmark.
The project distinguishes generating a response from processing an incoming prompt. For the NVIDIA PC, it reports prompt-processing rates of 1,620 to 2,650 tokens per second, measured with 32,000-token prompts; the listed answer-generation tests used 4,000-token answers. The table notes different engine versions for the NVIDIA results: version 0.1.36 for Q2_0 and 0.1.26 for the other listed configurations. Those separate workloads and software versions matter when interpreting or comparing the figures.
Running the model locally requires substantial memory and storage. Strata lists 32 GB or more of system RAM, about 80 GB of free disk space and a supported GPU with at least 12 GB of VRAM. Its setup downloads a model of about 70 GB, then loads roughly 35 to 55 GB into system memory while reserving some for the GPU. The project says the software runs on Windows or Linux and that model activity stays on the user’s PC.
Large Models on Gaming PCs
If the reported results hold on other supported systems, Strata could make a large language model available without sending prompts to a remote service. That can appeal to users who want local control over data, offline access after installation, or a way to connect a model to coding tools and other applications on their own computer. The project says Strata can chat, write code, interpret images and work with coding agents, though the supplied material does not provide independent evaluations of those capabilities.
The performance figures also clarify the trade-offs behind local use. 94 tokens per second is the reported peak for one RTX 5070 configuration, not a universal speed for every PC or model setting. Strata says smaller, more compressed model versions run faster, while larger versions retain more information and offer higher quality at a cost in speed and memory. The practical result depends on the user’s hardware, chosen model file, and workload.
As an affiliate, we earn on qualifying purchases.
What the Published Tests Measure
The request refers to running the model on an RTX 4090 at “100T/s.” In common performance reporting, tokens per second are abbreviated as tokens/s or tok/s; “100T/s” is ambiguous and, read literally as 100 trillion tokens per second, is not a result in Strata’s supplied tables. The documentation instead reports speeds in tokens per second, with the fastest listed result at 94 tokens per second on an RTX 5070.
Strata also makes a prediction about an RTX 3090 with 24 GB of VRAM, saying it should generate about 100 to 140 tokens per second. That is a projected speed, not a reported RTX 4090 measurement. The project’s table does not show an RTX 4090 result, and its listed systems are the RTX 5070 and RX 9070 XT machines. The evidence supplied therefore supports a report about local inference and those test systems, not a confirmed 4090 benchmark.
“Nothing leaves your PC.”
— Strata project documentation
high performance gaming PC with 12GB VRAM
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The RTX 4090 Claim Is Unverified
The supplied documentation does not say when the tests were conducted, provide raw benchmark logs, or report a run on an RTX 4090. It also does not define “100T/s.” If that means 100 tokens per second, the figure is not established for a 4090 by the included test table; if it means 100 trillion tokens per second, it is not supported by the reported units or results. No independent benchmark or quality comparison is included in the source material.
Other relevant details remain open as well. The table shows results for selected compressed model settings and particular prompt and answer lengths; it does not establish how speeds change across longer conversations, different software versions or other hardware. Strata points readers to additional documentation and community results, but those materials are not included here. Claims about privacy, image understanding and coding performance should likewise be read as project descriptions unless verified separately.
large memory gaming PC for AI models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
More Hardware Results Needed
Readers seeking to verify the headline claim would need a reproducible RTX 4090 benchmark that specifies the exact model quantization, Strata and engine versions, context length, prompt and output lengths, and measurement method. The project’s documentation directs users to detailed tables and community results, which may add tests beyond the two systems summarized in the supplied report.
For now, the confirmed development is that Strata publishes a way to run this large model locally on supported consumer hardware, alongside project-reported results for an RTX 5070 and an RX 9070 XT. Whether a 4090 achieves 100 tokens per second under a particular setup remains unconfirmed by the material provided; the “100T/s” wording should not be treated as a verified benchmark.
GPU with 12GB VRAM for AI processing
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Does the source confirm Qwen3.8 Flash Next running on an RTX 4090 at 100 tokens per second?
No. The published table supplied here covers an RTX 5070 and an RX 9070 XT, not an RTX 4090. It does not verify a 100-token-per-second result for the 4090.
What does “100T/s” mean in the headline claim?
The source material does not define it. Strata reports performance in tokens per second; read literally as 100 trillion tokens per second, “100T/s” is not supported by the project’s figures.
What speed did Strata report on an RTX 5070?
Strata’s table lists up to 94 tokens per second for answer generation on an RTX 5070 with the Q2_0 setting. It also lists slower results for other settings and separate prompt-processing speeds.
What hardware does Strata list as necessary?
The project specifies a supported NVIDIA or AMD GPU with at least 12 GB of VRAM, at least 32 GB of system RAM and about 80 GB of free disk space. It lists Windows and Linux support.
Does local operation mean the model has been independently shown to protect data?
Strata says model activity stays on the user’s PC. The supplied material does not include an independent security or privacy audit, so that description should not be mistaken for one.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
