AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

OpenTPU is an open-source FPGA accelerator project whose developers report running 10 AI models with real weights on an Inspur card built around a Xilinx Kintex-7. They say the hardware produced the same tokens as the project’s simulator, while measured throughput varied by model and configuration. The results are project-reported benchmarks, not an independent evaluation.

The open-source OpenTPU project has published results showing its accelerator running 10 AI models with real weights on an Inspur FPGA card, including configurations ranging from small models to models with tens of billions of parameters. The developers report that outputs from the physical card matched the project’s bit-exact simulator token for token; the results are project-reported and have not been independently verified in the supplied material.

OpenTPU runs on an Inspur YPCB-00338 card with a Xilinx Kintex-7 xc7k480t FPGA and two DDR3 memory channels. Its reported design uses a four-column systolic matrix unit, a streaming engine and LiteDRAM controllers. The developers give the memory’s peak bandwidth as 17.1 GB/s and say measured traffic while decoding reached 14.1–16.1 GB/s across listed configurations.

The published table includes LFM2.5-230M, Qwen3-0.6B, Gemma 4 E2B, Phi-4-mini and models up to Qwen3.5-4B and Gemma 4 E4B. Decode rates depend on model size, weight format and measurement method: the table reports 59.0 device tokens per second for LFM2.5-230M in int8 and 3.99 for Phi-4-mini in int8. For LFM2.5-230M using 4-bit weights with an int8 output head, it reports 85.8 device tokens per second. The project also supplies wall-clock rates, which include host work and are lower in some cases.

The benchmark method matters. The main table measures 64 greedy tokens after a 512-token prompt, with the host selecting each token; “device” counts accelerator cycles, while “wall” includes the host. Prefill is measured on the accelerator using the 512-token prompt. The project separately reports a card-controlled decode loop, where the FPGA selects each token, and a streamed-logits test. These figures use different procedures and should not be treated as directly interchangeable. The repository provides hardware, an instruction set, a simulator, a kernel language and compiler, and host software.

At a glance
reportWhen: Results dated September 29 to October 1…
The developmentThe OpenTPU project has published performance results for an AI-designed, open-source accelerator running language models on a Kintex-7 FPGA card.

What FPGA Results Show

The results offer a concrete example of an open hardware project running modern language models on a comparatively specialized, older FPGA platform rather than relying on a conventional graphics card. That matters to hardware researchers and developers who want to inspect how an accelerator is built, from its instruction set and memory system through to host software, and compare actual board behavior with a simulator.

OpenTPU’s figures also show the limits of the demonstration. The reported decode rates vary substantially with model size and quantization, and the project says decoding is DRAM-bound. For the listed tests, its measured memory traffic reaches 82% to 94% of the stated DDR3 peak. These numbers describe this card and the project’s measurement setup; they do not establish that the design will match commercial accelerators or perform similarly on other boards.

The work is also relevant as an experiment in AI-assisted hardware design. The project describes itself as bringing lessons from “auto-arch-tournament” to accelerators and poses the question of how far AI agents can go in hardware design. That framing is the project’s stated motivation, not evidence in the benchmark table that an AI system independently designed every part of the accelerator.

Amazon

FPGA AI accelerator card

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the OpenTPU Project

The repository presents OpenTPU as both an accelerator and a learning project. It places the SystemVerilog design, instruction set, bit-exact simulator, kernel tools and host software in one small codebase. The aim, according to the project, is to let readers follow operations from a Python matrix multiplication down to hardware signals, rather than treating the accelerator as a closed device.

The reported board image has changed over time. The project compares its newer production image with an earlier build that used Xilinx MIG, a two-column matrix unit and a 120.755 MHz clock. It says the current image runs at 133.33 MHz, calibrates both DDR3 channels at startup using a small CPU in the memory core, and completes that calibration in 12 seconds without host involvement. For prefill, the project reports gains of 1.3 to 2.0 times over the earlier image, depending on the model and configuration; it says decode performance stayed within 2.3% of the prior image.

Some configurations use compressed weights to fit models into the card’s memory. The project says Gemma 4 E2B’s per-layer embedding tables remain on the card and that its int8 configuration does not fit. For Gemma 4 E4B, a 2.95 GB table stays on the host, which transfers an 11 KB row to the card for each token. The repository documents these implementation choices alongside its benchmark results.

“An open-source AI accelerator, developed by AI.”

— OpenTPU project description

Amazon

Xilinx Kintex-7 FPGA development board

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Testing Remains Absent

The supplied source is the project’s GitHub repository, and the performance figures are developer-reported. It does not provide an independent reproduction or external review of the measurements. The report also does not establish how the accelerator compares with other hardware under a shared test setup, so direct performance rankings would be unsupported.

The phrase “developed by AI” is not accompanied in the supplied material by a detailed account of which design decisions or implementation stages were made by AI agents, what human review occurred, or how contributions were measured. The available figures show reported operation and simulator agreement, but do not by themselves answer those questions. The benchmarks cover named models and specified setups, not every model or use case.

Results for models that exceed the card’s 4 GiB capacity involve host-side transfers. The project reports that larger mixture-of-experts models can stream experts from host storage, but the supplied excerpt cuts off part of the final measurement description after giving a transfer rate for Qwen3.5-35B-A3B. Details of that test and its comparison conditions are therefore incomplete here.

Amazon

open-source AI hardware FPGA

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reproduction and Further Results

The project’s repository is the immediate place to check for revised measurements, build details and instructions for reproducing the tests. Its documentation includes board setup, model-specific notes and the benchmark script, giving technically equipped readers material to inspect alongside the headline figures.

For the results to support wider comparisons, independent testing would need to reproduce the same model versions, prompts, weight formats, board configuration and timing method. It would also help clarify the role of AI agents and human developers in the design. The supplied report does not announce a scheduled external evaluation or a future hardware release; whether either will happen remains unknown.

Amazon

FPGA machine learning accelerator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is OpenTPU?

OpenTPU is an open-source AI accelerator project. Its repository contains a SystemVerilog hardware design, an instruction set, a simulator, compiler tools and software for operating an FPGA card.

What hardware does it run on?

The reported tests use an Inspur YPCB-00338 card with a Xilinx Kintex-7 xc7k480t FPGA and two DDR3 channels. The project gives the memory’s peak bandwidth as 17.1 GB/s.

Did the FPGA match the simulator?

The project says the card produced the same tokens as its simulator, bit for bit, for the tested configurations. That is a claim from the project’s own report; the supplied source does not include independent verification.

Are the performance results independently verified?

Not in the supplied material. The figures come from OpenTPU’s repository and should be treated as project-reported benchmarks unless an independent reproduction is published.

Does “developed by AI” mean AI built the whole accelerator?

The project describes the accelerator as developed by AI and asks how far AI agents can go in hardware design. The supplied material does not detail which tasks AI agents performed or how much human design and review was involved.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Belongs On An AI Automation Desk In 2026?

A desk setup checklist names a Dell workstation and Jetson board, while leaving accelerators, servers, docks and storage to workload-based choices.

Opus 5.5 Agents Discover Two Room-temperature Magnetic Semiconductor Candidates

Vals AI reports two computationally predicted candidates for room-temperature magnetic semiconductors, but experimental confirmation is still needed.

What Rymvard’s SAP HANA Memory Means For The Capacity Ledger

Rymvard’s early-access 0.23 release adds HANA memory accounting and replication takeover checks, tested on a simulated estate rather than a customer deployment.

Claude, Change The “Add To Cart” Button To Blue

A recent trend indicates a change in the ‘Add to Cart’ button color to blue, sparking increased search interest; details remain unconfirmed.