AI in the Field · Oct 9, 2026
AI wrote everything from chip RTL to compiler, hitting 85 tokens per second on a used FPGA of about $300. What set the speed was its pass/fail verification
We followed FeSens' solo project, openTPU. The numbers are interesting, but on HN many commenters said "230M is too small" and asked "how much did a human do?"
Haru Misaki · Field Correspondent

Watch the video
Key points
- FeSens' openTPU puts AI-written RTL, an instruction set, a simulator, a compiler and PCIe driver software in a single repository. It runs 10 models, including Qwen3, on a used Kintex-7 card, and its output matches the simulator bit for bit
- LFM2.5-230M (4-bit) decodes at 85.8 tok/s. Models from 2.6B to 3.8B drop to 3.75–12.09 tok/s, and the project gives no comparison against CPUs or GPUs
- In the author's previous project, only 10 of 73 hypotheses were adopted, after passing formal verification, simulation and logic synthesis. The author's takeaway: "the verifier you plug into the loop is what decides the outcome"
Hi, I'm Haru. This was the project that made me say "wait, they had AI do all of that?" this week. Independent developer FeSens had AI write everything from the hardware design to the software that runs on it, then ran language models on a used card retired from a data center. I worked through the GitHub repository, the Hacker News thread and a roundup article. I haven't tried the hardware myself, so everything here comes from public information and what people have said online. 🔍
What was built? AI wrote everything from the hardware design to the compiler
First, what openTPU is. In short, the author had AI design a compute circuit built only for AI inference, which means running a trained model to get its output. According to the README, AI wrote the following:
- RTL written in SystemVerilog (the design description of how the circuits inside the chip behave)
- An ISA (the set of instructions the hardware understands, like the machine-code rules of a CPU)
- A simulator that matches the hardware bit for bit (software that imitates the real circuit on a computer)
- A kernel language and compiler (a dedicated language for describing computations, and software that turns it into instructions for the hardware)
- A profiler (a tool for finding where time is spent)
- Driver software for PCIe (the interface that connects a computer to expansion cards)
Chip design and the software that runs on the chip are usually handled by separate teams. Having them all in one repository is interesting on its own. The target is the Inspur YPCB-00338, a card retired from data centers, with a Xilinx Kintex-7 xc7k480t and 4GiB of DDR3 memory. Because it is an FPGA, a chip whose internal circuits can be rewritten any number of times, the AI-written design can be loaded straight onto it and tested.
In his Hacker News post, the author (fsbonetto on HN) wrote that the used board costs about $300 and is popular with hardware hobbyists. A separate roundup article put the price at $200. The project runs 10 models on it, including Qwen3, Qwen3.5 and LFM2.5, and the output matches the simulator bit for bit, according to the author. It does this on a single used card, not on dedicated hardware costing hundreds of thousands of yen, and I find that really appealing. 🛠️
openTPU in numbers: what's behind 85.8 tok/s
Here are the README's figures. tok/s is the number of tokens (chunks of text) produced per second. Decode is the speed at which the model writes its answer one token at a time. Prefill is the speed at which it reads in the prompt at the start.
- LFM2.5-230M (4-bit): decode 85.8 tok/s, prefill 335.4 tok/s
- LFM2.5-230M (int8): decode 59.0 tok/s, prefill 295.6 tok/s
- Qwen3-0.6B (4-bit): 31.3 tok/s
- Qwen3.5-0.8B: 16.3–23.3 tok/s
- Models from 2.6B to 3.8B: 3.75–12.09 tok/s
4-bit and int8 describe the precision at which the model's numbers are stored. 4-bit is coarser, so the data is smaller and less has to be read from memory, which makes it faster. The int8 figure of 59.0 tok/s was also reported by AI Weekly.
The hardware figures are detailed too. Memory is dual-channel DDR3-1066, with peak bandwidth (the maximum amount of data that can move to and from memory per second) of 17.1 GB/s. The design uses 82–94% of that, so it gets close to the limit. The clock (the rate at which the circuit ticks) is 133.33 MHz, and the timing margin, WNS, is +0.032 ns. That means signals only just arrive in time. The main branch has 1,439 commits. The repository was created on September 24, 2026, and as of October 8 it had 537 stars and 36 forks. That is a lot of work for about two weeks. 📊
The speed came from a pass/fail system, not from how clever the AI was
This is the part I read most closely. On HN, the author said he reused the approach from an earlier project in which he used AI to develop a RISC-V CPU core (RISC-V is an open CPU instruction set anyone can use). Through many rounds of agent-driven improvements, he took the system from a few tokens per second at first to more than 80 tokens per second on small models.
That earlier project is auto-arch-tournament. It starts from a textbook five-stage pipelined RV32IM core. In each round, Claude or Codex proposes a hypothesis ("changing this should make it faster") along with an implementation. Every proposal then has to pass three gates:
- Formal verification with riscv-formal (a mathematical check that the instruction-set rules are not broken)
- Simulation with Verilator (running the design to check the results)
- Logic synthesis with yosys/nextpnr (checking that it can actually be built as a circuit)
A change is merged only if it passes all three and also raises CoreMark/MHz, a benchmark of how much work is done per clock cycle. Of 73 hypotheses, 10 were adopted. The run took 9 hours 51 minutes, and CoreMark/MHz rose 92% to a final 2.91. More than 70% of proposals were rejected. openTPU extends this setup from a CPU to an inference accelerator.
The author's takeaway: "What decides the outcome isn't the loop itself, but the verifier you plug into it." To me this is exactly the harness engineering being talked about right now, meaning the design of systems that constrain and steer autonomous AI. However clever the AI's proposals are, if they can't be judged mechanically on whether they're correct and whether they're faster, you just pile up unusable changes. With a strict judge like a bit-exact simulator, though, the AI can fail dozens of times without harm. That really made sense to me. 💡
What isn't working, and the strongest doubts on HN
This is the important part. The thread reached 342 points and 400 comments, but the reaction wasn't all praise.
The first issue is model size. HN user Retro_Dev pointed out that the headline number comes from "a 230 million parameter model, which is very small." That's true: the README shows models from 2.6B to 3.8B dropping to 3.75–12.09 tok/s. Neither the thread nor the roundup article has any comparison with CPUs, GPUs or commercial inference hardware. So whether 80 tokens per second is fast honestly can't be judged from what's available now. 🤔
Next is the question of how much a human contributed. AI Weekly reported that the README doesn't say how much was written by AI versus humans, or which models the agents used. On HN, some commenters suggested that direction from an experienced person may have made a big difference. Others questioned the correctness of the floating-point math (arithmetic on decimal numbers), and one quipped: "Apparently LLMs don't need correct floating-point arithmetic."
The author lays out the limits clearly in the README. Decode is capped by DRAM bandwidth (82–85%), so no amount of circuit work can make memory reads keep up. The timing margin is only +0.032 ns. MoE models larger than 4GB need their experts streamed in from storage on the host. Honestly, I appreciated that he listed the weaknesses himself.
The discussion went further, to the idea of baking the weights directly into silicon. dmitrygr and others calculated that for a 27B-class model, the weights alone would far exceed the reticle limit (the maximum area that can be exposed onto a chip at once, about 800mm²). zdragnar wrote that by the time the chip was designed, the state-of-the-art models would already have been replaced.
My overall take: if you read openTPU as "AI built a fast chip," there aren't enough comparison numbers to judge it. Read as an example of how, with a strict judge in place, you can have AI build everything from the hardware design to the compiler in two weeks, it's very valuable. Next I'd like to see comparisons on the same board and against GPUs, along with a record of where humans stepped in. 🚀
Editorial cartoon
