FeSens recently unveiled openTPU, an open-source AI accelerator project where components of the hardware design were developed by AI. The project includes the RTL, ISA, a Python simulator, a compiler for a small Python kernel language, and a profiler. This accelerator can run models such as Qwen3, LFM2.5, and Qwen3.5 on a Kintex-7 PCIe card, achieving approximately 14 tokens per second with about 70% of peak memory bandwidth on an Inspur YPCB-00338 board.

The openTPU project aims to provide an accessible platform for individuals interested in learning about and experimenting with AI accelerator design. The entire environment, excluding the physical card, can operate on a laptop, allowing users to interact with models on the simulator without dedicated hardware. The project's creator noted that "cheap" Inspur boards are available for those who wish to implement the design on physical hardware.

This development aligns with a broader industry trend where AI is increasingly used in hardware design. Companies like OpenAI and Synopsys have partnered to create specialized AI models for semiconductor design, aiming to automate and accelerate the process of creating more sophisticated chips. OpenAI, for instance, utilized its own AI models to expedite the development of its custom Jalapeño Intelligence Processor, an ASIC specifically engineered for AI inference.

The application of AI in hardware design is complex and domain-specific, but it offers opportunities for innovation. Large language models and agentic systems are beginning to reshape the hardware design stack, with Register-Transfer Level (RTL) coding serving as a proving ground. These AI agents can integrate with existing hardware design tools, interpret hardware-specific languages, and automate various stages of the design and verification process.

The goal of these initiatives is to enhance chip design quality and productivity, allowing engineers to evaluate more design options and deliver more advanced silicon faster. This shift represents a move towards AI agents that can perform autonomous repair and synthesis using planning and feedback mechanisms.

The openTPU project is distinct from Google's Tensor Processing Unit (TPU), although it shares a similar name. The UCSB ArchLab also has an open-source re-implementation of Google's TPU, which is based on details from Google's published papers. Google's TPUs are custom ASICs designed to accelerate the inference phase of neural network computations, and they have been deployed in data centers since 2015.

The increasing demand for efficient AI inference hardware is driving many companies to develop custom silicon. This includes major players like Google with its TPUs, Amazon with Trainium and Inferentia, Meta with MTIA, and Microsoft with Maia. These custom chips aim to reduce reliance on general-purpose GPUs and optimize performance, cost, and energy efficiency for specific AI workloads, particularly large language model inference.