Siemens Digital Industries Software has announced Catapult™ AI NN software for high-level synthesis (HLS) of neural network accelerators on application-specific integrated circuits (ASICs) and systems-on-chips (SoCs). Catapult AI NN is a complete solution starting with a neural network description from an AI framework, translating it to C++ and synthesizing it into an RTL accelerator in Verilog or VHDL for implementation on silicon.
Catapult AI NN combines hls4ml, an open source package for machine learning hardware acceleration, with Siemens' Catapult™ HLS software for high-level synthesis. Developed in close collaboration with the U.S. Department of Energy laboratory Fermilab and other key contributors to hls4ml, Catapult AI NN addresses the unique requirements of machine learning accelerator design for power, performance and area on custom silicon.
“The process of handing off and manually converting neural network models to hardware implementations is highly inefficient, time-consuming and error-prone, especially when creating and validating hardware accelerator variants for specific performance, power and area,” said Mo Movahed, vice president and general manager of High Level Design, Verification and Power, Siemens Digital Industries Software. “Enabling scientists and AI experts to leverage industry-standard AI frameworks such as neural network model design, and seamlessly integrating these models into power-, performance- and area-optimized hardware designs opens up a whole new realm of possibilities for AI and machine learning software engineers. Our new Catapult AI NN solution now enables developers to automate and implement neural network models that achieve optimal PPA simultaneously during the software development process, ushering in a new era of efficiency and innovation in AI development.”
As runtime AI and machine learning tasks move from data centers to everything from consumer electronics to medical devices, the requirement for “right-sized” AI hardware is rapidly growing to minimize power consumption, reduce cost, and maximize end-product differentiation. However, most machine learning professionals are more comfortable working with tools like TensorFlow, PyTorch, or Keras than with synthesizable C++, Verilog, or VHDL. Traditionally, there has been no easy way for AI professionals to accelerate machine learning applications on right-sized ASIC or SoC implementations.
The hls4ml initiative aims to fill this gap by generating C++ from neural networks written in AI frameworks such as TensorFlow, PyTorch, and Keras. The generated C++ can then be deployed to FPGA, ASIC, or SoC implementations.
Catapult AI NN extends the capabilities of hls4ml to ASIC and SoC design. It includes a dedicated library of specialized C++ machine learning functions tailored for ASIC design. These functions allow designers to optimize PPA by making latency and resource trade-offs between alternative implementations of C++ code. Additionally, designers can now evaluate the impact of different neural net designs to determine the best neural network structure for their hardware.
“Particle detector applications have some extremely stringent edge AI constraints,” said Panagiotis Spentzolis, associate director for emerging technologies at Fermi National Accelerator Laboratory. “Through our collaboration with Siemens, we were able to develop Catapult AI NN, a synthesis framework that allows scientists and AI experts to leverage their expertise without having to become ASIC designers. What's more, this powerful new framework is also ideal for seasoned hardware experts.”
