Po2-QAT: Multiplier-Free Hardware Acceleration for Efficient AI

Faculty Mentor Information

Dr. Omiya Hassan, Boise State University

Presentation Date

7-15-2026

Abstract

The introduction of Google's TPU (Tensor Processor Unit) revolutionized AI acceleration by offering drastically improved energy efficiency over conventional GPUs. However, underlying arithmetic architectures leave room for optimization. This research investigates the performance and energy tradeoffs of utilizing a shifter-based systolic array to enable multiplier-free neural network inference. By leveraging power-of-two quantization, conventional Multiply-Accumulate (MAC) operations are replaced with more efficient shift-and-add computations. To evaluate this co-design approach, we designed and implemented a full-custom 16x16 shifter-based array supporting fixed-point and floating-point operations. Benchmarking this architecture against a custom 16x16 conventional MAC baseline yielded a 27% reduction in total power consumption, as well as a 37% decrease in silicon area. This work ultimately demonstrates that a shifter-based architecture can successfully run these optimized neural networks without sacrificing accuracy, delivering substantial power and area savings.

This document is currently not available here.

Share

COinS
 

Po2-QAT: Multiplier-Free Hardware Acceleration for Efficient AI

The introduction of Google's TPU (Tensor Processor Unit) revolutionized AI acceleration by offering drastically improved energy efficiency over conventional GPUs. However, underlying arithmetic architectures leave room for optimization. This research investigates the performance and energy tradeoffs of utilizing a shifter-based systolic array to enable multiplier-free neural network inference. By leveraging power-of-two quantization, conventional Multiply-Accumulate (MAC) operations are replaced with more efficient shift-and-add computations. To evaluate this co-design approach, we designed and implemented a full-custom 16x16 shifter-based array supporting fixed-point and floating-point operations. Benchmarking this architecture against a custom 16x16 conventional MAC baseline yielded a 27% reduction in total power consumption, as well as a 37% decrease in silicon area. This work ultimately demonstrates that a shifter-based architecture can successfully run these optimized neural networks without sacrificing accuracy, delivering substantial power and area savings.