Dynamic DRAM-Based Processing-In-Memory Architecture with Neuron Processing Elements

Invention Description
Machine learning (ML) is fast becoming a dominant paradigm of computing in almost every domain. ML algorithms are realized as parametric function graphs (i.e., deep neural networks–DNN) in which nodes represent the composition of inner products and non-linear functions, and connections represent function composition. The major internet companies like Google, Meta, Microsoft, and others, deploy DNNs with hundreds of billions of parameters, performing trillions of large dimensional matrix operations. Thus, DNNs are often both memory and compute-bound. Consequently, they require massive amounts of memory and large server farms with thousands of high-performance GPU processors. The electricity usage of such server farms is approaching that of whole industries and some nation states, and for such systems to be sustainable, at least one to two orders of magnitude improvements in energy efficiency are required.
 
Researchers at Arizona State University have developed a technology that introduces an innovative processing-in-memory (PIM) architecture that integrates neuron processing elements (NPEs) directly into DRAM to reduce data transfer overhead and increase parallelism. The configurable neurons perform arithmetic, logical, and predicate operations with support for multiple data formats and dynamic reconfiguration, all without incurring additional latency or power costs. By maximizing parallelism and minimizing area and power overhead, it enables scalable near-memory computing. Variable precision computing capability adapts to different neural network inference requirements, delivering substantial improvements in throughput and energy efficiency. Validated through CNN inference workloads, this architecture delivers significant improvements compared to existing solutions.
 
This novel DRAM-based processing-in-memory architecture leverages NPEs to significantly enhance energy efficiency and throughput for AI and machine learning workloads.
 
Potential Applications
  • AI and machine learning inference accelerators
  • Energy-efficient data centers focusing on deep learning workloads
  • Edge computing devices requiring low-power, high-throughput AI processing
  • Advanced memory solutions for high-performance computing systems
  • Hardware platforms for neural network training and inference optimization
Benefits and Advantages
  • Significantly higher throughput–up to 2.69× improvement compared to current solutions
  • Superior energy efficiency–up to 11.75× gains, reducing power consumption dramatically
  • Minimizes data transfer between memory and processor
  • Flexible support for multiple data formats and variable bit-widths tailored for AI workloads.
  • Dynamic reconfiguration without added latency or power costs.
  • Compact integration inside DRAM minimizes area footprint and maximizes parallelism
  • Dynamic reconfiguration of neuron processing elements without performance penalties
For more information about this opportunity, please see
Patent Information: