{"product_id":"artificial-intelligence-hardware-design-challenges-and-solutions-hardback-9781119810452","title":"Artificial Intelligence Hardware Design; Challenges and Solutions (Hardback) 9781119810452","description":"\u003cfont face=\"Georgia\"\u003e\r\n\u003cp\u003e\u003cfont size=\"6\"\u003eArtificial Intelligence Hardware Design\u003c\/font\u003e\u003cbr\u003e\r\n\u003cfont size=\"5\"\u003eChallenges and Solutions\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\r\n\r\n\r\n\u003cp\u003e\u003cfont size=\"4\"\u003eAlbert Chun-Chen Liu (Author), Oscar Ming Kin Law (Author)\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003e9781119810452, Wiley\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003eHardback, published 7 September 2021\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003e240 pages\u003cbr\u003e1 x 1 x 1 cm, 0.454 kg\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\r\n\r\n\r\n\r\n\u003cp align=\"justify\"\u003e\u003cstrong\u003e\u003cfont size=\"3\"\u003e\u003cb\u003eARTIFICIAL INTELLIGENCE HARDWARE DESIGN\u003c\/b\u003e \u003cp\u003e\u003cb\u003eLearn foundational and advanced topics in Neural Processing Unit design with real-world examples from leading voices in the field\u003c\/b\u003e \u003c\/p\u003e\n\u003cp\u003eIn \u003ci\u003eArtificial Intelligence Hardware Design: Challenges and Solutions\u003c\/i\u003e, distinguished researchers and authors Drs. Albert Chun Chen Liu and Oscar Ming Kin Law deliver a rigorous and practical treatment of the design applications of specific circuits and systems for accelerating neural network processing. Beginning with a discussion and explanation of neural networks and their developmental history, the book goes on to describe parallel architectures, streaming graphs for massive parallel computation, and convolution optimization. \u003c\/p\u003e\n\u003cp\u003eThe authors offer readers an illustration of in-memory computation through Georgia Tech’s Neurocube and Stanford’s Tetris accelerator using the Hybrid Memory Cube, as well as near-memory architecture through the embedded eDRAM of the Institute of Computing Technology, the Chinese Academy of Science, and other institutions. \u003c\/p\u003e\n\u003cp\u003eReaders will also find a discussion of 3D neural processing techniques to support multiple layer neural networks, as well as information like: \u003c\/p\u003e\n\u003cul\u003e\n\u003cli\u003eA thorough introduction to neural networks and neural network development history, as well as Convolutional Neural Network (CNN) models\u003c\/li\u003e \u003cli\u003eExplorations of various parallel architectures, including the Intel CPU, Nvidia GPU, Google TPU, and Microsoft NPU, emphasizing hardware and software integration for performance improvement\u003c\/li\u003e \u003cli\u003eDiscussions of streaming graph for massive parallel computation with the Blaize GSP and Graphcore IPU\u003c\/li\u003e \u003cli\u003eAn examination of how to optimize convolution with UCLA Deep Convolutional Neural Network accelerator filter decomposition\u003c\/li\u003e\n\u003c\/ul\u003e \u003cp\u003ePerfect for hardware and software engineers and firmware developers, \u003ci\u003eArtificial Intelligence Hardware Design\u003c\/i\u003e is an indispensable resource for anyone working with Neural Processing Units in either a hardware or software capacity.\u003c\/p\u003e\u003c\/font\u003e\u003c\/strong\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003e\u003cp\u003eAuthor Biographies xi\u003c\/p\u003e \u003cp\u003ePreface xiii\u003c\/p\u003e \u003cp\u003eAcknowledgments xv\u003c\/p\u003e \u003cp\u003eTable of Figures xvii\u003c\/p\u003e \u003cp\u003e\u003cb\u003e1 Introduction \u003c\/b\u003e\u003cb\u003e1\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e1.1 Development History 2\u003c\/p\u003e \u003cp\u003e1.2 Neural Network Models 4\u003c\/p\u003e \u003cp\u003e1.3 Neural Network Classification 4\u003c\/p\u003e \u003cp\u003e1.3.1 Supervised Learning 4\u003c\/p\u003e \u003cp\u003e1.3.2 Semi-supervised Learning 5\u003c\/p\u003e \u003cp\u003e1.3.3 Unsupervised Learning 6\u003c\/p\u003e \u003cp\u003e1.4 Neural Network Framework 6\u003c\/p\u003e \u003cp\u003e1.5 Neural Network Comparison 10\u003c\/p\u003e \u003cp\u003eExercise 11\u003c\/p\u003e \u003cp\u003eReferences 12\u003c\/p\u003e \u003cp\u003e\u003cb\u003e2 Deep Learning \u003c\/b\u003e\u003cb\u003e13\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e2.1 Neural Network Layer 13\u003c\/p\u003e \u003cp\u003e2.1.1 Convolutional Layer 13\u003c\/p\u003e \u003cp\u003e2.1.2 Activation Layer 17\u003c\/p\u003e \u003cp\u003e2.1.3 Pooling Layer 18\u003c\/p\u003e \u003cp\u003e2.1.4 Normalization Layer 19\u003c\/p\u003e \u003cp\u003e2.1.5 Dropout Layer 20\u003c\/p\u003e \u003cp\u003e2.1.6 Fully Connected Layer 20\u003c\/p\u003e \u003cp\u003e2.2 Deep Learning Challenges 22\u003c\/p\u003e \u003cp\u003eExercise 22\u003c\/p\u003e \u003cp\u003eReferences 24\u003c\/p\u003e \u003cp\u003e\u003cb\u003e3 Parallel Architecture \u003c\/b\u003e\u003cb\u003e25\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e3.1 Intel Central Processing Unit (CPU) 25\u003c\/p\u003e \u003cp\u003e3.1.1 Skylake Mesh Architecture 27\u003c\/p\u003e \u003cp\u003e3.1.2 Intel Ultra Path Interconnect (UPI) 28\u003c\/p\u003e \u003cp\u003e3.1.3 Sub Non-unified Memory Access Clustering (SNC) 29\u003c\/p\u003e \u003cp\u003e3.1.4 Cache Hierarchy Changes 31\u003c\/p\u003e \u003cp\u003e3.1.5 Single\/Multiple Socket Parallel Processing 32\u003c\/p\u003e \u003cp\u003e3.1.6 Advanced Vector Software Extension 33\u003c\/p\u003e \u003cp\u003e3.1.7 Math Kernel Library for Deep Neural Network (MKL-DNN) 34\u003c\/p\u003e \u003cp\u003e3.2 NVIDIA Graphics Processing Unit (GPU) 39\u003c\/p\u003e \u003cp\u003e3.2.1 Tensor Core Architecture 41\u003c\/p\u003e \u003cp\u003e3.2.2 Winograd Transform 44\u003c\/p\u003e \u003cp\u003e3.2.3 Simultaneous Multithreading (SMT) 45\u003c\/p\u003e \u003cp\u003e3.2.4 High Bandwidth Memory (HBM2) 46\u003c\/p\u003e \u003cp\u003e3.2.5 NVLink2 Configuration 47\u003c\/p\u003e \u003cp\u003e3.3 NVIDIA Deep Learning Accelerator (NVDLA) 49\u003c\/p\u003e \u003cp\u003e3.3.1 Convolution Operation 50\u003c\/p\u003e \u003cp\u003e3.3.2 Single Data Point Operation 50\u003c\/p\u003e \u003cp\u003e3.3.3 Planar Data Operation 50\u003c\/p\u003e \u003cp\u003e3.3.4 Multiplane Operation 50\u003c\/p\u003e \u003cp\u003e3.3.5 Data Memory and Reshape Operations 51\u003c\/p\u003e \u003cp\u003e3.3.6 System Configuration 51\u003c\/p\u003e \u003cp\u003e3.3.7 External Interface 52\u003c\/p\u003e \u003cp\u003e3.3.8 Software Design 52\u003c\/p\u003e \u003cp\u003e3.4 Google Tensor Processing Unit (TPU) 53\u003c\/p\u003e \u003cp\u003e3.4.1 System Architecture 53\u003c\/p\u003e \u003cp\u003e3.4.2 Multiply–Accumulate (MAC) Systolic Array 55\u003c\/p\u003e \u003cp\u003e3.4.3 New Brain Floating-Point Format 55\u003c\/p\u003e \u003cp\u003e3.4.4 Performance Comparison 57\u003c\/p\u003e \u003cp\u003e3.4.5 Cloud TPU Configuration 58\u003c\/p\u003e \u003cp\u003e3.4.6 Cloud Software Architecture 60\u003c\/p\u003e \u003cp\u003e3.5 Microsoft Catapult Fabric Accelerator 61\u003c\/p\u003e \u003cp\u003e3.5.1 System Configuration 64\u003c\/p\u003e \u003cp\u003e3.5.2 Catapult Fabric Architecture 65\u003c\/p\u003e \u003cp\u003e3.5.3 Matrix-Vector Multiplier 65\u003c\/p\u003e \u003cp\u003e3.5.4 Hierarchical Decode and Dispatch (HDD) 67\u003c\/p\u003e \u003cp\u003e3.5.5 Sparse Matrix-Vector Multiplication 68\u003c\/p\u003e \u003cp\u003eExercise 70\u003c\/p\u003e \u003cp\u003eReferences 71\u003c\/p\u003e \u003cp\u003e\u003cb\u003e4 Streaming Graph Theory \u003c\/b\u003e\u003cb\u003e73\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e4.1 Blaize Graph Streaming Processor 73\u003c\/p\u003e \u003cp\u003e4.1.1 Stream Graph Model 73\u003c\/p\u003e \u003cp\u003e4.1.2 Depth First Scheduling Approach 75\u003c\/p\u003e \u003cp\u003e4.1.3 Graph Streaming Processor Architecture 76\u003c\/p\u003e \u003cp\u003e4.2 Graphcore Intelligence Processing Unit 79\u003c\/p\u003e \u003cp\u003e4.2.1 Intelligence Processor Unit Architecture 79\u003c\/p\u003e \u003cp\u003e4.2.2 Accumulating Matrix Product (AMP) Unit 79\u003c\/p\u003e \u003cp\u003e4.2.3 Memory Architecture 79\u003c\/p\u003e \u003cp\u003e4.2.4 Interconnect Architecture 79\u003c\/p\u003e \u003cp\u003e4.2.5 Bulk Synchronous Parallel Model 81\u003c\/p\u003e \u003cp\u003eExercise 83\u003c\/p\u003e \u003cp\u003eReferences 84\u003c\/p\u003e \u003cp\u003e\u003cb\u003e5 Convolution Optimization \u003c\/b\u003e\u003cb\u003e85\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e5.1 Deep Convolutional Neural Network Accelerator 85\u003c\/p\u003e \u003cp\u003e5.1.1 System Architecture 86\u003c\/p\u003e \u003cp\u003e5.1.2 Filter Decomposition 87\u003c\/p\u003e \u003cp\u003e5.1.3 Streaming Architecture 90\u003c\/p\u003e \u003cp\u003e5.1.3.1 Filter Weights Reuse 90\u003c\/p\u003e \u003cp\u003e5.1.3.2 Input Channel Reuse 92\u003c\/p\u003e \u003cp\u003e5.1.4 Pooling 92\u003c\/p\u003e \u003cp\u003e5.1.4.1 Average Pooling 92\u003c\/p\u003e \u003cp\u003e5.1.4.2 Max Pooling 93\u003c\/p\u003e \u003cp\u003e5.1.5 Convolution Unit (CU) Engine 94\u003c\/p\u003e \u003cp\u003e5.1.6 Accumulation (ACCU) Buffer 94\u003c\/p\u003e \u003cp\u003e5.1.7 Model Compression 95\u003c\/p\u003e \u003cp\u003e5.1.8 System Performance 95\u003c\/p\u003e \u003cp\u003e5.2 Eyeriss Accelerator 97\u003c\/p\u003e \u003cp\u003e5.2.1 Eyeriss System Architecture 97\u003c\/p\u003e \u003cp\u003e5.2.2 2D Convolution to 1D Multiplication 98\u003c\/p\u003e \u003cp\u003e5.2.3 Stationary Dataflow 99\u003c\/p\u003e \u003cp\u003e5.2.3.1 Output Stationary 99\u003c\/p\u003e \u003cp\u003e5.2.3.2 Weight Stationary 101\u003c\/p\u003e \u003cp\u003e5.2.3.3 Input Stationary 101\u003c\/p\u003e \u003cp\u003e5.2.4 Row Stationary (RS) Dataflow 104\u003c\/p\u003e \u003cp\u003e5.2.4.1 Filter Reuse 104\u003c\/p\u003e \u003cp\u003e5.2.4.2 Input Feature Maps Reuse 106\u003c\/p\u003e \u003cp\u003e5.2.4.3 Partial Sums Reuse 106\u003c\/p\u003e \u003cp\u003e5.2.5 Run-Length Compression (RLC) 106\u003c\/p\u003e \u003cp\u003e5.2.6 Global Buffer 108\u003c\/p\u003e \u003cp\u003e5.2.7 Processing Element Architecture 108\u003c\/p\u003e \u003cp\u003e5.2.8 Network-on- Chip (NoC) 108\u003c\/p\u003e \u003cp\u003e5.2.9 Eyeriss v2 System Architecture 112\u003c\/p\u003e \u003cp\u003e5.2.10 Hierarchical Mesh Network 116\u003c\/p\u003e \u003cp\u003e5.2.10.1 Input Activation HM-NoC 118\u003c\/p\u003e \u003cp\u003e5.2.10.2 Filter Weight HM-NoC 118\u003c\/p\u003e \u003cp\u003e5.2.10.3 Partial Sum HM-NoC 119\u003c\/p\u003e \u003cp\u003e5.2.11 Compressed Sparse Column Format 120\u003c\/p\u003e \u003cp\u003e5.2.12 Row Stationary Plus (RS+) Dataflow 122\u003c\/p\u003e \u003cp\u003e5.2.13 System Performance 123\u003c\/p\u003e \u003cp\u003eExercise 125\u003c\/p\u003e \u003cp\u003eReferences 125\u003c\/p\u003e \u003cp\u003e\u003cb\u003e6 In-Memory Computation \u003c\/b\u003e\u003cb\u003e127\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e6.1 Neurocube Architecture 127\u003c\/p\u003e \u003cp\u003e6.1.1 Hybrid Memory Cube (HMC) 127\u003c\/p\u003e \u003cp\u003e6.1.2 Memory Centric Neural Computing (MCNC) 130\u003c\/p\u003e \u003cp\u003e6.1.3 Programmable Neurosequence Generator (PNG) 131\u003c\/p\u003e \u003cp\u003e6.1.4 System Performance 132\u003c\/p\u003e \u003cp\u003e6.2 Tetris Accelerator 133\u003c\/p\u003e \u003cp\u003e6.2.1 Memory Hierarchy 133\u003c\/p\u003e \u003cp\u003e6.2.2 In-Memory Accumulation 133\u003c\/p\u003e \u003cp\u003e6.2.3 Data Scheduling 135\u003c\/p\u003e \u003cp\u003e6.2.4 Neural Network Vaults Partition 136\u003c\/p\u003e \u003cp\u003e6.2.5 System Performance 137\u003c\/p\u003e \u003cp\u003e6.3 NeuroStream Accelerator 138\u003c\/p\u003e \u003cp\u003e6.3.1 System Architecture 138\u003c\/p\u003e \u003cp\u003e6.3.2 NeuroStream Coprocessor 140\u003c\/p\u003e \u003cp\u003e6.3.3 4D Tiling Mechanism 140\u003c\/p\u003e \u003cp\u003e6.3.4 System Performance 141\u003c\/p\u003e \u003cp\u003eExercise 143\u003c\/p\u003e \u003cp\u003eReferences 143\u003c\/p\u003e \u003cp\u003e\u003cb\u003e7 Near-Memory Architecture \u003c\/b\u003e\u003cb\u003e145\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e7.1 DaDianNao Supercomputer 145\u003c\/p\u003e \u003cp\u003e7.1.1 Memory Configuration 145\u003c\/p\u003e \u003cp\u003e7.1.2 Neural Functional Unit (NFU) 146\u003c\/p\u003e \u003cp\u003e7.1.3 System Performance 149\u003c\/p\u003e \u003cp\u003e7.2 Cnvlutin Accelerator 150\u003c\/p\u003e \u003cp\u003e7.2.1 Basic Operation 151\u003c\/p\u003e \u003cp\u003e7.2.2 System Architecture 151\u003c\/p\u003e \u003cp\u003e7.2.3 Processing Order 154\u003c\/p\u003e \u003cp\u003e7.2.4 Zero-Free Neuron Array Format (ZFNAf) 155\u003c\/p\u003e \u003cp\u003e7.2.5 The Dispatcher 155\u003c\/p\u003e \u003cp\u003e7.2.6 Network Pruning 157\u003c\/p\u003e \u003cp\u003e7.2.7 System Performance 157\u003c\/p\u003e \u003cp\u003e7.2.8 Raw or Encoded Format (RoE) 158\u003c\/p\u003e \u003cp\u003e7.2.9 Vector Ineffectual Activation Identifier Format (VIAI) 159\u003c\/p\u003e \u003cp\u003e7.2.10 Ineffectual Activation Skipping 159\u003c\/p\u003e \u003cp\u003e7.2.11 Ineffectual Weight Skipping 161\u003c\/p\u003e \u003cp\u003eExercise 161\u003c\/p\u003e \u003cp\u003eReferences 161\u003c\/p\u003e \u003cp\u003e\u003cb\u003e8 Network Sparsity \u003c\/b\u003e\u003cb\u003e163\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e8.1 Energy Efficient Inference Engine (EIE) 163\u003c\/p\u003e \u003cp\u003e8.1.1 Leading Nonzero Detection (LNZD) Network 163\u003c\/p\u003e \u003cp\u003e8.1.2 Central Control Unit (CCU) 164\u003c\/p\u003e \u003cp\u003e8.1.3 Processing Element (PE) 164\u003c\/p\u003e \u003cp\u003e8.1.4 Deep Compression 166\u003c\/p\u003e \u003cp\u003e8.1.5 Sparse Matrix Computation 167\u003c\/p\u003e \u003cp\u003e8.1.6 System Performance 169\u003c\/p\u003e \u003cp\u003e8.2 Cambricon-X Accelerator 169\u003c\/p\u003e \u003cp\u003e8.2.1 Computation Unit 171\u003c\/p\u003e \u003cp\u003e8.2.2 Buffer Controller 171\u003c\/p\u003e \u003cp\u003e8.2.3 System Performance 174\u003c\/p\u003e \u003cp\u003e8.3 SCNN Accelerator 175\u003c\/p\u003e \u003cp\u003e8.3.1 SCNN PT-IS-CP-Dense Dataflow 175\u003c\/p\u003e \u003cp\u003e8.3.2 SCNN PT-IS-CP-Sparse Dataflow 177\u003c\/p\u003e \u003cp\u003e8.3.3 SCNN Tiled Architecture 178\u003c\/p\u003e \u003cp\u003e8.3.4 Processing Element Architecture 179\u003c\/p\u003e \u003cp\u003e8.3.5 Data Compression 180\u003c\/p\u003e \u003cp\u003e8.3.6 System Performance 180\u003c\/p\u003e \u003cp\u003e8.4 SeerNet Accelerator 183\u003c\/p\u003e \u003cp\u003e8.4.1 Low-Bit Quantization 183\u003c\/p\u003e \u003cp\u003e8.4.2 Efficient Quantization 184\u003c\/p\u003e \u003cp\u003e8.4.3 Quantized Convolution 185\u003c\/p\u003e \u003cp\u003e8.4.4 Inference Acceleration 186\u003c\/p\u003e \u003cp\u003e8.4.5 Sparsity-Mask Encoding 186\u003c\/p\u003e \u003cp\u003e8.4.6 System Performance 188\u003c\/p\u003e \u003cp\u003eExercise 188\u003c\/p\u003e \u003cp\u003eReferences 188\u003c\/p\u003e \u003cp\u003e\u003cb\u003e9 3D Neural Processing \u003c\/b\u003e\u003cb\u003e191\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e9.1 3D Integrated Circuit Architecture 191\u003c\/p\u003e \u003cp\u003e9.2 Power Distribution Network 193\u003c\/p\u003e \u003cp\u003e9.3 3D Network Bridge 195\u003c\/p\u003e \u003cp\u003e9.3.1 3D Network-on-Chip 195\u003c\/p\u003e \u003cp\u003e9.3.2 Multiple-Channel High-Speed Link 195\u003c\/p\u003e \u003cp\u003e9.4 Power-Saving Techniques 198\u003c\/p\u003e \u003cp\u003e9.4.1 Power Gating 198\u003c\/p\u003e \u003cp\u003e9.4.2 Clock Gating 199\u003c\/p\u003e \u003cp\u003eExercise 200\u003c\/p\u003e \u003cp\u003eReferences 201\u003c\/p\u003e \u003cp\u003eAppendix A: Neural Network Topology 203\u003c\/p\u003e \u003cp\u003eIndex 205\u003c\/p\u003e\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003eSubject Areas: Computer science [\u003ca title=\"See our other books on Computer science\" href=\"https:\/\/freshlyprintedbooks.co.uk\/search?q=%22Computer%20science%20%5BUY%5D%22\"\u003eUY\u003c\/a\u003e]\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\r\n\u003c\/font\u003e","brand":"Wiley-IEEE Press","offers":[{"title":"Brand New","offer_id":52501146960152,"sku":"9781119810452","price":74.99,"currency_code":"GBP","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0730\/2037\/5320\/files\/9781119810452.jpg?v=1786212641","url":"https:\/\/freshlyprintedbooks.co.uk\/products\/artificial-intelligence-hardware-design-challenges-and-solutions-hardback-9781119810452","provider":"Freshly Printed Books","version":"1.0","type":"link"}