Share E-Book

Deep Learning with C++ Design and deploy neural networks using CUDA for high-performance AI in C++ (Bill Chen, Vikash Gupta)(Z-Library)

Author

,

C++
Language English

No Description

Format PDF
Size 6.6 MB
7
Views
(First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
(This page has no text content)
Page 2
Deep Learning with C++ Design and deploy neural networks using CUDA for high‑performance AI in C++ Bill Chen Vikash Gupta
Page 3
Deep Learning with C++ Copyright © 2026 Packt Publishing All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews. The author acknowledges the use of cutting-edge AI, such as ChatGPT, with the sole aim of enhancing the language and clarity within the book, thereby ensuring a smooth reading experience for readers. It’s important to note that the content itself has been crafted by the author and edited by a professional publishing team. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the authors, nor Packt Publishing or its dealers and distributors, will be held liable for any damages caused or alleged to have been caused directly or indirectly by this book. Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. Portfolio Director: Sunith Shetty Relationship Lead: Nilesh Kowadkar Project Manager: Yash Basil Content Engineer: Nathanya Dias Technical Editor: Sumant Jadhav Indexer: Rekha Nair Production Designer: Deepak Chavan Growth Lead: Merlyn M Shelley First published: April 2026 Production reference: 1170426 Published by Packt Publishing Ltd. Grosvenor House 11 St Paul’s Square Birmingham B3 1RB, UK. ISBN 978-1-83588-002-9 www.packtpub.com
Page 4
This book is dedicated to YOU, the reader—for your curiosity, your persistence, and your willingness to keep learning, building, and improving. – Bill Chen
Page 5
Foreword Deep learning has reshaped modern software, powering applications in vision, language, speech, recommendation, healthcare, and robotics. While Python has become the standard language for experimentation, production systems often depend on C++ for performance, efficiency, and control. When latency, memory use, and hardware optimization matter, C++ remains a critical part of the deep learning stack. That is what makes this book valuable. It gives readers a practical path from deep learning concepts to real implementation in C++. Rather than treating C++ as a secondary tool, it presents it as a serious environment for building, optimizing, and deploying modern deep learning systems. What I appreciate most is the book’s full-lifecycle approach. It begins with foundations such as environment setup, data preparation, and CUDA acceleration, then moves through core model families including neural networks, CNNs, RNNs, LSTMs, and generative models. It also covers distributed training, model compression, inference optimization, deployment, debugging, monitoring, and explainability. This makes the book useful for both C++ engineers entering deep learning and machine learning practitioners who want to move beyond Python prototypes into production-grade systems. It closes the gap between theory and deployment and shows how to build models that are not only accurate, but also efficient, reliable, and practical. I hope you enjoy the book. Bill Chen
Page 6
Contributors About the authors Bill Chen is a machine learning engineer at Meta specializing in deep learning, CUDA, and C++. He holds a PhD in Bioinformatics from the University of Kentucky and has worked in both production and instructional roles in applied AI. He has taught at the NVIDIA Deep Learning Institute, earned the NVIDIA-Certified Associate: Generative AI Multimodal credential, and served as part-time machine learning faculty at UCSC Silicon Valley Extension. His work includes Facebook group search modeling and surgical duration prediction. In this book, he combines industry experience and teaching to guide readers in building high-performance deep learning systems in C++. I would like to express my sincere gratitude to my editor, Nathanya, for the patience, flexibility, and support provided throughout the writing of this book, especially while I was serving on active duty in the US Army. Her understanding and encouragement made it possible for me to continue this project during a demanding period. I am also grateful to the reviewers for their valuable feedback and careful reading of the manuscript, and to the entire Packt team for their guidance and support throughout the publishing process. Vikash Gupta Ph.D., is a Senior Solutions Architect at Amazon Web Services (AWS), based in Seattle, Washington. He earned his Ph.D. in Computational Biology from INRIA, France, where his research centered on neuroimaging and statistical modeling. At AWS, he applies deep learning and artificial intelligence to advance medical imaging technologies, contributing to open-source initiatives such as the MONAI framework for healthcare. He also served as a research scientist at The Ohio State University and as an Assistant Professor at Mayo Clinic. He has authored more than 60 peer-reviewed publications.
Page 7
I would like to acknowledge the continued support of my family, particularly my wife, Garima Saraowgi, who has been incredibly supportive throughout the writing process. I would also like to thank my teachers, especially Prof. Xavier Pennec, who has consistently challenged and guided me throughout my career. Additionally, I want to recognize the contributions of the Packt team, specifically Nathanya Dias, for her dedicated review of this book’s content and for keeping me motivated and focused. Finally, this book is dedicated to all the students striving to learn and build new solutions—for those who come after. — Gustave, Claire Obscure
Page 8
About the reviewers Sriram M is an AI architect, consultant, and technical author with two decades of industry experience. He specializes in quantum AI, advanced machine learning, and secure model deployment. He is a co-author of Time Series Forecasting Using Generative AI and serves as a reviewer for technical books, bringing rigorous, research-driven evaluation to every manuscript. As a mentor and visiting faculty member, he teaches Python and AI to diverse learners and contributes actively to the academic community through publications, peer review, and ongoing research in quantum reservoir computing. He is also an invited speaker for industry events. Vishwas BV is a Senior Principal Data Scientist and AI Researcher with 13+ years of experience building scalable machine learning, generative AI, and agentic AI systems across manufacturing, energy, telecom, and retail. Vishwas specializes in translating complex business problems into production-grade AI architectures, including RAG, multi-agent workflows, and time series foundation models. He is a Springer-published author of two books and co-inventor of a granted U.S. patent on AI-driven emissions modeling that supported a multi-billion-dollar deal. Certified in AWS ML and Azure ML, he combines deep technical expertise with measurable business impact. 
Page 9
(This page has no text content)
Page 10
Table of Contents Preface xxv Free benefits with your book .......................................................................................... xxxi Part 1: Foundations of Deep Learning in C++ 1 Chapter 1: Introduction to Deep Learning with C++ and Environment Setup 3 Your purchase includes a free PDF copy + exclusive extras • 4 Technical requirements ...................................................................................................... 4 Introduction to deep learning concepts .............................................................................. 7 Neural networks – the building blocks of deep learning • 7 Types of learning paradigms • 9 Supervised learning • 9 Unsupervised learning • 10 Semi-supervised learning and reinforcement learning • 12 Advantages of using C++ for deep learning ....................................................................... 12 Setting up a C++ environment for deep learning ................................................................ 13 Using the free Google Colab • 13 Setting up the C++ development environment on-premises • 17 Key performance benefits of C++ in deep learning ............................................................ 22 Implementing a simple neural network in C++ • 24 Summary .......................................................................................................................... 26 Questions .......................................................................................................................... 26 Further reading ................................................................................................................. 26 Answers ............................................................................................................................ 27
Page 11
Table of Contentsx Chapter 2: Data Preparation and Preprocessing in C++ 29 Technical requirements .................................................................................................... 30 Understanding the importance of data preprocessing in DL .............................................. 31 The role of preprocessing • 32 Structured versus unstructured data – comprehensive techniques and examples ............ 32 Structured data • 33 Detailed examples • 33 Library-powered preprocessing (same techniques, less code) • 51 Unstructured data • 57 Text data • 57 Image data • 67 Audio/video data • 68 Exploring normalization and augmentation techniques .................................................. 69 Normalization strategies • 69 Augmentation techniques • 70 Integrating the PyTorch Dataset API for data pipelines ..................................................... 72 Managing large-scale datasets .......................................................................................... 73 Memory mapping • 73 Batch processing • 73 Summary .......................................................................................................................... 74 Questions .......................................................................................................................... 75 Further reading ................................................................................................................. 75 Answers ............................................................................................................................ 75 Chapter 3: CUDA for GPU Acceleration in Deep Learning with C++ 77 GPU-accelerated versus CPU-only applications ................................................................ 77 Installing and setting up CUDA ......................................................................................... 79 Installing the NVIDIA driver • 79 Downloading and installing the CUDA Toolkit • 80 Installing additional dependencies • 82 Using cloud-based CUDA environments • 82
Page 12
Table of Contents xi Optimizing C++ code with CUDA ...................................................................................... 86 CUDA programming model • 86 CPU baseline: array addition in C++ • 87 Converting to CUDA code • 89 Launching the CUDA kernel • 89 Full CUDA implementation • 90 Profiling CUDA performance • 92 CUDA thread hierarchy • 92 Speeding up with multiple threads • 93 Scaling up with multiple blocks • 96 How to debug in CUDA ................................................................................................... 100 Checking CUDA API calls • 100 Checking kernel launch errors • 101 Checking asynchronous execution errors • 101 CUDA error-handling function • 102 Summary ........................................................................................................................ 103 Questions ........................................................................................................................ 104 Further reading ............................................................................................................... 104 Answers .......................................................................................................................... 105 Part 2: Building and Training Neural Networks in C++ 107 Chapter 4: Building a Basic Neural Network in C++ 109 Understanding the basics ................................................................................................. 110 Building the forward pass and losses (e.g., BCE/cross-entropy for classification) ............. 111 Implementing neurons and activations in C++ with LibTorch ......................................... 116 Deriving/backpropagating gradients and updating weights with Stochastic Gradient Descent ............................................................................................................................ 121 Summary ......................................................................................................................... 125 Questions ......................................................................................................................... 125 Further reading ................................................................................................................ 125 Answers ........................................................................................................................... 126
Page 13
Table of Contentsxii Chapter 5: Multilayer Perceptron’s in C++ 127 Technical requirements .................................................................................................. 128 Understanding MLP architecture ..................................................................................... 129 Implementing MLP in C++ • 130 Layer class • 131 MultilayerPerceptron class • 132 Implementing MLP using CUDA • 135 Implementing MLP using LibTorch • 139 Choosing your MLP implementation approach: A strategic guide .................................... 141 Raw C++ with Eigen • 141 CUDA: maximum performance at significant cost • 141 LibTorch: practical default for production systems • 142 The hybrid approach • 142 Exploring deep learning training concepts ..................................................................... 144 Activation functions • 144 Sigmoid family • 145 ReLU family • 146 Other advanced functions • 150 Optimization techniques • 152 Gradient descent • 153 RMS Prop • 155 Momentum • 156 ADAM optimizer • 158 AdaGrad • 160 AdaDelta • 162 Summary ......................................................................................................................... 165 Further reading ............................................................................................................... 166
Page 14
Table of Contents xiii Chapter 6: Convolutional Neural Networks in C++ 169 History of CNNs .............................................................................................................. 170 Mathematical foundations of convolution ....................................................................... 171 Implementation in C++ .................................................................................................... 173 Forward pass implementation • 174 Integration with larger networks • 175 Matrix-based implementation with Eigen • 176 GPUS-accelerated implementation with CUDA • 178 ConvolutionLayer using PyTorch C++ API • 182 Building a C++-based neural network for image classification ....................................... 183 Image segmentation with U-Net ..................................................................................... 186 Image processing for deep learning ................................................................................. 191 Rotation • 191 Translation • 192 Cropping • 193 Scaling • 194 Zooming • 196 Flipping • 197 Padding • 199 Resampling • 200 Terminologies associated with CNN ............................................................................... 203 Pooling • 204 Stride • 205 Feature maps • 205 Summary ........................................................................................................................ 206 Further reading ............................................................................................................... 207
Page 15
Table of Contentsxiv Chapter 7: Recurrent Neural Networks and Long Short-Term Memory Networks in C++ 209 Understanding time series and sequential data .............................................................. 210 RNN architectures and mathematics ............................................................................... 211 Types of RNNs • 212 Mathematical foundations • 212 Memory dynamics and information flow • 214 C++ implementation for RNNs • 214 Eigen-based approach • 214 BPTT ................................................................................................................................ 216 Mathematical foundation for BPTT • 216 Gradient flow through time • 216 Parameter gradient accumulation • 217 Practical implementation considerations • 217 C++ implementation for BPTT • 218 Vector-based implementation • 219 Eigen-based implementation • 220 Libtorch-based implementation • 220 LSTM networks ................................................................................................................ 221 The cell state: LSTM’s memory highway • 222 LSTM mathematical formulation • 223 Information flow dynamics • 224 C++ implementation • 225 Vector-based implementation • 225 Matrix-based implementation • 227 Libtorch-based implementation • 231 LSTM backpropagation ................................................................................................... 234 Mathematical framework for LSTM backpropagation • 235 Gradient flow through gates • 236 Parameter gradient accumulation • 236 Practical implementation considerations • 236
Page 16
Table of Contents xv Text processing ............................................................................................................... 237 Enhanced string operations in modern C++ • 237 File I/O operations for text data • 238 Basic string manipulation and processing • 240 Language-specific considerations • 241 Convert to lowercase • 241 Remove punctuation and special characters • 242 Split sentences into words • 243 Stop words in sequential text processing • 246 Word embeddings: Bridging discrete tokens and continuous representations ................ 248 Word2Vec: Learning semantic representations • 249 Mathematical formulation • 249 C++ implementation with LibTorch • 250 Integration with LSTM networks • 253 Text processing pipeline for deep learning ...................................................................... 254 Applications of RNN on language tasks • 255 Building a sequence-to-sequence translator with LibTorch • 259 Vocabulary construction and token management • 260 The encoder-decoder forward pass • 261 Training with teacher forcing • 261 Inference through autoregressive decoding • 262 Summary ........................................................................................................................ 263 Further reading ............................................................................................................... 263 Chapter 8: Generative Networks, Autoencoders, and Large Language Models in C++ 265 Introduction to generative models .................................................................................. 266 Autoencoders: foundation of generative learning ........................................................... 267 Basic autoencoder architecture • 268 Loss functions and reconstruction error • 268 C++ implementation using LibTorch • 268
Page 17
Table of Contentsxvi Forward pass implementation • 270 Training loop • 270 VAEs ................................................................................................................................. 271 The reparameterization trick • 272 C++ implementation of VAE • 274 GANs – fundamentals ...................................................................................................... 279 GAN loss functions • 281 Discriminator loss • 281 Generator loss • 281 GAN training process • 282 Alternative loss functions • 283 C++ implementation • 283 Generator network architecture • 284 Discriminator network architecture • 285 Training a GAN • 285 Autoregressive generation and causal language modeling .............................................. 288 Temperature and sampling strategies • 290 Temperature scaling • 291 Sampling strategy • 292 Greedy decoding • 292 Beam search • 293 Top-K sampling • 295 TF-IDF • 300 N-Grams • 300 Evaluation metrics for generative models ....................................................................... 301 Categories of evaluation metrics • 302 Common statistical metrics • 302 Model-based metrics • 305 BERTScore • 305 Summary ........................................................................................................................ 306 Further reading ............................................................................................................... 307
Page 18
Table of Contents xvii Chapter 9: Transformers and Large Language Model Fine-Tuning in C++ 309 Introduction to transformer architecture ....................................................................... 310 The attention mechanism revolution • 311 Self-Attention mechanism • 312 Mathematical foundations of self-attention • 313 C++ implementation of self-attention • 314 Multi-Head attention • 316 C++ implementation of multi-head attention • 317 Positional encoding ......................................................................................................... 319 Learned position embeddings • 319 Sinusoidal position embedding • 320 RoPE • 324 Mathematical mechanism: pairing even and odd dimensions • 324 Why the dot product depends on relative position • 325 Alternative modern approaches • 327 Implementation considerations • 327 Complete transformer implementation .......................................................................... 328 Encoder stack implementation • 329 Processing flow • 329 Self-Attention • 329 Add and normalize • 330 Feed-Forward network • 330 Stacking the layers • 330 Decoder stack implementation • 330 Processing flow • 331 Multi-head cross-attention • 332 Feed-Forward network and layer stacking • 333 Putting it all together • 335 Training versus inference • 336
Page 19
Table of Contentsxviii Pre-trained language models .......................................................................................... 337 BERT: bidirectional context understanding • 337 GPT: autoregressive text generation • 338 Architectural differences between BERT and GPT • 339 Model parallelization ...................................................................................................... 339 DDP • 339 Understanding DDP implementation: code walkthrough • 340 DDPTrainer class • 341 Fully Sharded Data Parallel • 344 Understanding FSDP: code walkthrough • 344 Model compression techniques ........................................................................................ 351 Knowledge distillation • 352 The training process: dual objectives • 353 Pruning • 355 Pipeline overview • 357 Quantization • 359 Range selection strategies • 360 Post-Training quantization (PTQ) • 361 Quantization-Aware training (QAT) • 361 C++ implementation of quantization for neural networks • 362 Summary ........................................................................................................................ 367 Further reading ............................................................................................................... 368 Part 3: Deploying, Monitoring, and Explaining Deep Learning Systems in Production 371 Chapter 10: Deploying and Optimizing Models for Inference 373 Exporting models for C++ inference ................................................................................ 375 A compact model in the C++ Frontend • 376 Producing a TorchScript artifact (from C++) ................................................................... 378 Loading and running TorchScript in C++ • 378
Page 20
Table of Contents xix Parity: does the artifact match the native graph? ............................................................ 379 Consuming ONNX from C++ for portability • 381 What to keep consistent • 382 Deploying models to real targets ..................................................................................... 382 Packaging the binary (the one you ship) • 383 The heart of serving: a Micro-Batcher with Concurrency Control • 383 Concurrency controls (keeping p95/p99 honest) • 390 Cloud, On-Prem, Edge: what actually changes? • 390 Optimizing models for speed and cost ............................................................................. 391 Quantization: fewer bits, more throughput • 392 FP16 mixed precision (TorchScript, CUDA) • 392 INT8 (ONNX Runtime, Q/DQ graphs) • 393 Pruning: remove structure, not just values • 394 Distillation: deploy the student • 395 Runtime tuning: feeding the hardware efficiently • 397 Layout, threads, algorithms (LibTorch) • 397 Runtime tuning with ONNX Runtime • 398 Tie it together: a simple, repeatable benchmark • 399 Practical playbook • 399 Serving models in real time ............................................................................................ 400 Contract at the Boundary • 401 The Scheduling Layer: Micro-Batching with Concurrency Control • 402 HTTP and gRPC as the Service Edge • 405 HTTP (REST and Server-Sent Events) • 405 gRPC (Unary and Server-Streaming) • 405 Tail-Latency Discipline • 405 Cloud, On-Prem, and Edge • 406 A Minimal End-to-End Shape • 406 Operating and evolving production systems ................................................................... 407 Make Behavior Visible • 407 Detect Drift Before Users Do • 409
The above is a preview of the first 20 pages. Register to read the complete e-book.

Recommended for You

Loading recommended books...
Failed to load, please try again later

Tip the Site

Scan the WeChat Pay or Alipay code to tip. No login required.

WeChat Pay
Alipay
Back to List