(This page has no text content)
Engineering Generative AI-Based Software
This page intentionally left blank
Engineering Generative AI-Based Software Miroslaw Staron Computer Science and Engineering University of Gothenburg and Chalmers University of Technology Gothenburg, Sweden
Morgan Kaufmann is an imprint of Elsevier 50 Hampshire Street, 5th Floor, Cambridge, MA 02139, United States Copyright © 2026 Elsevier Inc. All rights are reserved, including those for text and data mining, AI training, and similar technologies. For accessibility purposes, images in electronic versions of this book are accompanied by alt text descriptions provided by Elsevier. For more information, see https://www.elsevier.com/about/accessibility. Books and Journals published by Elsevier comply with applicable product safety requirements. For any product safety concerns or queries, please contact our authorised representative, Elsevier B.V., at productsafety@elsevier.com. Publisher’s note: Elsevier takes a neutral position with respect to territorial disputes or jurisdictional claims in its published content, including in maps and institutional affiliations. MATLAB® is a trademark of The MathWorks, Inc. and is used with permission. The MathWorks does not warrant the accuracy of the text or exercises in this book. This book’s use or discussion of MATLAB® software or related products does not constitute endorsement or sponsorship by The MathWorks of a particular pedagogical approach or particular use of the MATLAB® software. No part of this publication may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopying, recording, or any information storage and retrieval system, without permission in writing from the publisher. Details on how to seek permission, further information about the Publisher’s permissions policies and our arrangements with organizations such as the Copyright Clearance Center and the Copyright Licensing Agency, can be found at our website: www.elsevier.com/permissions. This book and the individual contributions contained in it are protected under copyright by the Publisher (other than as may be noted herein). Notices Knowledge and best practice in this field are constantly changing. As new research and experience broaden our understanding, changes in research methods, professional practices, or medical treatment may become necessary. Practitioners and researchers must always rely on their own experience and knowledge in evaluating and using any information, methods, compounds, or experiments described herein. In using such information or methods they should be mindful of their own safety and the safety of others, including parties for whom they have a professional responsibility. To the fullest extent of the law, neither the Publisher nor the authors, contributors, or editors, assume any liability for any injury and/or damage to persons or property as a matter of products liability, negligence or otherwise, or from any use or operation of any methods, products, instructions, or ideas contained in the material herein. ISBN: 978-0-443-27606-4 For information on all Morgan Kaufmann publications visit our website at https://www.elsevier.com/books-and-journals Publisher: Mara E. Conner Acquisitions Editor: Chris Katsaropoulos Editorial Project Manager: Sonal Nagpal Production Project Manager: N. Kiruthigadevi Cover Designer: Mark Rogers Typeset by VTeX
To my family, students, and teachers. Without you, I would not be who I am.
This page intentionally left blank
Contents Biography xiii Foreword xv Preface xvii Acknowledgments xix 1. Introduction 1 1.1. Why now and why here? 1 1.2. Pre-training of transformer networks 2 1.3. Bottlenecks and latent spaces 7 1.4. Task-specific training 9 1.5. How this impacts software engineering of generative AI 11 1.6. Products that use this technology today 11 1.7. How this book is structured 12 1.8. Summary 12 2. Generative AI basics 15 2.1. Introduction 15 2.2. Instruct models 18 2.3. Training an instruct model 19 2.4. Zero-shot and one-shot prompting 23 2.5. Few-shot learning 24 2.6. Chain-of-thought learning and prompting 24 2.7. Solving a problem by creating a method 25 2.8. Securing Safe for Work content 28 2.9. Beyond models – what cannot be done in training 28 2.10. Summary 29 vii
viii Contents 3. Constructing generative AI software 31 3.1. The process 31 3.2. Typical machine learning development process 32 3.2.1. Agile software development principles 33 3.3. Phases of software development 34 3.3.1. Requirements 35 3.3.2. Software architecting and development 38 3.3.3. Software testing 41 3.3.4. Deployment and operations 44 3.4. New and classical roles in software development of generative AI software 45 3.5. Summary and where we go next 47 4. Functional and non-functional requirements for generative AI software 49 4.1. Functional and non-functional software requirements 49 4.2. Quality model for generative AI software 52 4.2.1. Data quality models 53 4.2.2. Software quality model 56 4.3. Functional requirements for generative AI systems 58 4.4. Non-functional requirements for generative AI software 59 4.5. Summary 66 5. Architecting generative AI software 67 5.1. Software architecture of generative AI software 67 5.2. Monolithic architecture 68 5.3. Model-view-controller in generative AI software 70 5.4. Microservice-based architecture 76 5.4.1. Ollama 81 5.4.2. LangChain 81 5.5. Embedded models 84 5.6. Other architectural styles 84 5.7. Non-functional properties and architectural styles 85
Contents ix 5.8. Architectural tactics for generative AI software 86 5.9. Summary 87 6. Implementation and quality assurance of generative AI software 89 6.1. Layers upon layers of frameworks 89 6.2. C and C++ for generative AI software 90 6.3. Even more language interoperability using web services 92 6.4. Portability of models 94 6.5. ONNX 95 6.6. DirectML and NPUs 97 6.7. Testing of models 100 6.8. Summary 104 7. Handling data for generative AI systems 105 7.1. Agents and why we need them 105 7.2. Agentic AI 106 7.3. Libraries to handle data for generative AI software 112 7.4. ChromaDB and embedding databases 118 7.5. ChromaDB and RAG empowered agentic AI 120 7.6. Final agents – access to tools 122 7.7. Summary 123 8. Deployment of generative AI software 125 8.1. Cloud services and datacenters 125 8.2. Patterns for deployment of generative AI software 126 8.3. Deployment process 128 8.3.1. Deployment of a web app 128 8.3.2. Deployment of a container 129 8.4. Deploying a capability – a model 131 8.5. Continuous integration and continuous deployment 134 8.6. Embedding generative AI in other products 135 8.6.1. Using ONNX Runtime to embed models 135 8.6.2. NanoLLM 138
x Contents 8.7. Performance difference between mini, medium, and large models of the same type 138 8.8. Summary 144 9. Generative AI ecosystems 147 9.1. Ecosystems 147 9.2. Services provided by generative AI software 148 9.3. Design of APIs 149 9.3.1. Different types of calls 150 9.3.2. Response codes 153 9.3.3. Error codes 154 9.3.4. Security 155 9.4. Co-development of elements of the software ecosystem 157 9.5. Coopetition: Cooperation and competition in software ecosystems 159 9.6. Summary 161 10. Summary and current trends 163 10.1. So, what we have learned so far is... 163 10.2. Interesting GitHub repositories to follow 164 10.2.1. Crew AI 164 10.2.2. Continue 164 10.2.3. Prompting 164 10.2.4. C/ua 164 10.2.5. More agent development guidelines 165 10.2.6. Playwright 165 10.2.7. Nvidia Dynamo – first AI OS 165 10.2.8. Self-hosting 165 10.2.9. Gatekeepers for prompts 165 10.3. Quo vadis generative AI? 166 10.4. What got us here will not keep us here 168 10.5. Multimodal models and linking models to other tools 169 10.5.1. Business models for generative AI 170
Contents xi 10.6. Beyond generative AI software 171 10.7. Open problems 171 10.8. A completely different challenge – regulatory and legal 172 10.8.1. Current and upcoming regulations 172 10.8.2. Intellectual property 172 10.8.3. Responsibility and liability 173 10.9. Final remarks 173 A. Additional online resources 175 References 177 Index 181
This page intentionally left blank
Biography Miroslaw Staron Miroslaw Staron is a professor of software engineering at the Department of Computer Sci- ence and Engineering, Chalmers, and the University of Gothenburg, Sweden. Prof. Staron is the Editor-in-Chief of Information and Software Technology and has a regular column in IEEE Software (Practitioner’s Digest). He is also on the editorial boards of IEEE Soft- ware and Elsevier’s Journal of Systems Architecture. Prof. Staron has been active in national bodies such as AI Sweden, AI Competence for Sweden, Swedsoft. His work has also been recognized with the title of Excellent Teacher at the University of Gothenburg. His research work focuses on software design, metrics, machine learning, and software quality. Prof. Staron authored five books and over 250 peer-reviewed articles in the area of software en- gineering. He lives in Gothenburg with his wife and three fantastic kids. His interests are computer games, reading, and long walks. He has always liked new technology and his mission is to make the technology work for the sustainable world. xiii
This page intentionally left blank
Foreword We are living in an incredibly exciting time as we are in the early days of a new general purpose technology being introduced and put in the hands of humankind: artificial in- telligence! We have been through this several times before, including the introduction of writing, agriculture, steam engines, electricity, and the Internet, just to mention a few. Ev- ery time a new general purpose technology was introduced and incorporated in society, the quality of life for humans improved remarkably. And there is no reason to assume that this time it won’t be the same! Within the field of artificial intelligence, it really has been Generative AI (GenAI) that has been stealing most of the limelight. Large Language Models (LLMs) and related models such as Visual Language Models and Multi-Modal Models have shown to provide outputs and interactions that are seemingly at human expert level or even beyond that. Of course, there are challenges such as hallucinations and these models make mistakes that seem obvious and simply stupid by our human standards, but the fact is the capabilities of these models make them extremely useful in a wide variety of fields. One of the fields that has been the most proactive and quick in adopting GenAI has been software engineering. It is very interesting and in some ways somewhat disconcerting that a field in which I have been a professor for almost 30 years suddenly is disrupted by this new technology. For most of my career, software, data, and digitalization in general affected many other fields but only created more opportunities for us software engineers. We had more work and opportunities, but could basically continue to work in the way that we always did. This is no longer the case. GenAI is helping companies generate code without or with limited human involve- ment. At the time of writing, upwards of 30% of code in systems is generated according to statements by Google, Microsoft, and others, and there is no reason to assume that this percentage will not continue to increase over time. With the fabulous amounts of invest- ment in AI, the technology will keep improving at a rapid pace for a good while longer. The challenge with new general purpose technologies is that every person has opin- ions and reflections about it, but most have very little idea of how things work under the hood. GenAI software is no different and with the multitude of tools available, the hype surrounding the field and the frenetic energy around GenAI for software, it is hard even for insiders to stay on top of things and to have a balanced, informed view of the field. The key difference separating professionals from amateurs in this is the engineering aspect. Building a prototype running in some container in the cloud is one thing. Devel- opment and evolution of reliable, safe, stable models as modules or subsystems in a larger system is an engineering challenge. One that is often underestimated. However, we need xv
xvi Foreword requirements, architecture, proper model training, quality assurance, regulatory compli- ance, deployment, monitoring, logging as well as many other aspects in place around the model to ensure a properly engineered system. This is where this book comes in. I am incredibly proud to call Miroslaw Staron, the author, my colleague, and friend for well over a decade. Through his research, he always dives in under the hood of any technology that crosses his path and GenAI is no different. Rather than relying on the opinions of others, he gets his hands dirty and starts to build systems simply for the purpose of understanding how the technology works in practice. The good, the bad, and the ugly. Warts and all. This book is no different: it takes you through the basic principles of GenAI for software, provides ample code samples and ways for you to get hands on with the technology, and helps you understand both the basics and the potential of this fabulous new technology. And he doesn’t stop there: he also addresses the engineering challenges associated with GenAI in software. Questions such as requirements engineering, software architec- ture implications, quality assurance, regulatory compliance among others, are discussed in a systematic, structured, and highly informative fashion. All in all, this book is, to use a colloquial term, the bee’s knees, in that it captures the essence of what should be top of mind for anyone in software. And one that you can’t miss! Concluding, artificial intelligence and GenAI in particular is the next general purpose technology that will have a profound, and highly positive, impact on humankind. However, for us to embrace and capitalize on the benefits provided by this technology, we need to understand it, identify its strengths and weaknesses, manage its engineering implications as well as its broader societal impact. Miroslaw has managed to capture all this in a gem of a book that I have enjoyed reading and that I will reread many times to come. I am honored to have been asked to provide a foreword to this book, but rather than reading my text, I encourage you, dear reader, to get going on the book. You won’t be dis- appointed! Jan Bosch Gothenburg, Sweden August 2025
Preface Engineering software changed when generative AI entered the mainstream. GitHub CoPi- lot and similar tools empowered software developers in a way we had never seen before. Today, I tell my students – In the future, there will be no place for bad or mediocre program- mers, only the fantastic ones who understand how to use generative AI. My students usually respond with the question – How do I become a great programmer then? This book answers the question of how to become great programmers, software devel- opers, and software engineers by learning how to design generative AI systems. Admittedly, not all software systems are based on generative AI today, but the number of generative AI-based systems will increase. Becoming great means that we need to master the new technology, we need to be able to utilize it in new products and our services. This book introduces generative AI from two perspectives: software developers’ and machine learning engineers’ perspectives. The first one introduces the typical architecture of generative AI software. The latter, on the other hand, introduces the statistical basics for generative AI software. By combining these two perspectives, we learn how generative AI learns from existing patterns and how to replicate them. Throughout the book, we gradually explore the engineering aspects of generative AI software. We dive deeper into non-functional requirements for generative software, in- cluding the concepts of NSFW (Not Safe For Work). We look deeper into the architecture of such software – both as part of monolith applications and as part of a service-oriented one. Finally, we explore frameworks that allow us to use open-source models with different programming languages and computer architectures. Miroslaw Staron Gothenburg, Sweden June 2025 xvii
This page intentionally left blank
Acknowledgments Every book project is a result of a lot of work and discussions. In my work I discussed a lot of ideas with my research group – SEAS – Software Engineering for Embedded and Auto- motive Software. The doctoral students and postdocs provided me with invaluable ideas, mostly through their smart questions. I would also like to thank my colleague Wilhelm Meding, who has encouraged me to write this book and supported me in the process. Without his questions and encourage- ment, this book would not be possible. Special thanks go to my family, Sylwia, Alexander, Viktoria, and Cornelia. They always believe in me and encourage me. Thank you for everything! A lot of thanks go to the person who enabled a lot of my research and supported me throughout the years – Anders Caspár†. I would like to thank my publisher – Chris Katsaropoulos – whose idea and feedback made this book project better already from the beginning. In writing this book, I utilized generative AI to refine grammar (Grammarly) and to condense and improve the readability of the code (GitHub CoPilot). These tools were in- valuable during the writing of this book. Thank you, Microsoft and Grammarly, for these fantastic tools! Miroslaw Staron Gothenburg, Sweden June 2025 xix
Loading comments...
Reply to Comment
Edit Comment