Share E-Book

AuthorBartosz Konieczny

Data projects are an intrinsic part of an organization’s technical ecosystem, but data engineers in many companies continue to work on problems that others have already solved. This hands-on guide shows you how to provide valuable data by focusing on various aspects of data engineering, including data ingestion, data quality, idempotency, and more. Author Bartosz Konieczny guides you through the process of building reliable end-to-end data engineering projects, from data ingestion to data observability, focusing on data engineering design patterns that solve common business problems in a secure and storage-optimized manner. Each pattern includes a user-facing description of the problem, solutions, and consequences that place the pattern into the context of real-life scenarios. Throughout this journey, you’ll use open source data tools and public cloud services to apply each pattern. You'll learn: Challenges data engineers face and their impact on data systems How these challenges relate to data system components Useful applications of data engineering patterns How to identify and fix issues with your current data components TTechnology-agnostic solutions to new and existing data projects, with open source implementation examples Bartosz Konieczny is a freelance data engineer who's been coding since 2010. He's held various senior hands-on positions that allowed him to work on many data engineering problems in batch and stream processing.

AI Reading Assistant

Summary and highlights from this book's index; jump to passages in the text

Passage locations
Tags
No tags
ISBN: 1098165780
Publish Year: 2024
Language: 英文
Pages: 393
File Format: PDF
File Size: 7.1 MB
Support Statistics
¥.00 · 0times
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Bartosz Konieczny Data Engineering Design Patterns Recipes for Solving the Most Common Data Engineering Problems
ISBN: 978-1-098-16581-9 US $79.99 CAN $99.99 DATA Data projects are an intrinsic part of an organization’s technical ecosystem, but data engineers in many companies continue to work on problems that others have already solved. This hands-on guide shows you how to provide valuable data by focusing on various aspects of data engineering, including data ingestion, data quality, idempotency, and more. Author Bartosz Konieczny guides you through the process of building reliable end-to-end data engineering projects, from data ingestion to data observability, focusing on data engineering design patterns that solve common business problems in a secure and storage-optimized manner. Each pattern includes a user-facing description of the problem, solutions, and consequences that place the pattern into the context of real-life scenarios. Throughout this journey, you’ll use open source data tools and public cloud services to apply each pattern. You’ll learn: • Challenges data engineers face and their impact on data systems • How these challenges relate to data system components • Useful applications of data engineering patterns • How to identify and fix issues with your current data components • Technology-agnostic solutions to new and existing data projects, with open source implementation examples Data Engineering Design Patterns “This book is the seminal work for the future of data engineering design patterns and should be required reading for any data professional. It is as important to the future of the profession as the Gang of Four’s Design Patterns was for software design.” Scott Haines, coauthor, Delta Lake: The Definitive Guide “Data engineering often feels like solving the same problems over and over. Bartosz Konieczny changes that with this book. Covering everything from idempotency to error handling and data observability, this is the definitive guide to building resilient data pipelines with reusable, proven design patterns.” Adi Polak, director, Confluent Bartosz Konieczny is a freelance data engineer who’s been coding since 2010. He’s held various senior hands- on positions that allowed him to work on many data engineering problems in batch and stream processing.
Praise for Data Engineering Design Patterns The book you now have in your hands is the seminal work for the future of data engineering design patterns. This should be required reading for any data professional, and is as important to the future of the profession as the Gang of Four’s design patterns were for software design. —Scott Haines, coauthor, Delta Lake: The Definitive Guide Data engineering often feels like solving the same problems over and over—Bartosz Konieczny changes that with this book. Covering everything from idempotency to error handing and data observability, this is the definitive guide to building resilient data pipelines with reusable, proven design patterns. —Adi Polak, director at Confluent Bartosz has made a great contribution to drive data engineering forward! Data engineering is the technical backbone on which the leading tech companies build their dominance, and this knowledge needs to spread beyond the technical elite. Data Engineering Design Patterns is a great step in that direction. It captures years of experience crafting solutions to common data engineering challenges. It gives names and concise descriptions to recurring architectural patterns that are folklore among data engineering veterans from the Hadoop age, and can now spread to a wider audience. —Lars Albertsson, data engineering entrepreneur Stoked to see that some of the data engineering principles I’ve advocated for in the past— like immutability, deterministic transformations, and idempotency—are not only taking root but getting expanded upon and developed to a whole new level in this book. A great resource for data engineers looking to build reliable, scalable pipelines. —Maxime Beauchemin, original creator of Apache Airflow and Superset
Data engineering suffers from a glut of complexity due to an ongoing proliferation of languages, frameworks and tools. This book provides clear roadmaps for solving data engineering problems regardless of the underlying technology being applied. —Matt Housely, coauthor, Fundamentals of Data Engineering A must-read for data engineers and architects, this book distills complex data engineering challenges into actionable, well-structured design patterns that empower teams to build scalable, maintainable data systems with confidence. —Keith Mascarenhas, data engineering leader A must-read for any data professional, this book transforms complex data engineering challenges into practical, actionable patterns. Offering a blueprint for mastering the art of building resilient data pipelines with ease, this is an indispensable guide that bridges the gap between theory and practice in data engineering. —Lipi Patnaik, senior software developer, Zeta This book is an invaluable resource for data engineers, offering practical patterns that simplify building robust, scalable data systems. An essential read for mastering the art of data pipeline design. —William Jamir Silva, senior software engineer, Adjust GmbH This book has to be one of the best data engineering design patterns books I’ve come across. The content is easy to understand and grasp, and the author has made every effort to cover the most critical and commonly used design patterns. Every design pattern problem and solution reminded me of challenges and solutions I came up with at work. This book is a must-read for data engineers looking to level up in their workplace. —Rahul Arulkumaran, Foundry Digital
Bartosz Konieczny Data Engineering Design Patterns Recipes for Solving the Most Common Data Engineering Problems
978-1-098-16581-9 [LSI] Data Engineering Design Patterns by Bartosz Konieczny Copyright © 2025 Bartosz Konieczny. All rights reserved. Printed in the United States of America. Published by O’Reilly Media, Inc., 1005 Gravenstein Highway North, Sebastopol, CA 95472. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Aaron Black Development Editor: Michele Cronin Production Editor: Jonathon Owen Copyeditor: Doug McNair Proofreader: Piper Content Partners Indexer: BIM Creatives, LLC Interior Designer: David Futato Cover Designer: Karen Montgomery Illustrator: Kate Dullea April 2025: First Edition Revision History for the First Edition 2025-04-14: First Release See http://oreilly.com/catalog/errata.csp?isbn=9781098165819 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. Data Engineering Design Patterns, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the author and do not represent the publisher’s views. While the publisher and the author have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the author disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights.
Table of Contents Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ix 1. Introducing Data Engineering Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 What Are Design Patterns? 1 Yet More Design Patterns? 3 Common Data Engineering Patterns 3 Case Study Used in This Book 5 Summary 6 2. Data Ingestion Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 Full Load 8 Pattern: Full Loader 8 Incremental Load 12 Pattern: Incremental Loader 12 Pattern: Change Data Capture 16 Replication 20 Pattern: Passthrough Replicator 20 Pattern: Transformation Replicator 24 Data Compaction 27 Pattern: Compactor 27 Data Readiness 30 Pattern: Readiness Marker 30 Event Driven 33 Pattern: External Trigger 33 Summary 37 3. Error Management Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39 Unprocessable Records 40 v
Pattern: Dead-Letter 40 Duplicated Records 46 Pattern: Windowed Deduplicator 46 Late Data 51 Pattern: Late Data Detector 51 Pattern: Static Late Data Integrator 58 Pattern: Dynamic Late Data Integrator 64 Filtering 70 Pattern: Filter Interceptor 70 Fault Tolerance 74 Pattern: Checkpointer 74 Summary 77 4. Idempotency Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 Overwriting 80 Pattern: Fast Metadata Cleaner 80 Pattern: Data Overwrite 86 Updates 89 Pattern: Merger 89 Pattern: Stateful Merger 94 Database 100 Pattern: Keyed Idempotency 100 Pattern: Transactional Writer 105 Immutable Dataset 110 Pattern: Proxy 110 Summary 113 5. Data Value Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115 Data Enrichment 116 Pattern: Static Joiner 116 Pattern: Dynamic Joiner 121 Data Decoration 125 Pattern: Wrapper 125 Pattern: Metadata Decorator 129 Data Aggregation 133 Pattern: Distributed Aggregator 133 Pattern: Local Aggregator 137 Sessionization 141 Pattern: Incremental Sessionizer 141 Pattern: Stateful Sessionizer 147 Data Ordering 153 Pattern: Bin Pack Orderer 153 vi | Table of Contents
Pattern: FIFO Orderer 158 Summary 162 6. Data Flow Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 165 Sequence 166 Pattern: Local Sequencer 166 Pattern: Isolated Sequencer 170 Fan-In 174 Pattern: Aligned Fan-In 175 Pattern: Unaligned Fan-In 178 Fan-Out 182 Pattern: Parallel Split 182 Pattern: Exclusive Choice 186 Orchestration 191 Pattern: Single Runner 191 Pattern: Concurrent Runner 193 Summary 195 7. Data Security Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197 Data Removal 198 Pattern: Vertical Partitioner 198 Pattern: In-Place Overwriter 203 Access Control 207 Pattern: Fine-Grained Accessor for Tables 207 Pattern: Fine-Grained Accessor for Resources 211 Data Protection 215 Pattern: Encryptor 215 Pattern: Anonymizer 219 Pattern: Pseudo-Anonymizer 222 Connectivity 226 Pattern: Secrets Pointer 226 Pattern: Secretless Connector 228 Summary 231 8. Data Storage Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 233 Partitioning 234 Pattern: Horizontal Partitioner 234 Pattern: Vertical Partitioner 240 Records Organization 243 Pattern: Bucket 243 Pattern: Sorter 246 Read Performance Optimization 251 Table of Contents | vii
Pattern: Metadata Enhancer 251 Pattern: Dataset Materializer 254 Pattern: Manifest 257 Data Representation 260 Pattern: Normalizer 260 Pattern: Denormalizer 266 Summary 271 9. Data Quality Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 273 Quality Enforcement 274 Pattern: Audit-Write-Audit-Publish 274 Pattern: Constraints Enforcer 281 Schema Consistency 284 Pattern: Schema Compatibility Enforcer 284 Pattern: Schema Migrator 290 Quality Observation 293 Pattern: Offline Observer 293 Pattern: Online Observer 298 Summary 302 10. Data Observability Design Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 305 Data Detectors 306 Pattern: Flow Interruption Detector 306 Pattern: Skew Detector 310 Time Detectors 314 Pattern: Lag Detector 314 Pattern: SLA Misses Detector 317 Data Lineage 321 Pattern: Dataset Tracker 321 Pattern: Fine-Grained Tracker 325 Summary 328 Afterword. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 331 Appendix: Summary of Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 333 Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 339 viii | Table of Contents
1 I was very into the topic of cloud and data engineering between 2020 and 2022. Following this interest, I self- published the book back in 2022. The first chapter is available for free on my site. Preface As a data engineer coming from the software engineering world, design patterns have always accompanied me on my journey. Adapter design pattern helped me write a backend with pluggable I/O abstractions, Template Method design pattern let me write an easily adaptable business logic, and thanks to the Builder design pattern, I could set up an easily maintainable unit tests layer. Having these great experiences in mind, I have been looking for similar standardized solutions since my first day in the data engineering space. Over time, in each new project, I found something that was similar to previous projects. By connecting these dots I first completed a list of data engineering patterns for cloud services.1 Meantime, I have been continuing to enrich my data engineering design patterns list, despite working on different business domains and with different technologies. That’s how, by summer 2023, I ended up with a quite solid list of data engineering design patterns that I included in the proposal for this book—which, since you are holding the book in your hands, was accepted. I hope the book will add a missing standardization piece each data engineer can rely on to identify a problem, its solu‐ tion, and warning points, and I also hope it will help data engineers work with the data engineering tools of tomorrow. ix
Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Constant width Used for program listings, as well as within paragraphs to refer to program ele‐ ments such as variable or function names, databases, data types, environment variables, statements, and keywords. This element signifies a general note. This element indicates a warning or caution. The Structure of This Book This book follows the workflow of a classical data engineering project that starts with data ingestion and ends with day-to-day monitoring. The steps of the project corre‐ spond to the main chapters, so you can easily identify the stage to which each pattern in a chapter applies. Additionally, each chapter has a two-level structure, with the levels being design pat‐ tern categories and the design patterns themselves. Why this two-level organization? First, a given data engineering problem can have at least two possible solutions, and it wouldn’t be possible to logically group them without having this first level of design pattern categories. Second, data engineering design patterns have their own names that sometimes sound mysterious, and design pattern categories provide extra appli‐ cation context that helps you know where to apply a particular pattern without requiring you to delve into details. Finally, for each pattern, you’ll find the following subsections: Problem This subsection provides a real-world example of when you can use the pattern. x | Preface
Solution This subsection describes the pattern in more technical detail. Usually, it starts with a high-level explanation followed by the technical implementation model. Consequences Patterns have their trade-offs, and this section explains what you should look for before implementing them. Whenever possible, each gotcha is completed with a mitigation solution. Examples In this final part, you’ll find code snippets explaining how to use the pattern within the modern data engineering tools. Unfortunately, it’s not technically pos‐ sible to share the pattern’s implementation in all existing data tools, so this book uses popular open source projects (Apache Spark, Apache Flink, Apache Airflow, PostgreSQL, and Delta Lake). Occasionally, the implementation extends the scope to the managed services in the public cloud. The code snippets are written in Python, SQL, and sometimes Scala or Java if the Python implementation is not available. At the end of this book, you will find a table summarizing all the described patterns. Also, the book has a GitHub repository that includes a glossary of terms that should give you the definitions of the most frequently used acronyms in the book. How to Use This Book It depends on your experience. If you’ve just started your data engineering journey, you probably haven’t seen most of the presented problems yet. In that case, reading this book cover to cover is a good approach. On the other hand, if you already have some significant experience and some of the problems look familiar, reading the book start to finish may not be the best idea. Instead, you can start by picking the patterns you haven’t heard about. To complete the picture, you can later go back to the patterns you know and see if you have imple‐ mented them in the same way. Also, no matter what your level of experience may be, code snippets will help you put the theory into practice and better understand what the pattern implementations can look like in your projects. Once you’ve read the book and played with the code snippets, you can consider it a reference book that’s almost like a cookbook that you can consult whenever you’re fac‐ ing a new data problem to solve (and not because of the flan recipe from Chapter 1). Preface | xi
What Should I Know Prior to Reading This Book? This book will not be a great resource if you have just started in data engineering and don’t have any commercial experience. In my opinion, a minimum of six months of commercial experience with data engineering should help you grasp the ideas more easily. Other than that, the minimum technical knowledge you’ll need to get the most out of this book is as follows: • Familiarity with data engineering concepts, such as extract, transform, load (ETL), extract, load, transform (ELT), data warehousing, data ingestion, and data orchestration. • Cloud awareness. Even though this book tends to favor open source technologies, there are places where cloud technology is more appropriate (e.g., data security). You don’t have to be a cloud expert, but you should at least be able to understand the basics, such as what a managed service is. • Hands-on experience with data processing logic in Java, Scala, Python, or SQL. Ideally, you have already deployed this logic in production. If you feel like there are gaps in your knowledge of the required topics, you should be able to easily fill in the gaps by reading Fundamentals of Data Engineering by Joe Reis and Matt Housley (O’Reilly, 2022). The book provides a comprehensive overview of the data engineering space that will not only help you understand the content in this book but also better prepare you to deal with the challenges you will face in your day- to-day work. Glossary and Code Examples Code examples are available for download on my GitHub page. To run them, you’ll need your favorite IDE and Docker installed with Docker Compose. The GitHub repo has the same organizational structure as the book, meaning each chapter has its own directory where you’ll find fully working examples of the patterns organized into subdirectories. Each directory has a dedicated README.md that will guide you through the demo. Also, the GitHub repository includes a glossary with the most important terms referred to in the book. It’s a complimentary resource that should help you recall some concepts if you haven’t heard about them for a while. This book is here to help you get your job done. In general, if example code is offered with this book, you may use it in your programs and documentation. You do not need to contact us for permission unless you’re reproducing a significant portion of the code. For example, writing a program that uses several chunks of code from this book does not require permission. Selling or distributing examples from O’Reilly xii | Preface
books does require permission. Answering a question by citing this book and quoting example code does not require permission. Incorporating a significant amount of example code from this book into your product’s documentation does require permission. We appreciate, but generally do not require, attribution. An attribution usually includes the title, author, publisher, and ISBN. For example: “Data Engineering Design Patterns by Bartosz Konieczny (O’Reilly). Copyright 2025 Bartosz Konieczny, 978-1-098-16581-9.” If you feel your use of code examples falls outside fair use or the permission given above, feel free to contact us at permissions@oreilly.com. O’Reilly Online Learning For more than 40 years, O’Reilly Media has provided technol‐ ogy and business training, knowledge, and insight to help companies succeed. Our unique network of experts and innovators share their knowledge and expertise through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and 200+ other publishers. For more information, visit https://oreilly.com. How to Contact Us Please address comments and questions concerning this book to the publisher: O’Reilly Media, Inc. 1005 Gravenstein Highway North Sebastopol, CA 95472 800-889-8969 (in the United States or Canada) 707-827-7019 (international or local) 707-829-0104 (fax) support@oreilly.com https://oreilly.com/about/contact.html We have a web page for this book, where we list errata, examples, and any additional information. You can access this page at https://oreil.ly/dataEngDesignPatterns. For news and information about our books and courses, visit https://oreilly.com. Find us on LinkedIn: https://linkedin.com/company/oreilly-media. Preface | xiii
Watch us on YouTube: https://youtube.com/oreillymedia. Acknowledgments Writing a book for O’Reilly has always been a dream. It turns out that I’ve made it come true, and although only my name is on the cover, my journey wouldn’t have been possible without the support and inspiration I have received from my loved ones and my data community. That’s why, before I let you discover the first patterns, I owe a few thank-yous! First and foremost, thanks to my family. To Sylwia, my wife, who has always been supportive, even when I had serious doubts about writing, freelancing, and program‐ ming. To Maja and Arthur, my lovely children, who have reminded me there is a life beyond the screens. Thank you for making the whole book-writing process more enjoyable and less monotonous, even though you unexpectedly shortened some of my miracle morning routines. ;) Also a big thank-you to Jarek and Hania, my parents, whose trust helped me grow and try impossible things, like writing this book! In addition to my family, I couldn’t have written this book without the inspiring peo‐ ple I have met in the data community. To start with, thank you to Jacek Laskowski, who has been a mentor to me since I saw his StackOverflow answers and read his The Internals of... books while I was learning Apache Spark. I’m very glad I could meet you in person and learn even more! Also, a big thank-you to Scott Haines, whose contributions to the data engineering space helped me grasp various stream processing and lakehouse aspects. And I prob‐ ably never would have tried to write this book without the feedback you shared with me at DAIS 2022. Thank you for that chat and all your involvement in the data com‐ munity that, I’m sure, has helped many more than just me! If you’re reading this book, it’s also thanks to the involvement of tech reviewers. Keith Mascharenas, Laura Uzcategui, Lipi Deepaakshi Patnaik, Matthew Housley, Scott Haines, and William Jamir Silva, thank you so much for your questions and all the technical details that helped me grow as a writer and engineer. I’m sure the readers will now appreciate the book even more! At this time, I’d also like to thank Rahul Arulkumaran, Leszek Michalak, and Jonathan Roussot for their valuable feedback after the first Early Release versions of this book. Besides the tech reviewers, Aaron Black and Michele Cronin from O’Reilly helped make this book come alive. Thank you, Aaron, for our discussions about the book and your guidance for the proposal writing. Michele, without your involvement, the book wouldn’t be here, for sure. I’m very glad that we could collaborate, and I wish all other authors could be supported by a development editor like you! xiv | Preface
Finally, I’d like to give a special mention to people who greatly contributed to my pro‐ fessional growth. To start with, thanks to Frank Pavageau, who taught me clean code principles before I knew there was a dedicated book on them! Also, a big thank-you to Jérôme Guibert, who taught me how to leverage automation to make my work eas‐ ier and more enjoyable! And besides the project knowledge shared by Frank and Jérôme, I’ve been learning a lot from the data community. Big thank-yous go to Adi Polak, Holden Karau, Itai Yaffe, and Jungtaek Lim, who have inspired and taught me over the years through their content contributions to the data engineering world! I wish that you, dear reader, can find such inspirational people around you. Preface | xv
(This page has no text content)
CHAPTER 1 Introducing Data Engineering Design Patterns Design patterns are well established in the software engineering space, but they have only recently begun getting traction in the data engineering world. Consequently, I owe you a few words of introduction and an explanation of what design patterns are in the context of data engineering. What Are Design Patterns? You may be surprised at how many times you rely on patterns in your daily life. Let’s take a look at an example involving cooking and one of my favorite desserts, flan; if you like creamy desserts and haven’t tried flan yet, I highly recommend it! When you want to prepare flan, you need to get all the ingredients and follow a list of prepara‐ tion steps. As an outcome, you get a tasty dessert. Why am I giving this cooking example as the introduction to a technical book about design patterns? It’s because a recipe is a great representation of what a design pattern should be: a predefined and customizable template for solving a problem. How does this flan example apply to this definition? • The ingredients and the list of preparation steps are the predefined template. They give you instructions but remain customizable, as you might decide to use brown sugar instead of white, for example. • There can be a single use or many uses. The flan can be a dessert you’ll share with family at teatime, or it can be a product that you’ll sell to make a living. This is the contextualization of a design pattern. Design patterns always respond to a spe‐ cific problem, which in this example is the problem of how to share a pleasant dessert with friends or how to produce the dessert to generate business revenue. 1
• You can decide to prepare this delicious dessert once or many times, if it happens to be your new favorite. For each new preparation, you won’t reinvent the wheel. Chances are, you’ll rely on the same successful recipe you tried before. That’s the reusability of the pattern. • But you must also be aware that preparing and eating flan has some implications for your life and health. If you prepare it every day, you’ll maybe have less time for sports practice, and as a result, you might have some health issues in the long run. These are the consequences of a pattern. • Finally, the recipe saves you time as it has been tested by many other people before. Additionally, it introduces a common dictionary that will make your life easier when discussing it with other people. Finding a recipe for flan is easier than finding one for caramel custard, which is a less popular name for flan. Now, how does all this relate to data engineering? Again, let’s use an example. You need to process a semi-structured dataset from a continuously running job. From time to time, you might be processing a record with a completely invalid format that will throw an exception and stop your job. But you don’t want your whole job to fail because of that simple malformed record. This is our contextualization. To solve this processing issue, you’ll apply a set of best practices to your data process‐ ing logic, such as wrapping the risky transformation with a try-catch block to cap‐ ture bad records and write them to another destination for analysis. That’s the predefined template. These are the rules you can adapt to your specific use case. For example, you could decide not to send these bad records to another database and instead, simply count their occurrences. Turns out that the example of handling erroneous records without breaking the pipe‐ line has a specific name, dead-lettering. Now, if you encounter the same problem again, but in a slightly different context—maybe while working on an ELT pipeline and performing the transformations in a data warehouse directly—you can apply the same logic. That’s the reusability of the pattern. The Dead-Letter pattern is one of the error management patterns detailed in Chapter 3. However, you shouldn’t follow the Dead-Letter pattern blindly. As with eating a flan every day, implementing the pattern has some consequences you should be aware of. Here, you add extra logic that adds some extra complexity to the codebase. You must be ready to accept this. Finally, a data engineering design pattern represents a holistic picture of a solution for a given problem. It then saves you time and also introduces a common language that can greatly simplify discussions with your teammates or data engineers you have just met. 2 | Chapter 1: Introducing Data Engineering Design Patterns