Share E-Book
Scan to open this page

Scan with your phone to open this page

AuthorJoe Reis, Matt Housley

Data engineering has grown rapidly in the past decade, leaving many software engineers, data scientists, and analysts looking for a comprehensive view of this practice. With this practical book, you'll learn how to plan and build systems to serve the needs of your organization and customers by evaluating the best technologies available in the framework of the data engineering lifecycle. Authors Joe Reis and Matt Housley walk you through the data engineering lifecycle and show you how to stitch together a variety of cloud technologies to serve the needs of downstream data consumers. You'll understand how to apply the concepts of data generation, ingestion, orchestration, transformation, storage, governance, and deployment that are critical in any data environment regardless of the underlying technology. This book will help you: • Get a concise overview of the entire data engineering landscape • Assess data engineering problems using an end-to-end framework of best practices • Cut through marketing hype when choosing data technologies, architecture, and processes • Use the data engineering lifecycle to design and build a robust architecture • Incorporate data governance and security across the data engineering lifecycle

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A vendor-neutral field guide to data engineering that replaces tool-chasing with a durable mental model: the data engineering lifecycle and its undercurrents. Best for software engineers, analysts, and data scientists who need an end-to-end view of how data systems are planned, built, and served. 【Book Arc】 - **Opening (~0%–10%)**: Defines data engineering and introduces the book's organizing idea — the data engineering lifecycle (generation, storage, ingestion, transformation, serving) plus its undercurrents (security, data management, DataOps, data architecture, orchestration, software engineering). Solves the "where do I even start?" problem. - **Early (~10%–30%)**: Establishes the human and technical foundations: the shifting role of the data engineer, the core languages (SQL, Python, a JVM language, bash), and the unreasonable effectiveness of SQL. Solves the "what skills actually matter" question. - **Middle (~30%–55%)**: Walks the lifecycle stages in depth — source systems and how they generate data, ingestion patterns, transformation (batch vs. streaming, cost/ROI, business rules), and storage. Solves the "how do I move and shape data reliably" problem. - **Late (~55%–80%)**: Turns to the undercurrents as first-class concerns — data management (lineage, integration, lifecycle/archival), DataOps, data architecture, orchestration, and software engineering practices. Solves the "how do I keep this system governable and maintainable" problem. - **Ending (~80%–100%)**: Focuses on serving data to consumers — analytics (business, operational, embedded), machine learning, and reverse ETL — plus serving mechanisms like file exchange, databases, streaming, query federation, data sharing, and semantic/metrics layers. Solves the "how does data create downstream value" problem. 【Key Takeaways】 - **The lifecycle, not the tool, is the unit of thinking** (Early): generation → storage → ingestion → transformation → serving gives you a stable frame for evaluating any technology, so you stop chasing hype and start reasoning about fit. - **Undercurrents run across every stage** (Early): security, data management, DataOps, data architecture, orchestration, and software engineering are not phases but cross-cutting concerns — ignoring them produces brittle systems. - **SQL remains the lingua franca of data** (Early): despite the MapReduce era, declarative SQL semantics now power warehouses, lakes, and streaming engines alike; proficiency here is non-negotiable. - **Transformation is where value begins** (Middle): from type-correcting and deduplication to normalization, aggregation, and ML featurization — but each transformation should be justified by cost, ROI, and business rules. - **Batch is a special case of streaming** (Middle): all data starts as a stream; treating batch and streaming as one continuum (rather than rival paradigms) future-proofs your architecture. - **Data management is what separates engineers from technicians** (Middle): lineage, integration/interoperability, and end-of-lifecycle archival/destruction are strategic, not clerical — especially as cloud pay-as-you-go costs make retention a CFO-visible decision. - **Integration increasingly happens through APIs, not direct connections** (Middle): pipelines orchestrate many systems via general-purpose APIs, which lowers per-system complexity but raises orchestration complexity. - **Serving is use-case-driven** (Ending): trust, user identity, self-service vs. curated access, and data-product thinking determine whether analytics, ML, or reverse ETL delivery actually lands. 【Reading Tips】 - **Deep-read the lifecycle and undercurrents chapters** (roughly the first half); skim the tool-name-dense passages — the book deliberately stays vendor-neutral, so extract principles, not product lists. - **Treat the serving chapter as a decision framework**: when you face an analytics or ML delivery question, return to the trust / use-case / self-service / data-definition checklist rather than the specific serving mechanisms. - **Watch for the recurring cost-agility-scalability-simplicity-reuse-interoperability trade-off**; it's the book's implicit scoring rubric for every architectural choice. - **If you're a software engineer or analyst transitioning in**, read the skills/languages section first to calibrate what to learn next, then loop back to the lifecycle. - **Keep a running list of undercurrents as you read each stage** — the payoff is seeing how security, governance, and orchestration thread through ingestion, transformation, and serving. 【Coverage Limits】 This guide is synthesized from stratified excerpts (front matter, table of contents, and selected lifecycle/management passages); later chapters on serving, ML, and specific undercurrents are only partially represented, so details there are inferred from headings and brief excerpts rather than full text.
Excerpt 1
3 Data Engineering Defined 4 The Data Engineering Lifecycle 5 Evolution of the Data Engineer 6 Data Engineering and Data Science 11 Data Engineering Skills a...
View in text
Excerpt 2
ves data engineers the holistic context to view their role. Figure 1-1. The data engineering lifecycle The data engineering lifecycle shifts the conversation...
View in text
Excerpt 3
e engineering support. ML engineers who are partially dedi‐ cated to research often rely on the same support teams for research and production. Data Engineer...
View in text
Excerpt 4
the process of integrating data across tools and processes. As we move away from a single-stack approach to analytics and toward a heterogeneous cloud enviro...
View in text
Excerpt 5
nifests similar characteristics. They possess the technical skills of a data engineer but no longer practice data engineering day to day; they mentor current...
View in text
Excerpt 6
y in tune and current with the state of technology and data. Gone are the days of ivory tower data architecture. In the past, architecture was largely orthog...
View in text
Excerpt 7
at users must learn to identify, accommodate, and optimize. Moving on-premises servers one by one to VMs in the cloud—known as simple lift and shift—is a per...
View in text
Excerpt 8
is nonsensical, but we see this all the time in benchmarks. Asymmetric Optimization The deceit of asymmetric optimization appears in many guises, but here’s...
View in text
Tags
AI categories
DataBig DataBackend
ISBN: 1098108302
Publisher: O'Reilly Media
Publish Year: 2022
Language: English
Pages: 446
File Format: PDF
File Size: 8.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…