Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Gwen Shapira, Todd Palino, Rajini Sivaram, Krit Petty

Rating No ratings yet

Every enterprise application creates data, whether it consists of log messages, metrics, user activity, or outgoing messages. Moving all this data is just as important as the data itself. With this updated edition, application architects, developers, and production engineers new to the Kafka streaming platform will learn how to handle data in motion. Additional chapters cover Kafka's AdminClient API, transactions, new security features, and tooling changes. Engineers from Confluent and LinkedIn responsible for developing Kafka explain how to deploy production Kafka clusters, write reliable event-driven microservices, and build scalable stream processing applications with this platform. Through detailed examples, you'll learn Kafka's design principles, reliability guarantees, key APIs, and architecture details, including the replication protocol, the controller, and the storage layer. You'll examine: Best practices for deploying and configuring Kafka Kafka producers and consumers for writing and reading messages Patterns and use-case requirements to ensure reliable data delivery Best practices for building data pipelines and applications with Kafka How to perform monitoring, tuning, and maintenance tasks with Kafka in production The most critical metrics among Kafka's operational measurements Kafka's delivery capabilities for stream processing systems

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practitioner's field manual for moving data in motion: this guide walks architects, developers, and production engineers from Kafka's core concepts through reliable producers, consumers, and the operational realities of running clusters at scale. Best suited to engineers who need to design, build, and operate real event-driven systems rather than just understand Kafka in theory. 【Book Arc】 - **Opening (~0%–10%)**: Frames the problem Kafka solves — modern companies are hundreds of interconnected applications, and data movement is as important as storage. Introduces topics, partitions, producers, consumers, and the stream-processing idea, plus the second-edition foreword on Kafka's explosive adoption. - **Early (~10%–35%)**: Gets you running. Covers installation, broker and cluster configuration, ZooKeeper ensemble sizing, retention and log-segment tuning, and OS-level concerns like page cache and swappiness — the groundwork for a stable deployment. - **Middle (~35%–60%)**: The client APIs. Producer design, `ProducerRecord` creation, error handling, partitioning, serializers (including why to prefer Avro/JSON/Protobuf over hand-rolled ones), interceptors, and the consumer side: consumer groups, rebalance, static membership, the poll loop, and the configuration knobs that govern them. - **Late (~60%–85%)**: Reliability and operations. Delivery guarantees, replication, the controller, storage-layer internals, plus AdminClient, transactions, security features, and the monitoring/tuning/maintenance tasks that keep production clusters healthy. - **Ending (~85%–100%)**: Stream processing and data pipelines — how Kafka's delivery capabilities support frameworks like Kafka Streams, Samza, and Storm, and how to assemble end-to-end pipelines and event-driven applications. 【Key Takeaways】 - **Partitions are the unit of everything** (Early): scalability, redundancy, and ordering all hinge on partitions; a "stream" is usually just a topic regardless of partition count. Understanding this unlocks the rest of the book. - **Producer reliability is a configuration choice** (Middle): the `acks` setting trades latency against durability — `acks=all` waits for in-sync replicas and is safest but slowest. Pick deliberately, not by default. - **Don't hand-roll serializers** (Middle): custom byte-format serializers are fragile and create cross-team compatibility nightmares; use Avro, JSON, Thrift, or Protobuf instead. - **Consumer groups and rebalance drive read behavior** (Middle): group membership, static membership, and the poll loop determine how data is distributed and how failures are handled — misconfiguring timeouts causes cascading rebalances. - **Retention operates on segments, not messages** (Early): `log.segment.bytes` and retention-by-size/time interact in surprising ways; the book recommends choosing either size- or time-based retention to avoid unexpected data loss. - **OS tuning matters more than you'd think** (Early): Kafka leans heavily on the page cache; low `vm.swappiness` (1, not 0 on modern kernels) and dirty-page handling directly affect producer response times. - **Interceptors enable observability and governance** (Middle): producer interceptors can count messages, add lineage headers, or redact sensitive data — a lightweight hook for cross-cutting concerns. - **Operations is a first-class topic** (Late): monitoring, the most critical metrics, tuning, and maintenance are treated as core skills, not afterthoughts. 【Reading Tips】 - **Skim the installation and ZooKeeper chapters** if you're on a managed Kafka service; deep-read the producer/consumer chapters even if you think you know them — the configuration rationale is the real value. - **Treat configuration parameters as the spine**: the book repeatedly explains *why* a default exists and when to change it. Take notes on `acks`, retention, and consumer timeout settings. - **Use the credit-card transaction example** in the producer chapter as a mental model for event-driven design; it recurs as a concrete pattern for pipelines. - **Slow down on reliability and replication**: the delivery-guarantee and controller material is where production incidents are prevented or caused. - **Pair the stream-processing chapter with your actual framework** (Kafka Streams, Samza, Storm) rather than reading it abstractly. 【Coverage Limits】 The excerpts are heavily weighted toward the front and middle of the book (setup, producers, consumers); the later chapters on transactions, security, AdminClient, and stream processing are referenced but their detailed content is not covered here. Chapter-level specifics beyond those named in the table of contents should be verified against the book itself.
Excerpt 1
55 acks 55 Message Delivery Time 56 linger.ms 59 buffer.memory 59 compression.type 59 batch.size 59 max.in.flight.requests.per.connection 60 max.request.size...
View in text
Excerpt 2
way of referring to messages is most common when discussing stream processing, which is when frameworks—some of which are Kafka Streams, Apache Samza, and St...
View in text
Excerpt 3
red until the log segment is closed, if log.retention.ms is set to 604800000 (1 week), there will actually be up to 17 days of messages retained until the cl...
View in text
Excerpt 4
ompatibility issues between different versions of serializ‐ ers and deserializers is fairly challenging: you need to compare arrays of raw bytes. To make mat...
View in text
Excerpt 5
ned by the previous poll. It doesn’t know which events were actually processed, so it is critical to always process all the events returned by poll() before...
View in text
Excerpt 6
c topic. We can specify any number of partitions and topics. If you call the command with null instead of a collection of partitions, it will trigger the ele...
View in text
Excerpt 7
• Checksum for validating that the batch is not corrupted. • Sixteen bits indicating different attributes: compression type, timestamp type (timestamp can be...
View in text
Excerpt 8
he number of events produced (usually as events per second). The consumers need to record the number of events consumed per unit or time, and the lag from th...
View in text
Tags
AI categories
DataBackendSoftware
ISBN: 1492043087
Publisher: O'Reilly Media
Publish Year: 2022
Language: English
Pages: 488
File Format: PDF
File Size: 6.0 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…