Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: 徐葳

No description

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Flink入门与实战 — Reading Guide ## 【One-Line Pitch】 A hands-on, code-first introduction to Apache Flink 1.6 that walks you from installation through production-grade features like HA clusters, Table API/SQL, and real-world ETL projects — ideal for Java/Scala developers new to stream processing who want to build working Flink applications rather than just learn theory. ## 【Book Arc】 - **Opening (~0%–4%)**: Flink's origins (Berlin TU research project → Apache top-level project), core characteristics (distributed, high-performance, HA, accurate), and the key insight that batch is just a special case of streaming. Covers the three data transmission models and typical use cases like real-time ETL. - **Early (~4%–14%)**: Quick-start chapter with development environment setup (IntelliJ IDEA, Maven with Aliyun mirror), a socket-based word count streaming example, then moves into installation — local mode, Standalone cluster, and Flink on YARN with detailed command analysis and JAR packaging via Maven. - **Early (~14%–25%)**: Deep dive into High Availability — ZooKeeper cluster setup, Standalone HA configuration across multiple nodes, and Flink on YARN HA leveraging YARN's task recovery plus checkpoint metadata in ZooKeeper. Ends with the Scala Shell for interactive experimentation. - **Middle (~25%–46%)**: The API core — DataStream sources (socket, collection, custom, connectors like Kafka/RabbitMQ/Redis), transformations (Map, FlatMap, KeyBy, Reduce, Union, Connect, Split/Select), partitioning strategies, sinks (Redis, files), then DataSet APIs and the relational layer: Table API and SQL with DataStream/DataSet/Table conversions. - **Middle (~46%–50%+)**: Advanced features — serialization options (Avro, Kryo, custom), broadcast variables for distributing lookup data, and accumulators for reliable counters across parallel tasks. Excerpts indicate these are covered with both Java and Scala examples. ## 【Key Takeaways】 - **Batch is a special case of streaming** (Opening): Flink treats all data as streams, with batch being the limit case — this unified model simplifies architecture and lets you use one framework for both processing styles. - **Cache timeout is your latency/throughput dial** (Opening): The buffer timeout value (0 to infinity) trades latency against throughput — set to 0 for lowest latency, infinity for maximum throughput, or anywhere between for custom balance. - **YARN deployment handles resource negotiation automatically** (Early): Flink on YARN uploads config and JARs to HDFS, starts JobManager+AM in one container, then dynamically allocates TaskManager containers — enabling parallel YARN sessions with port offsets per application. - **HA requires ZooKeeper for both Standalone and YARN modes** (Early): Standalone HA uses ZooKeeper for JobManager failover; YARN HA additionally relies on YARN's recovery mechanism plus checkpoint metadata stored in ZooKeeper — so ZooKeeper is non-negotiable for production. - **Connectors cover the ecosystem, but not symmetrically** (Early): Kafka, Kinesis, RabbitMQ, and NiFi provide both Source and Sink support; Cassandra, Elasticsearch, HDFS, and Redis are Sink-only; Twitter is Source-only — plan your data flows accordingly. - **Table API and SQL lower the barrier dramatically** (Middle): You can operate on data like MySQL tables without writing Java functions or manual tuning — SQL is non-programmer friendly and easy to adopt, though the API was still actively evolving with some unsupported features. - **Broadcast variables and accumulators solve distributed data problems** (Middle): Broadcast variables distribute lookup data to all tasks via RichFunction's open() method; accumulators provide accurate cross-parallelism counting that plain local variables can't guarantee. ## 【Reading Tips】 - **Skim Chapter 1** if you already know stream processing concepts — the key insight is the batch-as-stream model and the cache timeout trade-off; the rest is standard framework positioning. - **Deep-read Chapter 3 (installation/deployment)** if you'll run Flink yourself — the HA sections are the most valuable, showing real multi-node ZooKeeper + Flink setup with verification steps; the YARN internals explanation is worth studying even if you use a different deployment. - **Use Chapter 4 as a reference, not a cover-to-cover read** — the API sections are organized by category (Source/Transformation/Sink), so jump to what you need; the Java/Scala dual examples make it easy to follow in your preferred language. - **Pay attention to the custom Source example** (Early ~25%): the run()/cancel() pattern with a loop is the template for all custom sources — master this and you can integrate any data source. - **The final chapters (real-time ETL and reporting projects)** are where it all comes together — if you're short on time, prioritize these to see how the pieces fit in production scenarios. ## 【Coverage Limits】 This guide covers the book's first half in detail (through advanced features like broadcast variables and accumulators). The final two project chapters (real-time ETL and real-time reporting) are mentioned in the preface but not covered in the excerpts — expect hands-on integration of everything learned earlier. ##
Page 13
tware Foundation)的顶级项目之一。截至目前,Flink的版本经过了多 次更新,本书基于1.6版本写作。 Flink是一个开源的流处理框架,它具有以下特点。 分布式:Flink程序可以运行在多台机器上。 高性能:处理性能比较高。 高可用:由于Flink程序本身是稳定的,因此它支持高可用性(High...
View in text
Excerpt 2
如下。 3.4 Flink HA的介绍和使用  37 NameNode :Hadoop中HDFS的主节点进程名称。 DataNode :Hadoop中HDFS的从节点进程名称。 SecondaryNameNode :Hadoop中HDFS的辅助节点名称。 ResourceManager :Hadoop中YARN的...
View in text
Excerpt 3
ther {@link RedisDataType} works only with two variable i.e. name of the list and value to be added. * But for {@link RedisDataType#HASH} and {@link RedisDat...
View in text
Excerpt 4
ort org.apache.fl ink.table.api.java.BatchTableEnvironment; import org.apache.fl ink.table.sinks.CsvTableSink; * Created by xuwei.tech. public class SQLTest...
View in text
Excerpt 5
累加求和,结果为27。 第四次进来一条数据10,则立刻进行累加求和,结果为37。 第8章 Flink Time详解 本章主要针对 Flink Time中的 Event Time、Ingestion Time、Processing Time以及 Watermark进行详细讲解。 8.1 Time Stream数据中...
View in text
Excerpt 6
roducer011<String> myProducer = new FlinkKafkaProducer011<> (brokerList, topic, new SimpleStringSchema()); env.execute("StreamingKafkaSinkScala ") } } 10.3.2...
View in text
Excerpt 7
.sources.kafkaSource.interceptors.i1.type = regex_extractor # 设置正则表达式,匹配指定的数据,这样设置会在数据的header中增加topic参数 a1.sources.kafkaSource.interceptors.i1.regex = "type"...
View in text
Excerpt 8
dd HH:mm:ss"); return sdf.format(new Date()); } public static String getRandomArea(){ String[] types = {"AREA_US","AREA_AR","AREA_IN","AREA_ID"}; Random rand...
View in text
Tags
AI categories
Big DataBackendProgramming Language
Language: Chinese
File Format: PDF
File Size: 5.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…