Share E-Book

Data Engineering for Cybersecurity Build Secure Data Pipelines with Free and Open-Source Tools (James Bonifield)(Z-Library)

Author

Rating No ratings yet

Log in to rate

Cybersecurity
Language English

Security teams rely on telemetry—the continuous stream of logs, events, metrics, and signals that reveal what’s happening across systems, endpoints, and cloud services. But that data doesn’t organize itself. It has to be collected, normalized, enriched, and secured before it becomes useful. That’s where data engineering comes in. In this hands-on guide, cybersecurity engineer James Bonifield teaches you how to design and build scalable, secure data pipelines using free, open source tools such as Filebeat, Logstash, Redis, Kafka, and Elasticsearch and more. You’ll learn how to collect telemetry from Windows including Sysmon and PowerShell events, Linux files and syslog, and streaming data from network and security appliances. You’ll then transform it into structured formats, secure it in transit, and automate your deployments using Ansible.

Format PDF
Size 6.1 MB
10
Views
(First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
(This page has no text content)
Page 2
CONTENTS IN DETAIL ACKNOWLEDGMENTS INTRODUCTION Who This Book Is For What’s in This Book Prerequisites Online Resources PART I: FOUNDATIONS OF SECURE DATA ENGINEERING 1 DATA ENGINEERING BASICS Common Data Engineering Tasks Getting Data to the SIEM Managing Throughput and Latency Enriching and Standardizing Events Example Architectures A Basic Data Pipeline Proactive Data Retrieval Temporary Data Centralization Event Caching Data Serialization Formats JSON YAML Elastic Common Schema Virtual Machine Setup Windows File Transfer and Shell Access Tools Other Useful Tools Summary
Page 3
2 NETWORK ENCRYPTION TLS Components Interactions Mutual TLS Generating a Certificate Authority Root Configuration Root Generation Intermediate Configuration Intermediate Signing Chain Files Setting Up TLS for Logstash Flex Configuration Server Keys and Certificates Directory Preparation Logstash Configuration Firewall Ports SSH Enabling SSH Creating SSH Keys Starting the Agent Creating Configuration Files Distributing Public Keys Verifying Connectivity Summary 3 SOURCE AND CONFIGURATION MANAGEMENT The Git Version Control Workflow Installation Setting Up GitHub or GitLab Creating New Directories Initializing Locally Cloning the Remote Repository Updating README.md Creating a New File Ignoring Files Creating a Development Branch Adding the Branch Creating a New File Setting the Branch’s Remote Origin Merging the Branch into main Performing Optional Cleanup Rebasing and Resetting Your Code Creating a New Test Branch Modifying the Remote main Branch
Page 4
Temporarily Stashing Changes Cleanup Creating a Repository for the Book’s Project Code Summary PART II: LOG EXTRACTION AND MANAGEMENT 4 ENDPOINT AND NETWORK DATA Collecting Logs with Filebeat Installation Enabling TLS Creating a Configuration File Generating Certificate Signing Requests Configuration Input Sources Reading Local Files Applying Parsers and New Fields Listening to the Network Connecting to External Systems Processors Controlling Processors with Conditionals Screening Out Data Modules Sending Outputs Publishing to Kafka Publishing to Redis Writing Output Data to a File Sending Data to Logstash Pruning and Privatizing Data Filebeat for Windows Summary 5 WINDOWS LOGS Collecting Logs with Winlogbeat Event Log Types Application and System Logs The Security Log and Sysmon Other Log Sources Installation Enabling TLS
Page 5
Creating a Configuration File Generating a Key Pair and Client Certificate Configuration Starting the Service Collecting Logs from Sysmon and PowerShell Sysmon PowerShell Script Blocks PowerShell Modules Exploring Available Event Logs Working with Processors Summary 6 INTEGRATING AND STORING DATA An Elastic Agent Architecture Setting Up the Environment Enabling TLS Configuring DNS and Firewall Rules Installing Tools on Multiple Servers Viewing Events in Kibana Configuring Different Technology Integrations Network Metadata Iptables Firewall Logs Integration Guidelines Stand-Alone Agents Preparing the Virtual Machine Receiving a Logstash API Key Configuring a Stand-Alone Agent Policy Configuring Logstash Elasticsearch Ingest Pipelines Creating a Custom Ingest Pipeline Testing the Pipeline Summary 7 WORKING WITH SYSLOG DATA Logging Priorities Collecting Data with Rsyslog Installation Enabling TLS Custom Rsyslog Configurations Basic vs. Advanced Formats Adding and Testing New Configurations Advanced Format Syntax Global Properties Input Modules
Page 6
Output Modules Templates Rulesets Example Configurations A Client A Server Relay Summary PART III: DATA TRANSFORMATION AND STANDARDIZATION 8 DATA MANIPULATION PIPELINES Logstash Pipeline Components Logstash Installation and Setup Enabling TLS Configuring Java Virtual Machine Options Installing Codecs Inputs and Outputs HTTP POST Requests API Requests Reading and Writing Files Kafka Redis Amazon S3 and MinIO Logstash to Logstash Null Output Virtual Addressing in Pipelines Summary 9 TRANSFORMATION FILTERS Logstash Filters for Data Transformation Extracting Data from Messy Inputs Extracting Unstructured Data with grok Extracting Reliably Ordered Data with dissect Handling Inaccurate Syslog Timestamps Enriching Data Network Identification with the CIDR Filter Key-Value Lookups with translate Adding Directional Tags Field Comparison Equality and Inequality
Page 7
Field Existence Membership Regular Expressions Parsing Key-Value Pairs from Structured Syslog Data Splitting Fields and Avoiding Backslashes Input Mutate Output Writing Filters with Ruby Scripting Init and Code Array Comparisons Dropping Data to Save CPU Summary PART IV: DATA CENTRALIZATION, AUTOMATION, AND ENRICHMENT 10 CENTRALIZING SECURITY DATA Kafka Fundamentals Producers, Consumers, and Consumer Groups Brokers and Controllers Topics Partitions Data Replication Listeners and Metadata Streams and Connectors Creating a Kafka Cluster Setting Up the Network Creating Users and Directories Installing Kafka and a Java Development Kit Enabling TLS Configurations Configuring Server Properties Defining the Connection Settings Creating Cluster IDs Using Systemd for Automatic Startup Running Kafka Starting Brokers Creating Topics Publishing Messages Subscribing to Topics Creating Tool-Specific Topics Connecting to External Tools
Page 8
Rsyslog Filebeat Logstash Kafka Pipeline Design Considerations Summary 11 AUTOMATING TOOL CONFIGURATIONS Working with Ansible Inventories Commands, Playbooks, and Roles Gathering System Facts Templates Temporary Files SSH Requirements Installation Creating Inventories Configuration Files Running Ad Hoc Commands Enabling Check Mode Distributing Files Executing Commands Updating Packages Managing Services Rebooting Servers Playbooks for Task Execution Reusable Tasks Delegating Tasks Performing Tasks Locally Writing Dynamic Data to Files Modifying hosts Files Automatically Generating Encryption Keys Appending Text Summary 12 ANSIBLE TASKS AND PLAYBOOKS Allowing Firewall Traffic Automatic File Transfers Distributing the Kafka Package Downloading Files from the Internet Cloning Git Repositories Creating Kafka Configuration Templates Installing GPG Keys Responding to Prompts Accessing APIs and Web Services
Page 9
TLS Automation Setting Up the Directory Creating the Root CA Creating the Intermediate CA Generating the CA Chain Creating a Flex Certificate Adding Container Files Creating the TLS Playbook Troubleshooting with debug Summary 13 CACHING THREAT INTELLIGENCE DATA Speeding Up Lookups with Caching Installing Redis and Memcached Dividing the Project Servers into Nodes Enabling TLS Configuring Redis Configuring Memcached Working with Redis Working with Memcached Populating and Replicating Redis for Threat Lookups Creating the Leader Cache Receiving Threat Indicators with Logstash Defining Inputs Processing Indicators Handling Custom Uploads Populating Redis with Key-Value Pairs Creating Custom API Keys Outputting Data to Elasticsearch Reading Threat Intelligence Creating Redis Followers Configuring a Logstash Getter Deduplicating and Enriching Data Distributing Threat Intelligence with Memcached Configuring the Cyber Threat Intelligence Node Configuring the Data Node Preventing Drift Handling Analyst-Submitted Indicator Data Going Beyond Summary INDEX
Page 10
DATA ENGINEERING FOR CYBERSECURITY Build Secure Data Pipelines with Free and Open Source Tools by James Bonifield San Francisco
Page 11
DATA ENGINEERING FOR CYBERSECURITY. Copyright © 2025 by James Bonifield. All rights reserved. No part of this work may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopying, recording, or by any information storage or retrieval system, without the prior written permission of the copyright owner and the publisher. First printing 29 28 27 26 25    1 2 3 4 5 ISBN-13: 978-1-7185-0402-8 (print) ISBN-13: 978-1-7185-0403-5 (ebook) Published by No Starch Press®, Inc. 245 8th Street, San Francisco, CA 94103 phone: +1.415.863.9900 www.nostarch.com; info@nostarch.com Publisher: William Pollock Managing Editor: Jill Franklin Production Manager: Sabrina Plomitallo-González Production Editor: Sydney Cromwell Developmental Editor: Frances Saux Cover Illustrator: Josh Kemble Interior Design: Octopod Studios Technical Reviewers: Andrew Eastman and David Sanchez Copyeditor: Lisa McCoy Proofreader: Daniel Wolff Library of Congress Control Number: 2025007905 For customer service inquiries, please contact info@nostarch.com. For information on distribution, bulk sales, corporate sales, or translations: sales@nostarch.com. For permission to translate this work: rights@nostarch.com. To report counterfeit copies or piracy: counterfeit@nostarch.com. The authorized representative in the EU for product safety and compliance is EU Compliance Partner, Pärnu mnt. 139b-14, 11317 Tallinn, Estonia, hello@eucompliancepartner.com, +3375690241. No Starch Press and the No Starch Press iron logo are registered trademarks of No Starch Press, Inc. Other product and company names mentioned herein may be the trademarks of their respective owners. Rather than use a trademark symbol with every occurrence of a trademarked name, we are using the names only in an editorial fashion and to the benefit of the trademark owner, with no intention of infringement of the trademark. The information in this book is distributed on an “As Is” basis, without warranty. While every precaution has been taken in the preparation of this work, neither the author nor No Starch Press, Inc. shall have any liability to any person or entity with respect to any loss or damage caused or alleged to be caused directly or indirectly by the information contained in it.
Page 12
For Abby, Ruby, and Lily
Page 13
About the Author James Bonifield is a cybersecurity professional with over a decade of experience analyzing malicious activity, implementing data pipelines, and training others in the security industry. He has built enterprise-scale log solutions, automated many aspects of security operations, and led analyst teams investigating major cyber threat actors. He holds numerous certifications, including the OSCP, GXPN, and CISSP. He enjoys spending free time with his wife and kids, traveling, and tinkering with all things security and Python related. About the Technical Reviewers Andrew Eastman is a seasoned cybersecurity professional with over a decade of experience in threat detection, incident response, and network defense. He currently serves as a cybersecurity professional for Microsoft’s internal corporate blue team, where he focuses on safeguarding enterprise assets against advanced cyber threats. He holds a master’s degree in cybersecurity and information assurance from Western Governors University and has earned a suite of industry-recognized certifications, including CASP+, CySA+, CEH, and CCNA. David Sanchez is a consulting architect at Elastic, where he specializes in detection engineering, artificial intelligence and machine learning integration, and comprehensive cybersecurity solutions. He is adept at collaborating with clients to design and implement tailored, scalable security strategies that address unique challenges and drive business success.
Page 14
ACKNOWLEDGMENTS First, I want to thank my amazing wife, Abby, for always supporting my projects. Thank you for listening to me talk about this book all day, every day. I also want to thank my girls, Ruby and Lily, for being great kids. Thanks to my parents, Jeff and Leigh, and my brother, Jerry, for encouragement along the way. Thank you to Bill Pollock and Jill Franklin for extending an offer to write for No Starch Press. Thanks to my developmental editor, Frances Saux; to my production editor, Sydney Cromwell; to my copyeditor, Lisa McCoy; and to the rest of the amazing team at No Starch Press. Thank you to Andrew Eastman and David Sanchez, the technical reviewers for this book. I’ve had the pleasure of knowing these gentlemen for many years, and I appreciate their insights.
Page 15
INTRODUCTION Data engineering is the practice of creating infrastructure that can receive data from multiple sources, format or consolidate that data, then store it in databases for analysis. Many organizations require data engineers to track information like customer engagement metrics, financial transactions, and stock market movements. But when it comes to cybersecurity, data engineers are critical to connecting sensors to the analysts who can detect incidents. Although generally invisible to all but those who manage them, the tools that ship data from one place to another are foundational to cybersecurity programs. Before an analyst can identify a problem, the relevant alerts and event logs must leave the organization’s workstations and servers and land in a central database. In this book, you’ll learn how to extract, transform, standardize, and streamline events using free and open source tools. You’ll learn how to centralize log management, perform data enrichment, and distribute events to various technologies and databases for analysis.
Page 16
Who This Book Is For This book is for anyone tasked with creating a cybersecurity monitoring system or centralizing a business’s logs. It may be particularly useful for those operating with constrained budgets, but it is relevant to those operating with enterprise-level funding as well. Network defenders can use this book to transform events taken from a multitude of systems and tools into a standard format before storing them. Administrators and engineers can also use this book to manage the many device health logs flowing through the network. Offensive testers can also use it to read and transform the variety of outputs from hacking tools to store them for client reports, automate the comparison of results, and dispatch additional tool actions. Those seeking to automate defensive or offensive actions may find centralizing and standardizing logs useful as well. What’s in This Book We’ll cover a data engineer’s toolkit for tasks including encrypting network communications, extracting data from hosts, and transforming and cleaning that data. We’ll also cover tools that ship data to other places, automation software, and caching tools. Here's a breakdown of the book's parts and what you'll learn in each chapter: Part I: Foundations of Secure Data Engineering This first section lays the foundation for the rest of the book. You’ll configure your own TLS infrastructure, prepare SSH keys, and use Git to centralize and manage your project files. Chapter 1: Data Engineering Basics Learn about the goals of a data engineer: building pipelines that centralize, standardize, and enrich data. Explore common architectures for an organization’s logging infrastructure. Then, survey the JSON, ECS, and YAML formats and install tools we’ll use throughout the book. Chapter 2: Network Encryption Encrypting network data is important for protecting data, but some people find it daunting to configure, so they bypass it when learning a new tool. In this chapter,
Page 17
you’ll become comfortable with using network encryption and set up a TLS infrastructure that you can use to encrypt tool traffic in subsequent chapters. You’ll also become familiar with configuring SSH to connect to remote servers. Chapter 3: Source and Configuration Management Learn to use Git for source management so that, as you create configurations throughout this book, you can track any changes you’ve made and revert them if needed. You’ll use Git locally to manage files and remotely to act as a backup and a hub for your work. Part II: Log Extraction and Management To begin creating data pipelines, you’ll have to extract data from various hosts and the network. This section covers tools and approaches for doing so. You’ll use tools for Linux and Windows as well as ones that receive network data from appliances or sensors. Chapter 4: Endpoint and Network Data Work with Filebeat to extract data from Linux hosts and connect to network services to receive logs. You’ll interact with local files, receive data over the network, and proactively reach out to external sources for data. Chapter 5: Windows Logs Use Winlogbeat to read the Windows event log. Enable and use Windows security features such as Sysmon and enhanced PowerShell logging, and explore uncommon event logs containing useful security data. Chapter 6: Integrating and Storing Data Explore the Elasticsearch database and its frontend user interface, Kibana. Learn to work with Elastic Agent, which combines features from Filebeat, Winlogbeat, and other tools in a browser-based GUI. Read host-level data, then collect network metadata directly from the host. Also learn about ingest pipelines and assets needed to transform log data inside Elasticsearch. Chapter 7: Working with Syslog Data Use Rsyslog to read and write local files and to receive and transmit data over the network. Explore Rsyslog’s plug-ins for OpenSSL and Kafka. Part III: Data Transformation and Standardization
Page 18
Standardization is an important part of data engineering. You’ll learn how to convert and parse multiple data formats from various technologies into one standardized naming convention using JSON structures. This data transformation step is useful for security analysts, automated response tools, and machine learning. Chapter 8: Data Manipulation Pipelines Learn about Logstash’s inputs and outputs, connect systems to Logstash, and link the tool with services such as APIs, Amazon S3, and Redis. Chapter 9: Transformation Filters Use Logstash’s filters to transform and manipulate incoming events to correct timestamps, enrich and delete data, standardize key-value data pair field names, extract values, and write custom Ruby code. Part IV: Data Centralization, Automation, and Enrichment In this part, you’ll use Kafka to centralize your data feeds, enabling you to distribute and analyze your logging events in real time. You’ll learn to automate many of the processes described in previous chapters using Ansible. Finally, you’ll use caching tools to add a cyber threat intelligence enrichment process to your data pipeline. Chapter 10: Centralizing Security Data Configure Kafka to act as central plumbing for your data pipelines. Create topics to keep data organized and then use tools such as Filebeat and Logstash to send and pull data from Kafka. Chapter 11: Automating Tool Configurations Explore Ansible, which allows you to script and automate system administration tasks across multiple hosts at once. Chapter 12: Ansible Tasks and Playbooks Learn various ways of using Ansible to install and configure the tools discussed in this book. Automate your TLS certificate creation and data pipeline deployment, and run multiple automated actions in sequence. Chapter 13: Caching Threat Intelligence Data Use the caching tools Redis and Memcached to store data using an in-memory cache, which speeds up data lookups and enrichment activities. Create a
Page 19
distributed cyber threat intelligence enrichment process directly in your data pipeline. Prerequisites This book assumes a basic familiarity with the Linux operating system and its command line. We’ll walk through all configuration steps and commands to run, but you’ll find it easiest to follow along if you’re used to working in the terminal. Here are some other topics you might want to familiarize yourself with before moving on: Virtual machines Virtual machines are virtualized computers running inside of other computers. For example, you might run a Linux virtual machine on your Windows desktop or laptop. The computer that hosts the virtual machine is called a hypervisor. Virtual machines are helpful for practicing new tools and concepts, as you can take a snapshot of their current state and then return the machine to that state if anything goes wrong. TCP and UDP The tools covered throughout this book make network connections using the Transmission Control Protocol (TCP) and User Datagram Protocol (UDP). TCP is a protocol that requires systems to acknowledge receipt of the data they exchange, making it resilient to connectivity issues. TCP is useful for tasks like file transfers and email delivery, where it’s important not to lose any data. UDP doesn’t require an acknowledgment from a recipient, and it’s useful for tasks like watching videos or listening to audio without constantly being interrupted. If some of the bytes drop while traveling from the server to your screen, the show goes on without a hiccup. HTTP and TLS You’ll use the Hypertext Transfer Protocol (HTTP) to request and receive data from the web. HTTP is an unencrypted protocol, meaning an overly curious or malicious actor on the same network could position themselves to inspect the content of your communications
Page 20
with a web server. Today, most websites use HTTPS, which encrypts traffic with Transport Layer Security (TLS), the “S” in “HTTPS.” We’ll cover TLS further in Chapter 2 and use it to encrypt all network traffic when able, including for data fetched from the web and communications between tools. SSH Secure Shell (SSH) is a technology that creates an encrypted tunnel, or network connection, from one computer to another. It’s often used to remotely control a computer as if you’re using its local keyboard. You’ll configure several virtual machines in this book and then make extensive use of SSH to interact with them. You’ll also explore an automation tool, Ansible, that uses SSH to orchestrate actions on many hosts at once. Chapter 2 discusses SSH in more detail. Scripting languages In this book, you’ll write a few scripts using the Python programming language and use the Ruby programming language to transform data. We’ll walk through these scripts’ contents in detail. That said, you could also automate many of the commands we’ll cover using shell scripts that execute Linux commands. Some experience creating shell, Python, or Ruby scripts is recommended, but not necessary, to follow along. Online Resources You’ll find all of the commands and configurations in this book on GitHub, along with useful tidbits, at https://github.com/bonifield/data-engineering- for-cybersecurity. Each chapter has a corresponding directory containing its code and files. This repository will also contain any updates or corrections, if necessary.
The above is a preview of the first 20 pages. Register to read the complete e-book.

Recommended for You

Loading recommended books...
Failed to load, please try again later

Tip the Site

Scan the WeChat Pay or Alipay code to tip. No login required.

WeChat Pay
Alipay
← Back to List