Page
1
Jason Morgan & Flynn Linkerd: Up & Running A Guide to Operationalizing a Kubernetes-Native Service Mesh
Page
2
CLOUD COMPUTING “An easy-to-read and easy-to-understand tutorial-style introduction to Linkerd. Provides actionable advice for running Linkerd in production.” —Benjamin Muschko Independent Consultant and Trainer Linkerd: Up & Running linkedin.com/company/oreilly-media youtube.com/oreillymedia With the massive adoption of microservices, operators and developers face far more complexity in their applications today. Service meshes can help you manage this problem by providing a unified control plane to secure, manage, and monitor your entire network. This practical guide shows you how the Linkerd service mesh enables cloud native developers—including platform and site reliability engineers—to solve the thorny issue of running distributed applications in Kubernetes. Jason Morgan and Flynn draw on their years of experience at Buoyant—the creators of Linkerd—to demonstrate how this service mesh can help ensure that your applications are secure, observable, and reliable. You’ll understand why Linkerd, the original service mesh, can still claim the lowest time to value of any mesh option available today. • Learn how Linkerd works and which tasks it can help you accomplish • Install and configure Linkerd in an imperative and declarative manner • Secure interservice traffic and set up secure multicluster links • Create a zero trust authorization strategy in Kubernetes clusters • Organize services in Linkerd to override error codes, set custom retries, and create timeouts • Use Linkerd to manage progressive delivery and pair this service mesh with the ingress of your choice Jason Morgan is a DevOps practitioner who has helped many organizations on their cloud native journeys. Jason helps teams adopt cloud native ways of working so they can deliver for their customers and learn how to go fast forever. Jason has given talks, written a number of articles, and contributes to the CNCF. Flynn is a Linkerd tech evangelist at Buoyant, the original author of Emissary-ingress, a co-lead of the GAMMA initiative, and a frequent speaker and writer in the field. He’s been a part of the cloud native world since 2016 and has spent more than 40 years in computing, with a focus on communications and security. US $65.99 CAN $82.99 ISBN: 978-1-098-14231-5
Page
3
Jason Morgan and Flynn Linkerd: Up and Running A Guide to Operationalizing a Kubernetes-Native Service Mesh Boston Farnham Sebastopol TokyoBeijing
Page
4
978-1-098-14231-5 [LSI] Linkerd: Up and Running by Jason Morgan and Flynn Copyright © 2024 Jason Morgan and Kevin Hood. All rights reserved. Printed in the United States of America. Published by O’Reilly Media, Inc., 1005 Gravenstein Highway North, Sebastopol, CA 95472. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: John Devins Development Editor: Angela Rufino Production Editor: Gregory Hyman Copyeditor: Penelope Perkins Proofreader: Rachel Head Indexer: Sue Klefstad Interior Designer: David Futato Cover Designer: Karen Montgomery Illustrator: Kate Dullea April 2024: First Edition Revision History for the First Edition 2024-04-11: First Release See http://oreilly.com/catalog/errata.csp?isbn=9781098142315 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. Linkerd: Up and Running, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the authors and do not represent the publisher’s views. While the publisher and the authors have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the authors disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights. This work is part of a collaboration between O’Reilly and Buoyant. See our statement of editorial independence.
Page
5
Table of Contents Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ix 1. Service Mesh 101. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 Basic Mesh Functionality 2 Security 4 Reliability 5 Observability 6 How Do Meshes Actually Work? 9 So Why Do We Need This? 11 Summary 11 2. Intro to Linkerd. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 Where Does Linkerd Come From? 13 Linkerd1 14 Linkerd2 15 The Linkerd Proxy 15 Linkerd Architecture 15 mTLS and Certificates 17 Certifying Authorities 18 The Linkerd Control Plane 19 Linkerd Extensions 20 Summary 24 3. Deploying Linkerd. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 Considerations 25 Linkerd Versioning 25 Workloads, Pods, and Services 26 TLS certificates 27 iii
Page
6
Linkerd Viz 28 Deploying Linkerd 29 Required Tools 29 Provisioning a Kubernetes Cluster 29 Installing Linkerd via the CLI 30 Installing Linkerd via Helm 31 Configuring Linkerd 34 Cluster Networks 35 Linkerd Control Plane Resources 35 Opaque and Skip Ports 35 Summary 36 4. Adding Workloads to the Mesh. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 Workloads Versus Services 37 What Does It Mean to Add a Workload to the Mesh? 38 Injecting Individual Workloads 40 Injecting All Workloads in a Namespace 40 linkerd.io/inject Values 40 Why Might You Decide Not to Add a Workload to the Mesh? 41 Other Proxy Configuration Options 42 Protocol Detection 42 When Protocol Detection Goes Wrong 44 Opaque Ports Versus Skip Ports 44 Configuring Protocol Detection 45 Default Opaque Ports 46 Kubernetes Resource Limits 47 Summary 47 5. Ingress and Linkerd. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49 Ingress Controllers with Linkerd 53 The Ingress Controller Is Just Another Meshed Workload 53 Linkerd Is (Mostly) Invisible 55 Use Cleartext Within the Cluster 55 Route to Services, Not Endpoints 56 Ingress Mode 58 Specific Ingress Controller Examples 60 Emissary-ingress 60 NGINX 60 Envoy Gateway 61 Summary 61 iv | Table of Contents
Page
7
6. The Linkerd CLI. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63 Installing the CLI 63 Updating the CLI 64 Installing a Specific Version 64 Alternate Ways to Install 64 Using the CLI 64 Selected Commands 67 linkerd version 67 linkerd check 68 linkerd inject 70 linkerd identity 72 linkerd diagnostics 73 Summary 78 7. mTLS, Linkerd, and Certificates. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 Secure Communications 80 TLS and mTLS 80 mTLS and Certificates 82 Linkerd and mTLS 83 Certificates and Linkerd 83 The Linkerd Trust Anchor 85 The Linkerd Identity Issuer 85 Linkerd Workload Certificates 86 Certificate Lifetimes and Rotation 87 Certificate Management in Linkerd 89 Automatic Certificate Management with cert-manager 89 Summary 96 8. Linkerd Policy: Overview and Server-Based Policy. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97 Linkerd Policy Overview 97 Linkerd Default Policy 98 Linkerd Policy Resources 100 Server-Based Policy Versus Route-Based Policy 102 Server-Based Policy with the emojivoto Application 103 Configuring the Default Policy 103 Configuring Dynamic Policy 106 Summary 119 9. Linkerd Route-Based Policy. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 Route-Based Policy Overview 121 The booksapp Sample Application 122 Installing booksapp 124 Table of Contents | v
Page
8
Configuring booksapp Policy 125 Infrastructure Policy 125 Read-Only Access 127 Enabling Write Access 135 Allowing Writes to books 137 Reenabling the Traffic Generator 138 Summary 140 10. Observing Your Platform with Linkerd. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 Why Do We Need This? 141 How Does Linkerd Help? 141 Observability in Linkerd 142 Setting Up Your Cluster 142 Tap 143 Service Profiles 144 Topology 148 Linkerd Viz 149 Audit Trails and Access Logs 150 Access Logging: The Good, the Bad, and the Ugly 150 Enabling Access Logging 151 Summary 151 11. Ensuring Reliability with Linkerd. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153 Load Balancing 153 Retries 154 Retry Budgets 155 Configuring Retries 155 Configuring the Budget 159 Timeouts 159 Configuring Timeouts 160 Traffic Shifting 162 Traffic Shifting, Gateway API, and the Linkerd SMI Extension 162 Setting Up Your Environment 163 Weight-Based Routing (Canary) 165 Header-Based Routing (A/B Testing) 168 Traffic Shifting Summary 169 Circuit Breaking 170 Enabling Circuit Breaking 170 Tuning Circuit Breaking 171 Summary 172 vi | Table of Contents
Page
9
12. Multicluster Communication with Linkerd. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173 Types of Multicluster Setups 173 Gateway-Based Multicluster 173 Pod-to-Pod Multicluster 174 Gateways Versus Pod-to-Pod 175 Multicluster Certificates 176 Cross-Cluster Service Discovery 176 Setting Up for Multicluster 178 Continuing with a Gateway-Based Setup 181 Continuing with a Pod-to-Pod Setup 182 Multicluster Gotchas 183 Deploying and Connecting an Application 183 Checking Traffic 187 Policy in Multicluster Environments 187 Summary 188 13. Linkerd CNI Versus Init Containers. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 189 Kubernetes sans Linkerd 189 Nodes, Pods, and More 189 Networking in Kubernetes 191 The Role of the Packet Filter 193 The Container Networking Interface 195 The Kubernetes Pod Startup Process 196 Kubernetes and Linkerd 196 The Init Container Approach 196 The Linkerd CNI Plugin Method 197 Races and Ordering 199 Summary 200 14. Production-Ready Linkerd. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201 Linkerd Community Resources 201 Getting Help 201 Responsible Disclosure 202 Kubernetes Compatibility 202 Going to Production with Linkerd 202 Stable or Edge? 202 Preparing Your Environment 202 Configuring Linkerd for High Availability 204 Monitoring Linkerd 207 Certificate Health and Expiration 207 Control Plane 208 Data Plane 208 Table of Contents | vii
Page
10
Metrics Collection 208 Linkerd Viz for Production Use 208 Accessing Linkerd Logs 210 Upgrading Linkerd 211 Upgrading via Helm 212 Upgrading via the CLI 213 Readiness Checklist 214 Summary 215 15. Debugging Linkerd. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 217 Diagnosing Data Plane Issues 217 “Common” Linkerd Data Plane Failures 217 Setting Proxy Log Levels 221 Debugging the Linkerd Control Plane 222 Linkerd Control Plane and Availability 222 The Core Control Plane 222 Linkerd Extensions 225 Summary 225 Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 227 viii | Table of Contents
Page
11
Preface Service meshes need a little reputational rehab. Many cloud native practitioners seem to have in mind that meshes are frightening, complex things, things to be avoided until examined as a last resort to save a dying application. We’d love to change that: service meshes are incredible tools for making developing and operating cloud native applications dramatically easier than it would otherwise be. And, of course, we think Linkerd is the best mesh out there at making things easy for people. So if you’ve been tearing your hair out trying to understand a misbehaving applica‐ tion based just on its logs, or if you’ve spent months trying to get some other mesh running and you just want things to work, or if you’re trying to explain to yet another developer why they really don’t need to worry about coding retries and mTLS into their microservice…you’re in the right place. We’re glad you’re here. Who Should Read This Book This book is meant to help anyone who thinks it’s easier to get things done when creating, running, or debugging microservices applications, and is looking to Linkerd to help with that. While we think that the book will benefit people who are interested in Linkerd for its own sake, Linkerd—like computing itself—is ultimately a means, not an end. This book reflects that. Beyond that, it doesn’t matter to us whether you’re an application developer, a cluster operator, a platform engineer, or whatever; there should be something in here to help you get the most out of Linkerd. Our goal is to give you everything you need to get Linkerd up and running to help you get things done. You’ll need some basic knowledge of Kubernetes, the overall concept of running things in containers, and the Unix command line to get the most out of this book. ix
Page
12
Some familiarity with Prometheus, Helm, Jaeger, etc. will also be helpful, but isn’t really critical. Why We Wrote This Book We’ve both worked in the cloud native world for years and in software for many more years before that. Across all that time, the challenge that has never gone away is education; the coolest new thing on the block isn’t much good until people really, truly understand what it is and how to use it. Service meshes really should be pretty well understood by now, but of course every month there are people who need to sort out the latest and greatest changes in the meshes, and every month there are more people migrating to what is, to them, the entirely new cloud native world. We wrote this book, and we’ll keep updating it, to help all these people out. Navigating This Book Chapter 1, “Service Mesh 101”, is an introduction to service meshes: what they do, what they can help with, and why you might want to use one. This is a must-read for folks who aren’t familiar with meshes. Chapter 2, “Intro to Linkerd”, takes a deep dive into Linkerd’s architecture and history. If you’re familiar with Linkerd already, this may be mostly recap. Chapter 3, “Deploying Linkerd”, and Chapter 4, “Adding Workloads to the Mesh”, are all about getting Linkerd running in a cluster and getting your application working with Linkerd. These two chapters cover the basic nuts and bolts of actually using Linkerd. Chapter 5, “Ingress and Linkerd”, continues by talking about the ingress problem, how to manage it, and how Linkerd interacts with ingress controllers. Chapter 6, “The Linkerd CLI”, talks about the linkerd CLI, which you can use to control and examine a Linkerd deployment. Chapter 7, “mTLS, Linkerd, and Certificates”, dives deep into Linkerd mTLS and the way it uses X.509 certificates. Chapter 8, “Linkerd Policy: Overview and Server-Based Policy”, and Chapter 9, “Linkerd Route-Based Policy”, continue by exploring how Linkerd can use those mTLS identities to enforce policy in your cluster. Chapter 10, “Observing Your Platform with Linkerd”, is all about Linkerd’s application-wide observability mechanisms. Chapter 11, “Ensuring Reliability with Linkerd”, in turn, covers how to use Linkerd to improve reliability within your application, and Chapter 12, “Multicluster Communication with Linkerd”, talks about extending a Linkerd mesh across multiple Kubernetes clusters. x | Preface
Page
13
Chapter 13, “Linkerd CNI Versus Init Containers”, addresses the thorny topic of how, exactly, you’ll have Linkerd interact with the low-level networking configuration of your cluster. Unfortunately, this may be a necessary topic of discussion as you con‐ sider taking Linkerd to production, which is the topic of Chapter 14, “Production- Ready Linkerd”. Finally, Chapter 15, “Debugging Linkerd”, discusses how to troubleshoot Linkerd itself, should you find things misbehaving (even though we hope you won’t!). Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Constant width Used for program listings, as well as within paragraphs to refer to program elements such as variable or function names, databases, data types, environment variables, statements, and keywords. Constant width italic Shows text that should be replaced with user-supplied values or by values deter‐ mined by context. This element signifies a general note. This element indicates a warning or caution. Using Code Examples Supplemental material (code examples, exercises, etc.) is available for download at https://oreil.ly/linkerd-code. If you have a technical question or a problem using the code examples, please send email to support@oreilly.com. This book is here to help you get your job done. In general, if example code is offered with this book, you may use it in your programs and documentation. You Preface | xi
Page
14
do not need to contact us for permission unless you’re reproducing a significant portion of the code. For example, writing a program that uses several chunks of code from this book does not require permission. Selling or distributing examples from O’Reilly books does require permission. Answering a question by citing this book and quoting example code does not require permission. Incorporating a significant amount of example code from this book into your product’s documentation does require permission. We appreciate, but generally do not require, attribution. An attribution usually includes the title, author, publisher, and ISBN. For example: “Linkerd: Up and Run‐ ning by Jason Morgan and Flynn (O’Reilly). Copyright 2024 Jason Morgan and Kevin Hood, 978-1-098-14231-5.” If you feel your use of code examples falls outside fair use or the permission given above, feel free to contact us at permissions@oreilly.com. O’Reilly Online Learning For more than 40 years, O’Reilly Media has provided technol‐ ogy and business training, knowledge, and insight to help companies succeed. Our unique network of experts and innovators share their knowledge and expertise through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and 200+ other publishers. For more information, visit https://oreilly.com. How to Contact Us Please address comments and questions concerning this book to the publisher: O’Reilly Media, Inc. 1005 Gravenstein Highway North Sebastopol, CA 95472 800-889-8969 (in the United States or Canada) 707-827-7019 (international or local) 707-829-0104 (fax) support@oreilly.com https://www.oreilly.com/about/contact.html We have a web page for this book, where we list errata, examples, and any additional information. You can access this page at https://oreil.ly/linkerd-up-and-running. xii | Preface
Page
15
For news and information about our books and courses, visit https://oreilly.com. Find us on LinkedIn: https://linkedin.com/company/oreilly-media Watch us on YouTube: https://youtube.com/oreillymedia Acknowledgments Many, many thanks to the fine folks who helped us develop this book, including (but not limited to!): • Our editor, Angela Rufino • Technical reviewers Daniel Bryant, Ben Muschko, and Swapnil Shevate, who provided amazing feedback that made the book worlds better • The unsung heroes at O’Reilly who got everything into publishable shape • Last but very much not least, the Linkerd maintainers and the fine folks at Buoyant who created the thing that we’re writing about From Flynn, a big shout out to SC and RAH for putting up with him during the year it took to put this together. Many, many thanks. Preface | xiii
Page
16
(This page has no text content)
Page
17
CHAPTER 1 Service Mesh 101 Linkerd is the first service mesh—in fact, it’s the project that coined the term “service mesh.” It was created in 2015 by Buoyant, Inc., as we’ll discuss more in Chapter 2, and for all that time it’s been focused on making it easier to produce and operate truly excellent cloud native software. But what, exactly, is a service mesh? We can start with the definition from the CNCF Glossary: In a microservices world, apps are broken down into multiple smaller services that communicate over a network. Just like your wifi network, computer networks are intrinsically unreliable, hackable, and often slow. Service meshes address this new set of challenges by managing traffic (i.e., communication) between services and adding reliability, observability, and security features uniformly across all services. The cloud native world is all about computing at a huge range of scales, from tiny clusters running on your laptop for development up through the kind of massive infrastructure that Google and Amazon wrangle. This works best when applications use the microservices architecture, but the microservices architecture is inherently more fragile than a monolithic architecture. Fundamentally, service meshes are about hiding that fragility from the application developer—and, indeed, from the application itself. They do this by taking several features that are critical when creating robust applications and moving them from the application into the infrastructure. This allows application developers to focus on what makes their applications unique, rather than having to spend all their time worrying about how to provide the critical functions that should be the same across all applications. 1
Page
18
In this chapter, we’ll take a high-level look at what service meshes do, how they work, and why they’re important. In the process, we’ll provide the background you need for our more detailed discussions about Linkerd in the rest of the book. Basic Mesh Functionality The critical functions provided by services meshes fall into three broad categories: security, reliability, and observability. As we examine these three categories, we’ll be comparing the way they play out in a typical monolith and in a microservices application. Of course, “monolith” can mean several different things. Figure 1-1 shows a diagram of the “typical” monolithic application that we’ll be considering. Figure 1-1. A monolithic application The monolith is a single process within the operating system, which means that it gets to take advantage of all the protection mechanisms offered by the operating sys‐ tem; other processes can’t see anything inside the monolith, and they definitely can’t modify anything inside it. Communications between different parts of the monolith are typically function calls within the monolith’s single memory space, so again there’s no opportunity for any other process to see or alter these communications. It’s true that one area of the monolith can alter the memory in use by other parts—in fact, this is a huge source of bugs!—but these are generally just errors, rather than attacks. 2 | Chapter 1: Service Mesh 101
Page
19
Multiple Processes Versus Multiple Machines “But wait!” we hear you cry. “Any operating system worthy of the name can provide protections that do span more than one process! What about memory-mapped files or System V shared memory segments? What about the loopback interface and Unix domain sockets (to stretch the point a bit)?” You’re right: these mechanisms can allow multiple processes to cooperate and share information while still being protected by the operating system. However, they must be explicitly coded into the application, and they only function on a single machine. Part of the power of cloud native orchestration systems like Kubernetes is that they’re allowed to schedule Pods on any machine in your cluster, and you won’t know which machine ahead of time. This is tremendously flexible, but it also means that mechanisms that assume everything is on a single machine simply won’t work in the cloud native world. In contrast, Figure 1-2 shows the corresponding microservices application. Figure 1-2. A microservices application With microservices, things are different. Each microservice is a separate process, and microservices communicate only over the network—but the protection mechanisms provided by the operating system function only inside a process. These mechanisms aren’t enough in a world where any information shared between microservices has to travel over the network. This reliance on communications over the unreliable, insecure network raises a lot of concerns when developing microservices applications. Basic Mesh Functionality | 3
Page
20
Security Let’s start with the fact that the network is inherently insecure. This gives rise to a number of possible issues, some of which are shown in Figure 1-3. Figure 1-3. Communication is a risky business Some of the most significant security issues are eavesdropping, tampering, identity theft, and overreach: Eavesdropping Evildoers may be able to intercept communications between two microservices, reading communications not intended for them. Depending on what exactly an evildoer learns, this could be a minor annoyance or a major disaster. The typical protection against eavesdropping is encryption, which scrambles the data so that only the intended recipient can understand it. Tampering An evildoer might also be able to modify the data in transit over the network. At its simplest, the tampering attack would simply corrupt the data in transit; at its most subtle, it would modify the data to be advantageous to the attacker. It’s extremely important to understand that encryption alone will not protect against tampering! The proper protection is to use integrity checks like check‐ sums; all well-designed cryptosystems include integrity checks as part of their protocols. 4 | Chapter 1: Service Mesh 101