Page
1
Production Kubernetes Building Successful Application Platforms Josh Rosso, Rich Lander, Alexander Brand & John Harris
Page
2
(This page has no text content)
Page
3
Josh Rosso, Rich Lander, Alexander Brand, and John Harris Production Kubernetes Building Successful Application Platforms Boston Farnham Sebastopol TokyoBeijing
Page
4
978-1-492-09230-8 [LSI] Production Kubernetes by Josh Rosso, Rich Lander, Alexander Brand, and John Harris Copyright © 2021 Josh Rosso, Rich Lander, Alexander Brand, and John Harris. All rights reserved. Printed in the United States of America. Published by O’Reilly Media, Inc., 1005 Gravenstein Highway North, Sebastopol, CA 95472. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: John Devins Development Editor: Jeff Bleiel Production Editor: Christopher Faucher Copyeditor: Kim Cofer Proofreader: Piper Editorial Consulting, LLC Indexer: Ellen Troutman Interior Designer: David Futato Cover Designer: Karen Montgomery Illustrator: Kate Dullea March 2021: First Edition Revision History for the First Edition 2021-03-16: First Release See http://oreilly.com/catalog/errata.csp?isbn=9781492092308 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. Production Kubernetes, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the authors, and do not represent the publisher’s views. While the publisher and the authors have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the authors disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights. This work is part of a collaboration between O’Reilly and VMware Tanzu. See our statement of editorial independence.
Page
5
Table of Contents Foreword. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . xiii Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . xv 1. A Path to Production. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 Defining Kubernetes 1 The Core Components 2 Beyond Orchestration—Extended Functionality 4 Kubernetes Interfaces 5 Summarizing Kubernetes 7 Defining Application Platforms 7 The Spectrum of Approaches 8 Aligning Your Organizational Needs 10 Summarizing Application Platforms 11 Building Application Platforms on Kubernetes 12 Starting from the Bottom 13 The Abstraction Spectrum 15 Determining Platform Services 16 The Building Blocks 17 Summary 21 2. Deployment Models. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 Managed Service Versus Roll Your Own 24 Managed Services 24 Roll Your Own 24 Making the Decision 25 Automation 26 Prebuilt Installer 26 iii
Page
6
Custom Automation 27 Architecture and Topology 28 etcd Deployment Models 28 Cluster Tiers 29 Node Pools 31 Cluster Federation 32 Infrastructure 35 Bare Metal Versus Virtualized 36 Cluster Sizing 39 Compute Infrastructure 41 Networking Infrastructure 42 Automation Strategies 44 Machine Installations 46 Configuration Management 46 Machine Images 46 What to Install 47 Containerized Components 49 Add-ons 50 Upgrades 52 Platform Versioning 52 Plan to Fail 53 Integration Testing 54 Strategies 55 Triggering Mechanisms 60 Summary 61 3. Container Runtime. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63 The Advent of Containers 64 The Open Container Initiative 65 OCI Runtime Specification 65 OCI Image Specification 67 The Container Runtime Interface 69 Starting a Pod 70 Choosing a Runtime 72 Docker 73 containerd 74 CRI-O 75 Kata Containers 76 Virtual Kubelet 77 Summary 78 iv | Table of Contents
Page
7
4. Container Storage. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 Storage Considerations 80 Access Modes 80 Volume Expansion 81 Volume Provisioning 81 Backup and Recovery 81 Block Devices and File and Object Storage 82 Ephemeral Data 83 Choosing a Storage Provider 83 Kubernetes Storage Primitives 83 Persistent Volumes and Claims 83 Storage Classes 86 The Container Storage Interface (CSI) 87 CSI Controller 88 CSI Node 89 Implementing Storage as a Service 89 Installation 90 Exposing Storage Options 92 Consuming Storage 94 Resizing 96 Snapshots 97 Summary 99 5. Pod Networking. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101 Networking Considerations 102 IP Address Management 102 Routing Protocols 104 Encapsulation and Tunneling 106 Workload Routability 108 IPv4 and IPv6 109 Encrypted Workload Traffic 109 Network Policy 110 Summary: Networking Considerations 112 The Container Networking Interface (CNI) 112 CNI Installation 114 CNI Plug-ins 116 Calico 117 Cilium 120 AWS VPC CNI 123 Multus 125 Additional Plug-ins 126 Summary 126 Table of Contents | v
Page
8
6. Service Routing. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 Kubernetes Services 128 The Service Abstraction 128 Endpoints 135 Service Implementation Details 138 Service Discovery 148 DNS Service Performance 151 Ingress 152 The Case for Ingress 153 The Ingress API 154 Ingress Controllers and How They Work 156 Ingress Traffic Patterns 157 Choosing an Ingress Controller 161 Ingress Controller Deployment Considerations 162 DNS and Its Role in Ingress 165 Handling TLS Certificates 166 Service Mesh 169 When (Not) to Use a Service Mesh 169 The Service Mesh Interface (SMI) 170 The Data Plane Proxy 173 Service Mesh on Kubernetes 175 Data Plane Architecture 179 Adopting a Service Mesh 181 Summary 184 7. Secret Management. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187 Defense in Depth 188 Disk Encryption 189 Transport Security 190 Application Encryption 190 The Kubernetes Secret API 191 Secret Consumption Models 193 Secret Data in etcd 196 Static-Key Encryption 198 Envelope Encryption 201 External Providers 203 Vault 203 Cyberark 203 Injection Integration 204 CSI Integration 208 Secrets in the Declarative World 210 Sealing Secrets 211 vi | Table of Contents
Page
9
Sealed Secrets Controller 211 Key Renewal 214 Multicluster Models 215 Best Practices for Secrets 215 Always Audit Secret Interaction 215 Don’t Leak Secrets 216 Prefer Volumes Over Environment Variables 216 Make Secret Store Providers Unknown to Your Application 216 Summary 217 8. Admission Control. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 219 The Kubernetes Admission Chain 220 In-Tree Admission Controllers 222 Webhooks 223 Configuring Webhook Admission Controllers 225 Webhook Design Considerations 227 Writing a Mutating Webhook 228 Plain HTTPS Handler 229 Controller Runtime 231 Centralized Policy Systems 234 Summary 241 9. Observability. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 243 Logging Mechanics 244 Container Log Processing 244 Kubernetes Audit Logs 247 Kubernetes Events 249 Alerting on Logs 250 Security Implications 251 Metrics 251 Prometheus 251 Long-Term Storage 253 Pushing Metrics 253 Custom Metrics 253 Organization and Federation 254 Alerts 255 Showback and Chargeback 257 Metrics Components 260 Distributed Tracing 269 OpenTracing and OpenTelemetry 269 Tracing Components 270 Application Instrumentation 272 Table of Contents | vii
Page
10
Service Meshes 272 Summary 272 10. Identity. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 273 User Identity 274 Authentication Methods 275 Implementing Least Privilege Permissions for Users 285 Application/Workload Identity 288 Shared Secrets 289 Network Identity 289 Service Account Tokens (SAT) 293 Projected Service Account Tokens (PSAT) 297 Platform Mediated Node Identity 299 Summary 311 11. Building Platform Services. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 313 Points of Extension 314 Plug-in Extensions 314 Webhook Extensions 315 Operator Extensions 316 The Operator Pattern 317 Kubernetes Controllers 317 Custom Resources 318 Operator Use Cases 323 Platform Utilities 323 General-Purpose Workload Operators 324 App-Specific Operators 324 Developing Operators 325 Operator Development Tooling 325 Data Model Design 329 Logic Implementation 331 Extending the Scheduler 347 Predicates and Priorities 348 Scheduling Policies 348 Scheduling Profiles 350 Multiple Schedulers 350 Custom Scheduler 350 Summary 351 12. Multitenancy. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 353 Degrees of Isolation 354 Single-Tenant Clusters 354 viii | Table of Contents
Page
11
Multitenant Clusters 355 The Namespace Boundary 357 Multitenancy in Kubernetes 358 Role-Based Access Control (RBAC) 358 Resource Quotas 360 Admission Webhooks 361 Resource Requests and Limits 363 Network Policies 368 Pod Security Policies 370 Multitenant Platform Services 374 Summary 375 13. Autoscaling. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 377 Types of Scaling 378 Application Architecture 379 Workload Autoscaling 380 Horizontal Pod Autoscaler 380 Vertical Pod Autoscaler 384 Autoscaling with Custom Metrics 387 Cluster Proportional Autoscaler 388 Custom Autoscaling 389 Cluster Autoscaling 389 Cluster Overprovisioning 393 Summary 395 14. Application Considerations. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397 Deploying Applications to Kubernetes 398 Templating Deployment Manifests 398 Packaging Applications for Kubernetes 399 Ingesting Configuration and Secrets 400 Kubernetes ConfigMaps and Secrets 400 Obtaining Configuration from External Systems 403 Handling Rescheduling Events 404 Pre-stop Container Life Cycle Hook 404 Graceful Container Shutdown 405 Satisfying Availability Requirements 407 State Probes 408 Liveness Probes 409 Readiness Probes 410 Startup Probes 411 Implementing Probes 412 Pod Resource Requests and Limits 413 Table of Contents | ix
Page
12
Resource Requests 413 Resource Limits 414 Application Logs 415 What to Log 415 Unstructured Versus Structured Logs 416 Contextual Information in Logs 416 Exposing Metrics 416 Instrumenting Applications 417 USE Method 419 RED Method 419 The Four Golden Signals 419 App-Specific Metrics 419 Instrumenting Services for Distributed Tracing 420 Initializing the Tracer 420 Creating Spans 421 Propagate Context 422 Summary 423 15. Software Supply Chain. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 425 Building Container Images 426 The Golden Base Images Antipattern 428 Choosing a Base Image 429 Runtime User 430 Pinning Package Versions 430 Build Versus Runtime Image 431 Cloud Native Buildpacks 432 Image Registries 434 Vulnerability Scanning 435 Quarantine Workflow 437 Image Signing 438 Continuous Delivery 439 Integrating Builds into a Pipeline 440 Push-Based Deployments 443 Rollout Patterns 445 GitOps 446 Summary 448 16. Platform Abstractions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 449 Platform Exposure 450 Self-Service Onboarding 451 The Spectrum of Abstraction 453 Command-Line Tooling 454 x | Table of Contents
Page
13
Abstraction Through Templating 455 Abstracting Kubernetes Primitives 458 Making Kubernetes Invisible 462 Summary 464 Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 465 Table of Contents | xi
Page
14
(This page has no text content)
Page
15
Foreword It has been more than six years since we publicly released Kubernetes. I was there at the start and actually submitted the first commit to the Kubernetes project. (That isn’t as impressive as it sounds! It was a maintenance task as part of creating a clean repo for public release.) I can confidently say that the success Kubernetes has seen is some‐ thing we had hoped for but didn’t really expect. That success is based on a large com‐ munity of dedicated and welcoming contributors along with a set of practitioners who bridge the gap to the real world. I’m lucky enough to have worked with the authors of Production Kubernetes at the startup (Heptio) that I cofounded with the mission to bring Kubernetes to typical enterprises. The success of Heptio is, in large part, due to my colleagues’ efforts in creating a direct connection with real users of Kubernetes who are solving real prob‐ lems. I’m grateful to each one of them. This book captures that on-the-ground experi‐ ence to give teams the tools they need to really make Kubernetes work in a production environment. My entire professional career has been based on building systems aimed at applica‐ tion teams and developers. It started with Microsoft Internet Explorer and then con‐ tinued with Windows Presentation Foundation and then moved to cloud with Google Compute Engine and Kubernetes. Again and again I’ve seen those building platforms suffer from what I call “The Platform Builder’s Curse.” The people who are building the platforms are focused on a longer time horizon and the challenge of building a foundation that will, hopefully, last decades. But that focus creates a blind spot to the problems that users are having right now. Oftentimes we are so busy building a thing we don’t have the time and problems that lead us to actually use the thing we are building. xiii
Page
16
The only way to defeat the platform builder’s curse is to actively seek information from outside our platform-builder bubble. This is what the Heptio Field Engineering team (and later the VMware Kubernetes Architecture Team—KAT) did for me. Beyond helping a wide variety of customers across industries be successful with Kubernetes, the team is a critical window into the reality of how the “theory” of our platform is applied. This problem is only exacerbated by the thriving ecosystem that has been built up around Kubernetes and the Cloud Native Computing Foundation (CNCF). This includes both projects that are part of the CNCF and those that are in the larger orbit. I describe this ecosystem as “beautiful chaos.” It is a rainforest of projects with vary‐ ing degrees of overlap and maturity. This is what innovation looks like! But, just like exploring a rainforest, exploring this ecosystem requires dedication and time, and it comes with risks. New users to the world of Kubernetes often don’t have the time or capacity to become experts in the larger ecosystem. Production Kubernetes maps out the parts of that ecosystem, when individual tools and projects are appropriate, and demonstrates how to evaluate the right tool for the problems the reader is facing. This advice goes beyond just telling readers to use a particular tool. It is a larger framework for understanding the problem a class of tools solves, knowing whether you have that problem, being familiar with the strengths and weaknesses to different approaches, and offering practical advice for getting going. For those looking to take Kubernetes into production, this information is gold! In conclusion, I’d like to send a big “Thank You” to Josh, Rich, Alex, and John. Their experience has made many customers directly successful, has taught me a lot about the thing that we started more than six years ago, and now, through this book, will provide critical advice to countless more users. — Joe Beda Principal Engineer for VMware Tanzu, Cocreator of Kubernetes, Seattle, January 2021 xiv | Foreword
Page
17
Preface Kubernetes is a remarkably powerful technology and has achieved a meteoric rise in popularity. It has formed the basis for genuine advances in the way we manage soft‐ ware deployments. API-driven software and distributed systems were well estab‐ lished, if not widely adopted, when Kubernetes emerged. It delivered excellent renditions of these principles, which are foundational to its success, but it also deliv‐ ered something else that is vital. In the recent past, software that autonomously con‐ verged on declared, desired state was possible only in giant technology companies with the most talented engineering teams. Now, highly available, self-healing, autoscaling software deployments are within reach of every organization, thanks to the Kubernetes project. There is a future in front of us where software systems accept broad, high-level directives from us and execute upon them to deliver desired out‐ comes by discovering conditions, navigating changing obstacles, and repairing prob‐ lems without our intervention. Furthermore, these systems will do it faster and more reliably than we ever could with manual operations. Kubernetes has brought us all much closer to that future. However, that power and capability comes at the cost of some additional complexity. The desire to share our experiences helping others navi‐ gate that complexity is why we decided to write this book. You should read this book if you want to use Kubernetes to build a production-grade application platform. If you are looking for a book to help you get started with Kuber‐ netes, or a text on how Kubernetes works, this is not the right book. There is a wealth of information on these subjects in other books, in the official documentation, and in countless blog posts and the source code itself. We recommend pairing the consump‐ tion of this book with your own research and testing for the solutions we discuss, so we rarely dive deeply into step-by-step tutorial style examples. We try to cover as much theory as necessary and leave most of the implementation as an exercise to the reader. xv
Page
18
Throughout this book, you’ll find guidance in the form of options, tooling, patterns, and practices. It’s important to read this guidance with an understanding of how the authors view the practice of building application platforms. We are engineers and architects who get deployed across many Fortune 500 companies to help them take their platform aspirations from idea to production. We have been using Kubernetes as the foundation for getting there since as early as 2015, when Kubernetes reached 1.0. We have tried as much as possible to focus on patterns and philosophy rather than on tools, as new tooling appears quicker than we can write! However, we inevi‐ tably have to demonstrate those patterns with the most appropriate tool du jour. We have had major successes guiding teams through their cloud native journey to completely transform how they build and deliver software. That said, we have also had our doses of failure. A common reason for failure is an organization’s misconcep‐ tion of what Kubernetes will solve for. This is why we dive so deep into the concept early on. Over this time we’ve found several areas to be especially interesting for our customers. Conversations that help customers get further on their path to produc‐ tion, or even help them define it, have become routine. These conversations became so common that we decided maybe it’s time to write a book! While we’ve made this journey to production with organizations time and time again, there is only one key consistency across them. This is that the road never looks the same, no matter how badly we sometimes want it to. With this in mind, we want to set the expectation that if you’re going into this book looking for the “5-step program” for getting to production or the “10 things every Kubernetes user should know,” you’re going to be frustrated. We’re here to talk about the many decision points and the traps we’ve seen, and to back it up with concrete examples and anecdotes when appropriate. Best practices exist but must always be viewed through the lens of prag‐ matism. There is no one-size-fits-all approach, and “It depends” is an entirely valid answer to many of the questions you’ll inevitably confront on the journey. That said, we highly encourage you to challenge this book! When working with clients we’re always encouraging them to challenge and augment our guidance. Knowledge is fluid, and we are always updating our approaches based on new features, informa‐ tion, and constraints. You should continue that trend; as the cloud native space con‐ tinues to evolve, you’ll certainly decide to take alternative roads from what we recommended. We’re here to tell you about the ones we’ve been down so you can weigh our perspective against your own. xvi | Preface
Page
19
Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Constant width Used for program listings, as well as within paragraphs to refer to program ele‐ ments such as variable or function names, databases, data types, environment variables, statements, and keywords. Constant width bold Shows commands or other text that should be typed literally by the user. Constant width italic Shows text that should be replaced with user-supplied values or by values deter‐ mined by context. Kubernetes kinds are capitalized, as in Pod, Service, and StatefulSet. This element signifies a tip or suggestion. This element signifies a general note. This element indicates a warning or caution. Using Code Examples Supplemental material (code examples, exercises, etc.) is available for download and discussion at https://github.com/production-kubernetes. If you have a technical question or a problem using the code examples, please send email to bookquestions@oreilly.com. Preface | xvii
Page
20
This book is here to help you get your job done. In general, if example code is offered with this book, you may use it in your programs and documentation. You do not need to contact us for permission unless you’re reproducing a significant portion of the code. For example, writing a program that uses several chunks of code from this book does not require permission. Selling or distributing examples from O’Reilly books does require permission. Answering a question by citing this book and quoting example code does not require permission. Incorporating a significant amount of example code from this book into your product’s documentation does require permission. We appreciate, but generally do not require, attribution. An attribution usually includes the title, author, publisher, and ISBN. For example: “Production Kubernetes by Josh Rosso, Rich Lander, Alexander Brand, and John Harris (O’Reilly). Copyright 2021 Josh Rosso, Rich Lander, Alexander Brand, and John Harris, 978-1-492-09231-5.” If you feel your use of code examples falls outside fair use or the permission given above, feel free to contact us at permissions@oreilly.com. O’Reilly Online Learning For more than 40 years, O’Reilly Media has provided technol‐ ogy and business training, knowledge, and insight to help companies succeed. Our unique network of experts and innovators share their knowledge and expertise through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and 200+ other publishers. For more information, visit http://oreilly.com. How to Contact Us Please address comments and questions concerning this book to the publisher: O’Reilly Media, Inc. 1005 Gravenstein Highway North Sebastopol, CA 95472 800-998-9938 (in the United States or Canada) 707-829-0515 (international or local) 707-829-0104 (fax) xviii | Preface