Page
1
Infrastructure and Application Performance Monitoring Prometheus Up & Running Julien Pivotto & Brian Brazil Second EditionSECOND EDITION 2nd Edition
Page
2
SOF T WARE DEVELOPMENT “Julien and Brian have been absolutely key in developing Prometheus, and their deep expertise is reflected in this book. It offers invaluable practical advice for deploying and using Prometheus in real-world scenarios.” —Julius Volz Cofounder of Prometheus and founder of PromLabs “With best practices and instrumentation advice directly from core Prometheus developers, this book will help you monitor your services with confidence.” — TJ Hoplock Senior Observability SRE, NS1 Prometheus: Up & Running Twitter: @oreillymedia linkedin.com/company/oreilly-media youtube.com/oreillymedia Get up to speed with Prometheus, the metrics-based monitoring system used in production by thousands of organizations. This updated edition provides site reliability engineers, Kubernetes administrators, and software developers with a hands-on introduction to the most important aspects of Prometheus, including dashboarding and alerting, direct code instrumentation, and metric collection from third-party systems with exporters. Prometheus server maintainer Julien Pivotto and core developer Brian Brazil demonstrate how to use the system for application and infrastructure monitoring. You’ll learn the Prometheus setup, Node Exporter, and Alertmanager, and discover how to use these tools for application and infrastructure monitoring. You’ll understand why this open source system has continued to gain popularity. • Monitor your infrastructure with Node Exporter and use collectors for network, disks, and pressure metrics • Discover where and how much instrumentation to apply to your application code • Learn about Grafana, a popular tool for building dashboards • Gain views of your machines and services using discovery, including the new HTTP SD mechanism • Use Prometheus with Kubernetes and examine exporters you can use with containers • Explore Prometheus’s improvements and features, including trigonometry functions • Learn how Prometheus supports security features including TLS and basic authentication Julien Pivotto has been a leading contributor to the Prometheus server and the CNCF ecosystem since 2017. Currently, he is a principal software architect and cofounder at O11y. Brian Brazil is the founder of Robust Perception and a Prometheus core developer. He’s well known in the community, having given countless conference presentations and writing a widely read blog. US $65.99 CAN $82.99 ISBN: 978-1-098-13114-2 SECOND EDITION
Page
3
Julien Pivotto and Brian Brazil Prometheus: Up & Running Infrastructure and Application Performance Monitoring SECOND EDITION Boston Farnham Sebastopol TokyoBeijing
Page
4
978-1-098-13114-2 [LSI] Prometheus: Up & Running, Second Edition by Julien Pivotto and Brian Brazil Copyright © 2023 Julien Pivotto. All rights reserved. Printed in the United States of America. Published by O’Reilly Media, Inc., 1005 Gravenstein Highway North, Sebastopol, CA 95472. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: John Devins Development Editor: Rita Fernando Production Editor: Ashley Stussy Copyeditor: Kim Cofer Proofreader: Sonia Saruba Indexer: Ellen Troutman-Zaig Interior Designer: David Futato Cover Designer: Karen Montgomery Illustrator: Kate Dullea July 2018: First Edition April 2023: Second Edition Revision History for the Second Edition 2023-04-04: First Release See http://oreilly.com/catalog/errata.csp?isbn=9781098131142 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. Prometheus: Up & Running, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the authors, and do not represent the publisher’s views. While the publisher and the authors have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the authors disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights.
Page
5
Table of Contents Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . xi Part I. Introduction 1. What Is Prometheus?. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 What Is Monitoring? 4 A Brief and Incomplete History of Monitoring 6 Categories of Monitoring 7 Prometheus Architecture 11 Client Libraries 12 Exporters 13 Service Discovery 14 Scraping 14 Storage 15 Dashboards 15 Recording Rules and Alerts 16 Alert Management 16 Long-Term Storage 17 What Prometheus Is Not 17 2. Getting Started with Prometheus. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 Running Prometheus 19 Using the Expression Browser 23 Running the Node Exporter 27 Alerting 31 iii
Page
6
Part II. Application Monitoring 3. Instrumentation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 A Simple Program 41 The Counter 43 Counting Exceptions 45 Counting Size 47 The Gauge 47 Using Gauges 48 Callbacks 50 The Summary 50 The Histogram 52 Buckets 53 Unit Testing Instrumentation 56 Approaching Instrumentation 57 What Should I Instrument? 57 How Much Should I Instrument? 59 What Should I Name My Metrics? 60 4. Exposition. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 Python 66 WSGI 66 Twisted 67 Multiprocess with Gunicorn 68 Go 71 Java 72 HTTPServer 73 Servlet 74 Pushgateway 76 Bridges 79 Parsers 80 Text Exposition Format 80 Metric Types 81 Labels 82 Escaping 82 Timestamps 82 check metrics 83 OpenMetrics 83 Metric Types 84 Labels 85 Timestamps 85 iv | Table of Contents
Page
7
5. Labels. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87 What Are Labels? 87 Instrumentation and Target Labels 88 Instrumentation 88 Metric 90 Multiple Labels 90 Child 91 Aggregating 93 Label Patterns 94 Enum 94 Info 96 When to Use Labels 98 Cardinality 99 6. Dashboarding with Grafana. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103 Installation 104 Data Source 106 Dashboards and Panels 107 Avoiding the Wall of Graphs 109 Time Series Panel 109 Time Controls 111 Stat Panel 113 Table Panel 115 State Timeline Panel 117 Template Variables 118 Part III. Infrastructure Monitoring 7. Node Exporter. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 CPU Collector 126 Filesystem Collector 127 Diskstats Collector 128 Netdev Collector 129 Meminfo Collector 130 Hwmon Collector 130 Stat Collector 131 Uname Collector 132 OS Collector 132 Loadavg Collector 132 Pressure Collector 133 Textfile Collector 134 Table of Contents | v
Page
8
Using the Textfile Collector 135 Timestamps 137 8. Service Discovery. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139 Service Discovery Mechanisms 140 Static 141 File 142 HTTP 145 Consul 146 EC2 148 Relabeling 149 Choosing What to Scrape 150 Target Labels 153 How to Scrape 162 metric_relabel_configs 164 Label Clashes and honor_labels 166 9. Containers and Kubernetes. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169 cAdvisor 169 CPU 170 Memory 171 Labels 171 Kubernetes 172 Running in Kubernetes 172 Service Discovery 174 kube-state-metrics 184 Alternative Deployments 185 10. Common Exporters. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187 Consul 187 MySQLd 189 Grok Exporter 191 Blackbox 194 ICMP 195 TCP 199 HTTP 201 DNS 204 Prometheus Configuration 205 11. Working with Other Monitoring Systems. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209 Other Monitoring Systems 209 InfluxDB 211 vi | Table of Contents
Page
9
StatsD 212 12. Writing Exporters. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215 Consul Telemetry 215 Custom Collectors 219 Labels 223 Guidelines 224 Part IV. PromQL 13. Introduction to PromQL. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 229 Aggregation Basics 229 Gauge 229 Counter 231 Summary 232 Histogram 233 Selectors 235 Matchers 235 Instant Vector 237 Range Vector 238 Subqueries 240 Offset 241 At Modifier 242 HTTP API 242 query 242 query_range 245 14. Aggregation Operators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249 Grouping 249 without 250 by 251 Operators 252 sum 252 count 253 avg 254 group 255 stddev and stdvar 255 min and max 256 topk and bottomk 256 quantile 257 count_values 258 Table of Contents | vii
Page
10
15. Binary Operators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 261 Working with Scalars 261 Arithmetic Operators 261 Trigonometric Operator 263 Comparison Operators 263 Vector Matching 265 One-to-One 266 Many-to-One and group_left 268 Many-to-Many and Logical Operators 271 Operator Precedence 275 16. Functions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 277 Changing Type 277 vector 278 scalar 278 Math 279 abs 279 ln, log2, and log10 279 exp 280 sqrt 280 ceil and floor 281 round 281 clamp, clamp_max, and clamp_min 281 sgn 282 Trigonometric Functions 282 Time and Date 283 time 283 minute, hour, day_of_week, day_of_month, day_of_year, days_in_month, month, and year 284 timestamp 285 Labels 286 label_replace 286 label_join 286 Missing Series, absent, and absent_over_time 287 Sorting with sort and sort_desc 288 Histograms with histogram_quantile 288 Counters 289 rate 289 increase 291 irate 291 resets 292 Changing Gauges 293 viii | Table of Contents
Page
11
changes 293 deriv 293 predict_linear 294 delta 294 idelta 294 holt_winters 295 Aggregation Over Time 295 17. Recording Rules. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 297 Using Recording Rules 297 When to Use Recording Rules 300 Reducing Cardinality 300 Composing Range Vector Functions 302 Rules for APIs 302 How Not to Use Rules 303 Naming of Recording Rules 304 Part V. Alerting 18. Alerting. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 311 Alerting Rules 312 for 314 Alert Labels 316 Annotations and Templates 318 What Are Good Alerts? 321 Configuring Alertmanagers in Prometheus 322 External Labels 323 19. Alertmanager. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 325 Notification Pipeline 325 Configuration File 326 Routing Tree 327 Receivers 334 Inhibitions 344 Alertmanager Web Interface 345 Part VI. Deployment 20. Server-Side Security. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 351 Security Features Provided by Prometheus 351 Table of Contents | ix
Page
12
Enabling TLS 351 Advanced TLS Options 353 Enabling Basic Authentication 354 21. Putting It All Together. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 357 Planning a Rollout 357 Growing Prometheus 359 Going Global with Federation 360 Long-Term Storage 363 Running Prometheus 365 Hardware 365 Configuration Management 367 Networks and Authentication 368 Planning for Failure 370 Alertmanager Clustering 372 Meta- and Cross-Monitoring 373 Managing Performance 374 Detecting a Problem 375 Finding Expensive Metrics and Targets 375 Reducing Load 376 Horizontal Sharding 377 Managing Change 379 Getting Help 379 Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 381 x | Table of Contents
Page
13
Preface This book describes in detail how to use the Prometheus monitoring system to moni‐ tor, graph, and alert on the performance of your applications and infrastructure. This book is intended for application developers, system administrators, and everyone in between. Expanding the Known When it comes to monitoring, knowing that the systems you care about are turned on is important, but that’s not where the real value is. The big wins are in understand‐ ing the performance of your systems. By performance we don’t only mean the response time of and CPU used by each request, but the broader meaning of performance. How many requests to the data‐ base are required for each customer order that is processed? Is it time to purchase higher throughput networking equipment? How many machines are your cache misses costing? Are enough of your users interacting with a complex feature in order to justify its continued existence? These are the sorts of questions that a metrics-based monitoring system can help you answer, and beyond that help you dig into why the answer is what it is. We see monitoring as getting insight from throughout your system, from high-level overviews down to the nitty-gritty details that are useful for debugging. A full set of monitoring tools for debugging and analysis includes not only metrics, but also logs, traces, and profiling; but metrics should be your first port of call when you want to answer systems-level questions. Prometheus encourages you to have instrumentation liberally spread across your systems, from applications all the way down to the bare metal. With instrumentation you can observe how all your subsystems and components are interacting, and convert unknowns into knowns. xi
Page
14
The Evolution of Prometheus As Prometheus has crossed the 10-year mark, this second edition brings new devel‐ opments across all sections. Prometheus has continued to evolve and expand, offering even more options for scraping, storing, and querying data. This progress is a result of the dedicated community of users and contributors who use Prometheus across a wide and growing range of industries and applications. The second edition of this book provides coverage of the many new PromQL func‐ tions, service discovery providers, and Alertmanager receivers that have been added since the first edition. A new dedicated chapter covers server-side security features, such as TLS, that have been added to Prometheus and some of the exporters. Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Constant width Used for program listings, as well as within paragraphs to refer to program elements such as variable or function names, databases, data types, environment variables, statements, and keywords. Constant width bold Shows commands or other text that should be typed literally by the user. Constant width italic Shows text that should be replaced with user-supplied values or by values deter‐ mined by context. This element signifies a tip or suggestion. This element signifies a general note. xii | Preface
Page
15
This element indicates a warning or caution. Using Code Examples Supplemental material (code examples, configuration files, etc.) is available for down‐ load at https://github.com/prometheus-up-and-running-2e/examples. If you have a technical question or a problem using the code examples, please send email to bookquestions@oreilly.com. This book is here to help you get your job done. In general, if example code is offered with this book, you may use it in your programs and documentation. You do not need to contact us for permission unless you’re reproducing a significant portion of the code. For example, writing a program that uses several chunks of code from this book does not require permission. Selling or distributing examples from O’Reilly books does require permission. Answering a question by citing this book and quoting example code does not require permission. Incorporating a significant amount of example code from this book into your product’s documentation does require permission. We appreciate, but generally do not require, attribution. An attribution usually includes the title, author, publisher, and ISBN. For example: “Prometheus: Up & Running, Second Edition by Julien Pivotto and Brian Brazil (O’Reilly). Copyright 2023 Julien Pivotto, 978-1-098-13114-2.” If you feel your use of code examples falls outside fair use or the permission given above, feel free to contact us at permissions@oreilly.com. O’Reilly Online Learning For more than 40 years, O’Reilly Media has provided technol‐ ogy and business training, knowledge, and insight to help companies succeed. Our unique network of experts and innovators share their knowledge and expertise through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and 200+ other publishers. For more information, visit https://oreilly.com. Preface | xiii
Page
16
How to Contact Us Please address comments and questions concerning this book to the publisher: O’Reilly Media, Inc. 1005 Gravenstein Highway North Sebastopol, CA 95472 800-998-9938 (in the United States or Canada) 707-829-0515 (international or local) 707-829-0104 (fax) We have a web page for this book, where we list errata, examples, and any additional information. You can access this page at https://oreil.ly/prometheus-up-running-2e. Email bookquestions@oreilly.com to comment or ask technical questions about this book. For news and information about our books and courses, visit https://oreilly.com. Find us on LinkedIn: https://linkedin.com/company/oreilly-media Follow us on Twitter: https://twitter.com/oreillymedia Watch us on YouTube: https://youtube.com/oreillymedia Acknowledgments This book would not have been possible without all the work of the Prometheus team, and the hundreds of contributors to Prometheus and its ecosystem. A special thanks to Julius Volz, Richard Hartmann, Carl Bergquist, Andrew McMillan, and Greg Stark for providing feedback on initial drafts of the first revision of this book. Thanks to Brian Brazil, Bartłomiej Płotka, Carl Bergquist, TJ Hoplock, and Richard Hartmann for their feedback on the second edition. xiv | Preface
Page
17
PART I Introduction This section will introduce you to monitoring in general, and Prometheus more specifically. In Chapter 1 you will learn about the many different meanings of monitoring and approaches to it, the metrics approach that Prometheus takes, and the architecture of Prometheus. In Chapter 2 you will get your hands dirty running a simple Prometheus setup that scrapes machine metrics, evaluates queries, and sends alert notifications.
Page
18
(This page has no text content)
Page
19
1 Kubernetes was the first member. CHAPTER 1 What Is Prometheus? Prometheus is an open source, metrics-based monitoring system. Of course, Prome‐ theus is far from the only one of those out there, so what makes it notable? Prometheus does one thing and it does it well. It has a simple yet powerful data model and a query language that lets you analyze how your applications and infrastructure are performing. It does not try to solve problems outside of the metrics space, leaving those to other more appropriate tools. Since its beginnings with no more than a handful of developers working in Sound‐ Cloud in 2012, a community and ecosystem has grown around Prometheus. Prome‐ theus is primarily written in Go and licensed under the Apache 2.0 license. There are hundreds of people who have contributed to the project itself, which is not controlled by any one company. It is always hard to tell how many users an open source project has, but we estimate that as of 2022, hundreds of thousands of organizations are using Prometheus in production. In 2016 the Prometheus project became the second member1 of the Cloud Native Computing Foundation (CNCF). For instrumenting your own code, there are client libraries in all the popular languages and runtimes, including Go, Java/JVM, C#/.Net, Python, Ruby, Node.js, Haskell, Erlang, and Rust. Many popular applications are already instrumented with Prometheus client libraries, like Kubernetes, Docker, Envoy, and Vault. For third- party software that exposes metrics in a non-Prometheus format, there are hundreds of integrations available. These are called exporters, and include HAProxy, MySQL, PostgreSQL, Redis, JMX, SNMP, Consul, and Kafka. A friend of Brian’s even added support for monitoring Minecraft servers, as he cares a lot about his frames per second. 3
Page
20
2 Next to the simple text format, a more standardized version, slightly different, called OpenMetrics has been created out of the Prometheus text format. A simple text format2 makes it easy to expose metrics to Prometheus. Other monitor‐ ing systems, both open source and commercial, have added support for this format. This allows all of these monitoring systems to focus more on core features, rather than each having to spend time duplicating effort to support every single piece of software a user like you may wish to monitor. The data model identifies each time series not just with a name, but also with an unordered set of key-value pairs called labels. The PromQL query language allows aggregation across any of these labels, so you can analyze not just per process but also per datacenter and per service or by any other labels that you have defined. These can be graphed in dashboard systems such as Grafana and Perses. Alerts can be defined using the exact same PromQL query language that you use for graphing. If you can graph it, you can alert on it. Labels make maintaining alerts easier, as you can create a single alert covering all possible label values. In some other monitoring systems you would have to individually create an alert per machine/application. Relatedly, service discovery can automatically determine what applications and machines should be scraped from sources such as Kubernetes, Con‐ sul, Amazon Elastic Compute Cloud (EC2), Azure, Google Compute Engine (GCE), and OpenStack. For all these features and benefits, Prometheus is efficient and simple to run. A single Prometheus server can ingest millions of samples per second. It is a single, statically linked binary with a configuration file. All components of Prometheus can be run in containers, and they avoid doing anything fancy that would get in the way of configuration management tools. It is designed to be integrated into the infrastructure you already have and built on top of, not to be a management platform itself. Now that you have an overview of what Prometheus is, let’s step back for a minute and look at what is meant by “monitoring” in order to provide some context. Follow‐ ing that, we will look at what the main components of Prometheus are, and what Prometheus is not. What Is Monitoring? In secondary school, one of Brian’s teachers told him that if you were to ask 10 econ‐ omists what economics means, you’d get 11 answers. Monitoring has a similar lack of consensus as to what exactly it means. When he tells others what he does, people think his job entails everything from keeping an eye on temperature in factories, to 4 | Chapter 1: What Is Prometheus?