Training: June 20–21, 2016
Tutorials: June 21, 2016
Keynotes & Sessions: June 22–23, 2016
Santa Clara, CA

A practical guide to monitoring and alerting with time series at scale

Jamie Wilkinson (Google)
11:20am–12:00pm Wednesday, 06/22/2016
First time at Velocity Santa Clara, Measuring the right things
Location: Ballroom GH Level: Intermediate
Average rating: ***..
(3.27, 15 ratings)

Prerequisite knowledge

Attendees should have basic programming and arithmetic experience.

Description

Monitoring is the foundational bedrock of site reliability yet is the bane of most sysadmins’ lives. Why? Monitoring sucks when the cost of maintenance scales proportionally with the size of the system being monitored. Recently, tools like Riemann and Prometheus have emerged to address this problem by scaling out monitoring configurations sublinearly with the size of the system.

In a talk complementing the Google SRE book chapter “Practical Alerting from Time Series Data,” Jamie Wilkinson explores the theory of alert design and time series-based alerting methods and offers practical examples in Prometheus that you can deploy in your environment today to reduce the amount of alert spam and help operators keep a healthy level of production hygiene.

Photo of Jamie Wilkinson

Jamie Wilkinson

Google

Jamie Wilkinson has been a site reliability engineer at Google for over 11 years but is still trying to automate himself out of a job.