Build & maintain complex distributed systems
October 1–2, 2017: Training
October 2–4, 2017: Tutorials & Conference
New York, NY

Have you tried turning it off and turning it on again?

Tanya Reilly (Squarespace)
2:25pm3:05pm Wednesday, October 4, 2017
DevOps & Tools, Systems Engineering
Location: Gramercy
Average rating: *****
(5.00, 5 ratings)

Who is this presentation for?

  • Systems engineers and architects

Prerequisite knowledge

  • A basic familiarity with distributed systems

What you'll learn

  • Explore disaster recovery best practices

Description

Most of us have a backup strategy, many of us have a restore strategy, and several of us have even fully tested these strategies. But even simple sites may be difficult to recover after a disaster. Tanya Reilly explains why backups are not enough. Complex systems are much harder to reason about and can even be coupled together in ways that make them unrecoverable.

Tanya explores the parts of disaster recovery you might be less prepared for and the dependencies that you might not think about until one day when you really do turn an entire service, entire site, or (perish the thought) an entire company off and on again. You’ll learn why the best laid fallback plans tend to go wrong and why you should start deliberately managing your dependencies long before you think you need to. Along the way, Tanya also covers the dependency cycles that make it difficult or impossible to restart groups of systems—like where do you store the documentation on how to recover the documentation server?

Photo of Tanya Reilly

Tanya Reilly

Squarespace

Tanya Reilly is a principal software engineer at Squarespace working on infrastructure and site reliability. Before Squarespace she spent 12 years in Site Reliability Engineering at Google. Originally from Ireland, she is now an enthusiastic New Yorker. She likes raspberry pi, coding on trains and building systems that are hard to break. She blogs at http://noidea.dog.