Files
nexus/sreweekly/markdown/393/01-github-teivah-sre-roadmap-an-opinionated-roadmap-to-become-an-sre-conc.md
2026-09-12 17:23:01 +08:00

224 lines
5.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GitHub – teivah/sre-roadmap: An Opinionated Roadmap to Become an SRE (Concepts > Tools)
- **期号**: SRE Weekly Issue #393(2023-10-08)
- **作者**: Teiva Harsanyi
- **链接**: https://github.com/teivah/sre-roadmap
## 简介
This repo contains a path to learn SRE, in the form of a list of concepts to familiarize oneself with.
## 正文
An opinionated roadmap to become an SRE (Concepts > Tools)
- *Concepts* - Fallacies of distributed computing
- Synchronous vs. asynchronous
- Event log vs. message queue
- Exactly-once delivery
- Different types of message failure
- Orchestration vs. choreography
- Causality
- CDN
- *Hashing* - Consistent hashing
- Geohashing
- Perfect hashing
- Read-heavy vs. write-heavy impacts
- Federation
- *Latency* - Latency, throughput, goodput
- Latency numbers every programmer should know
- How to prevent latency variability
- Tail latency
- How to reduce sharing
- Idempotency
- *Load balancer* - Concepts
- Layer 4 vs. layer 7 load balancer
- Liveness vs. safety properties
- Microservices: pros and cons
- REST
- gRPC
- Service mesh
- Source of truth
- Stateful vs. stateless
- Total vs. partial order
- Why can't we rely on the system clock in distributed systems
- Vector clock
- *Cache* - When to use a cache
- Cache-aside vs. read-through
- Eviction policy
- Refresh-ahead
- Write-through vs. write-back
- Distributed cache
- Performance cache vs. capacity cache
- *Databases* - *Different types of databases* - NoSQL vs. SQL databases
- Relational vs. document
- Column-oriented databases
- Graph databases
- Vector database
- Objects-based storage
- ACID
- *Partitioning* - Criteria
- Methods
- Replication vs. partition
- Hotspot
- CALM theorem
- CAP theorem
- PACELC theorem
- Cardinality
- Chain replication
- Consensus
- Concurrency control
- Consistency models
- Isolation levels
- Serializability
- Linearizability
- CRDT
- *Indexes* - Tradeoff
- Primary vs. secondary indexes
- Denormalization
- View & materialized view
- Transaction
- Distributed transactions downsides
- Strategies to handle rebalancing
- Leader election
- MVCC
- N+1 select problem
- Quorum
- Raft
- Read repair
- Single-leader, multi-leader, leaderless replication
- Split-brain
- 2PC
- 3PC
- WAL
- Write and read amplification
- *Data structure* - *Probabilistic data structures* - Bloom filter
- Count-min sketch
- HyperLogLog
- *Storage* - LSM tree
- B-tree
- SSTable
- *Concepts* - Difference between availability, resiliency, robustness, fault-tolerance, and reliability
- Why is it wrong to target 100% availability
- Blast radius
- Failure domain
- Cascading failures
- Hard vs. soft dependencies
- *Scalability* - Concepts
- Knee point
- Ceiling
- Number one source of outages
- Tail tolerance
- Toil
- *Patterns/Anti-patterns* - Bulkhead pattern
- Circuit breaker
- Exponential backoff
- Jitter
- Graceful degradation
- Load shedding
- Retry amplification
- Backpressure
- Rate limiting
- Request hedging
- *Practices* - Chaos engineering
- *Concepts* - What's the difference between monitoring and observability
- Trace vs. metric vs. log
- Golden signals
- Observer effect
- Percentile
- Streetlight anti-method
- Time-series based monitoring lies
- USE method
- Main metrics for cache
- Why should we be careful about average performance metrics
- *Alerting* - Alerting strategy
- Alerting fatigue concept
- Characteristic of a good alert
- Slow vs. fast burn alert
- *Concepts* - Bake time
- Feature flag
- Feature freeze
- Rollout supervision
- *Rollout types* - Blue green rollout
- Canary rollout
- Progressive rollout
- Shadow rollout
- *Concepts* - SLI vs. SLO vs. SLA
- Error budget
- *SLO* - Difference between KPIs and SLOs
- Benefits of having alerts based on SLOs
- Why is exceeding an SLO not necessarily a good thing
- SLO for data (freshness, completeness, consistency, etc.)
- SLO for mobiles
- SLO for services
- Container
- Container orchestration
- Scripting
- Filesystem
- Memory
- Processes
- Resource utilization
- Network
- ARP protocol
- Bandwidth
- BGP
- CoDel
- CORS
- DNS
- Ping vs. heartbeat
- *TCP* - TCP vs. UDP
- Congestion control
- Connection backlog
- Flow control
- Handshake
- HTTP
- HTTP/2
- Head of line blocking
- Health checks: passive vs. active
- Internet model
- NTP
- OSI model
- Routers
- Switch
- Network topologies
- What happens if you type google.com in your browser
- Authentication
- Certificate
- Certificate authority
- Cipher
- Confidentiality
- Encryption
- TLS
- PKI
- Signature
- Core analysis loop
- Correlation vs. causation
- First principle
- Five whys technique
- *Incident management* - How to address an incident (assess, mitigate, resolve)
- Incident roles
- How to write a postmortem
- 3C principles (Coordinate, Communicate, maintain Control)
- SRE role
- Version control
- *Communication* - Writing
- Oral
- Presentation
- The XY problem
- Collaboration
- Problem solving
- Curiosity
- Navigating ambiguity
- Staying humble