Module 01 · 3 lessons
Evidence foundations
Define trustworthy telemetry identity, quality, and Spring instrumentation before building operational views.
17 reading lessons · 1 deterministic capstone
Follow one Spring order service from instrumentation and telemetry quality through JVM signals, SLO alerts, incident command, and verified recovery.
Five ordered modules · evidence first
Each lesson separates emitted telemetry, observed evidence, user impact, and durable business outcomes.
Module 01 · 3 lessons
Define trustworthy telemetry identity, quality, and Spring instrumentation before building operational views.
Module 02 · 9 lessons
Use bounded logs, metrics, traces, JVM recordings, GC, thread, and JIT evidence according to each signal's boundary.
Module 03 · 3 lessons
Operate the telemetry chain, define user-centered objectives, and route actionable alerts to owned runbooks.
Module 04 · 2 lessons
Coordinate response, communicate facts, learn without blame, and verify recovery with independent evidence.
Module 05 · Capstone lab
Compare modeled user impact with observed evidence, SLO burn, alert windows, telemetry gaps, and recovery.
Browser-only capstone
The lab is not a benchmark, capacity calculator, SLO compliance audit, incident replay, alerting dry-run, MTTA/MTTR calculator, or production verification.
Course library
Upgraded lessons keep their canonical slugs and browser-local progress keys.
System Design Aug 2, 2026 12 min read
Learn how logs, metrics, traces, and events work together to diagnose backend failures, protect telemetry quality, and verify recovery.
System Design Aug 19, 2026 4 min read
Build trustworthy telemetry contracts for identity, schemas, units, timestamps, missing data, duplicates, and bounded dimensions.
Backend Aug 19, 2026 4 min read
Instrument a Spring order service with Actuator, Micrometer Observation, tracing, OTLP, bounded dimensions, and tested propagation.
System Design Aug 2, 2026 11 min read
Design searchable JSON logs and correlation IDs that connect requests, messages, traces, and failures without leaking secrets or exploding cost.
System Design Aug 2, 2026 11 min read
Use RED, USE, and the four golden signals to diagnose request and resource failures with PromQL, safe labels, histograms, and exemplars.
System Design Aug 2, 2026 10 min read
Trace backend work across HTTP and message queues with W3C context, safe baggage, meaningful spans, sampling boundaries, and broken-trace tests.
Java Aug 19, 2026 4 min read
Use Java 25 Flight Recorder to investigate CPU, allocation, GC, locks, I/O, and virtual threads without overstating profile evidence.
Java Aug 19, 2026 5 min read
Read Java 25 garbage collection logs to distinguish allocation pressure, live-set growth, concurrent work, pauses, and Full GC safely.
Java Aug 19, 2026 5 min read
Diagnose Java 25 G1 young, marking, mixed, evacuation, and humongous-object behavior before making minimal evidence-backed tuning changes.
Java Aug 19, 2026 5 min read
Diagnose generational ZGC allocation headroom, concurrent work, stalls, heap sizing, and latency-throughput tradeoffs on Java 25.
Java Aug 19, 2026 4 min read
Diagnose Java 25 platform and virtual threads, locks, safepoints, CPU saturation, pinning, and dependency waits with bounded evidence.
Java Aug 19, 2026 5 min read
Diagnose Java 25 tiered compilation, compilation queues, deoptimization, code cache pressure, warmup, JFR, and AOT profile boundaries.
System Design Aug 2, 2026 11 min read
Build a resilient vendor-neutral telemetry pipeline with OTLP receivers, memory limits, batching, redaction, retries, queues, and self-monitoring.
System Design Aug 2, 2026 11 min read
Turn user journeys into measurable SLIs, realistic SLOs, error budgets, and multi-window burn-rate alerts that page on sustained impact.
Operations Aug 19, 2026 4 min read
Route symptom-based SLO alerts through actionable pages, tickets, ownership, escalation, handoff, and recovery-focused runbooks.
Operations Aug 19, 2026 4 min read
Coordinate incident command, operations, communication, handoffs, factual timelines, blameless learning, and owned follow-up work.
System Design Aug 2, 2026 11 min read
Follow an evidence-driven incident runbook from SLO alert and triage through metrics, traces, logs, mitigation, recovery verification, and postmortem.
Reliability Labs Aug 19, 2026 4 min read
Run six deterministic browser scenarios that compare modeled user impact with observed telemetry, alert windows, and recovery evidence.
FAQ
No. Every command and query is explanatory text, and the capstone is a deterministic browser-only model with no network, credentials, storage, telemetry query, or analytics write.
No. Recovery requires fresh user outcomes, trustworthy evidence delivery, dependency capacity, and durable order-state verification.
No. Sampling, redaction, pipeline loss, backend retention, and instrumentation gaps define what a trace set can support.
No. Requests, steps, coverage, burn, and recovery are explicit teaching units rather than measurements, compliance results, forecasts, MTTA, or MTTR.