Reliability Labs

SLO Incident Triage Lab

Compare modeled user impact with incomplete observed evidence, windowed burn-rate alerts, reversible mitigation, and recovery verification.

Local workbench
Input stays in this browserRun with Ctrl/⌘ + Enter

Deterministic browser-only evidence model

SLO Incident Triage Lab

Compare modeled user impact with the evidence an alert can actually observe, then test whether a mitigation has enough follow-up evidence to support recovery.

This is not a benchmark, capacity calculator, SLO compliance audit, incident replay, alerting dry-run, MTTA/MTTR calculator, or production verification. Steps, requests, coverage, burn rates, and recovery are teaching units rather than measured production behavior.

Six fixed teaching scenarios

Choose an evidence and impact pattern

Selecting a card loads its inputs and runs the same deterministic pure function.

01 Workload and SLO

Set the offered work, bounded service capacity, baseline outcomes, and teaching objective.

02 Incident

Place a recent change, incident window, degraded outcomes, and optional reversible mitigation.

03 Evidence

Vary incident-time telemetry coverage and trace retention without rewriting missing evidence as success.

04 Detection

Set bounded short and long evidence windows and a teaching burn-rate threshold.

Ctrl/⌘ + Enter to run

Healthy bounded evidence is ready. Nothing leaves this browser.

Deterministic result

Actual impact and observed evidence timeline

No modeled evidence warning

StepOfferedAdmittedCompletedFailedSlowRejectedBadObservedObserved badSampledSaturationActual burnObserved burnAlertPhaseEvidenceRisk flags

Metrics

Quantify impact and coverage

    Traces

    Inspect retained causal paths

      Logs

      Correlate bounded events

        Changes

        Test recent-change hypotheses

          Dependencies

          Cross the owning boundary

            How to use it

            Choose one of six deterministic scenarios or enter a bounded teaching configuration. Identical inputs always return the same timeline, summary, and evidence plan.

            Read the lab examples and evidence guide, then continue with Production Incident Troubleshooting with Logs, Metrics, and Traces.

            This is not a benchmark, capacity calculator, SLO compliance audit, incident replay, alerting dry-run, MTTA/MTTR calculator, or production verification. Nothing leaves this browser.

            Related tools

            Keep the workflow moving

            These local-first tools often pair well with SLO Incident Triage Lab.

            Available tool

            Java Concurrency Budget Lab

            Model fixed pools, virtual threads, resource guards, deadlines, and cooperative cancellation with a deterministic browser-only timeline.

            Open tool

            FAQ

            SLO Incident Triage Lab questions

            Does this lab query a telemetry backend?

            No. It is a pure deterministic browser model with no network, telemetry query, credential, storage, remote execution, or analytics access.

            Can this lab prove SLO compliance or production recovery?

            No. Requests, steps, coverage, burn rate, and recovery are explicit teaching units rather than measured service evidence.

            Does a firing alert identify the root cause?

            No. The model separates user impact, observed evidence, alert state, recent changes, and the next evidence to inspect.