Diagrammatic

Design a Self-Healing Infrastructure Platform — System Design Interview Practice

Design a platform that automatically detects, diagnoses, and remediates infrastructure issues without human intervention, using playbook automation and ML-driven decision making. Work through the requirements, architecture trade-offs, and an interactive design review.

Concepts and architecture decisions to consider

  • aiopsConcept to explore
  • self healingConcept to explore
  • automationConcept to explore
  • remediationConcept to explore
  • infrastructureConcept to explore
  • reliabilityConcept to explore

Interview prompt

Design a guarded self-healing infrastructure platform that detects incidents, diagnoses likely causes, executes canary remediations, and proves recovery without creating cascading failures.

  • Define incident signals, evidence, diagnosis confidence, playbook versions, blast-radius limits, approvals, and remediation state transitions.
  • Require idempotent actions, canary execution, health verification, automatic rollback, cooldowns, circuit breakers, and human escalation.
  • Separate detection and decisioning from privileged execution; preserve an immutable audit trail and replayable incident evidence.
  • Explain noisy alerts, conflicting actions, partial failure, secrets, authorization, observability, and a safe manual-only mode.

Requirements and scale assumptions

  • Ingest health signals, correlate incidents, rank hypotheses, select an approved playbook, and execute a bounded canary remediation.
  • Show evidence, confidence, action progress, verification results, rollback status, ownership, and escalation history.
  • Support playbook review/versioning, dry runs, maintenance windows, incident replay, action cancellation, and audit export.
  • Start an approved remediation within five minutes of a high-confidence detection and verify each action before expansion.
  • Handle thousands of concurrent services and incidents without a single hot key or unbounded synchronous work.
  • Do not lose committed state; make retries and duplicate events safe.
  • Degrade safely when downstream workers, caches, or external dependencies fail.
  • Thousands of services with bursty incident fan-out
  • Partition by the primary tenant, user, item, or geographic key and isolate hot partitions.
  • Keep serving state bounded; retain raw events or durable records for replay and auditing.
  • Peak scale: 10k services; 1k concurrent incidents — Capacity assumption that drives partitioning and backpressure.
  • Latency target: start < 5m; canary verified before rollout — User-facing budget for the primary request or read path.
  • Durable boundary: Committed before async — Incident evidence and approved playbook executions are authoritative; derived diagnoses can be recomputed.
  • Async boundary: At-least-once workers — Keep Use decision trees for remediation selection, Implement runbook automation with Ansible/Terraform, Use canary remediation (fix one instance, verify, then all) off the synchronous path.

Key entities

  • ResourceSpecresourceId, tenantId, desiredState, version, policyVersion, updatedAt

    Versioned desired state for a self healing infrastructure platform managed resource.

  • OperationoperationId, resourceId, requestHash, step, attempt, status

    Durable self healing infrastructure platform reconciliation operation with per-step progress.

  • PolicyVersionpolicyId, scope, version, rules, effectiveAt, status

    Auditable self healing infrastructure platform policy evaluated before provisioning or mutation.

  • ReconciliationCheckpointresourceId, provider, observedVersion, cursor, lastError, updatedAt

    Provider-specific self healing infrastructure platform observation and recovery cursor.

Data flow

  1. 1. Accept a desired-state commandThe self healing infrastructure platform control plane authenticates the tenant, validates policy and quotas, checks the expected version, and records the desired state.
  2. 2. Plan a safe operationA planner turns self healing infrastructure platform desired state into ordered, bounded steps with dependency checks, blast-radius limits, and rollback metadata.
  3. 3. Reconcile providers asynchronouslyWorkers apply self healing infrastructure platform operations through provider adapters, persist checkpoints, rate-limit calls, and treat unknown outcomes as observable state.
  4. 4. Publish observed healthThe serving projection joins desired and observed self healing infrastructure platform state with operation status, policy version, freshness, and actionable errors.
  5. 5. Recover and auditRetries, dead letters, drift detection, and operator approvals repair self healing infrastructure platform resources without losing the original command or provider evidence.

Deep dives and trade-offs

  • Desired versus observed stateKeep self healing infrastructure platform desired state separate from provider-observed state and show both to operators. Make every reconciliation step conditional and resumable so a worker crash does not restart unsafe effects. Version policy and resource state so old operations cannot overwrite newer intent.
  • Provider failures and unknown outcomesUse provider-specific idempotency tokens and query-after-timeout behavior for self healing infrastructure platform operations. Bound retries with exponential backoff, circuit breakers, and per-provider quotas. Route irreconcilable drift to an approval or quarantine path instead of retrying forever.
  • Blast radius and operationsPartition self healing infrastructure platform work by tenant, region, cluster, or resource class and cap concurrent mutations. Audit who changed desired state, which policy allowed it, and what provider evidence was observed. Alert on drift age, operation backlog, failed steps, policy denials, and stale observations.
  • Push versus pull reconciliationUse event triggers for fast response and periodic scans for missed events, drift, and recovery. A push-only self healing infrastructure platform controller silently misses changes when a provider event is lost.
  • Central control plane versus provider-native controllersKeep policy, intent, and audit centralized while isolating provider-specific application logic behind adapters. A monolithic controller becomes hard to scale and couples unrelated provider failure domains.
  • Automatic repair versus approvalAutomate low-risk, reversible self healing infrastructure platform changes and require approval for destructive or high-blast-radius operations. Full automation without policy or blast-radius controls can turn a transient signal into a widespread outage.
Diagrammatic — system design practice and architecture review.