Diagrammatic

Distributed File Storage — System Design Interview Practice

Design a distributed file storage system like Google Drive or Dropbox for storing and sharing files. Work through the requirements, architecture trade-offs, and an interactive design review.

Concepts and architecture decisions to consider

  • file storageConcept to explore
  • synchronizationConcept to explore
  • versioningConcept to explore

Interview prompt

Design upload, version, sync, sharing, and durable file storage so users can upload and share a file reliably at scale.

  • Define the source of truth for object versions and permissions and make retries idempotent.
  • Use bounded, partitioned state to meet 100M users and 1PB of new objects per day and metadata p95 <=150ms.
  • Separate the critical request path from chunk assembly, virus scanning, indexing, and replication.
  • Explain consistency, failure recovery, authorization, observability, and a degraded mode.

Requirements and scale assumptions

  • Support the core workflow to upload and share a file.
  • Expose status, results, and freshness appropriate to upload, version, sync, sharing, and durable file storage.
  • Support authorization, validation, updates, deletion, and recovery semantics.
  • Meet metadata p95 <=150ms under normal load.
  • Scale to 100M users and 1PB of new objects per day without a single hot key or unbounded synchronous work.
  • Do not lose committed state; make retries and duplicate events safe.
  • Degrade safely when downstream workers, caches, or external dependencies fail.
  • 100M users and 1PB of new objects per day
  • Partition by the primary tenant, user, item, or geographic key and isolate hot partitions.
  • Keep serving state bounded; retain raw events or durable records for replay and auditing.
  • Peak scale: 100M users — Capacity assumption that drives partitioning and backpressure.
  • Latency target: metadata p95 <=150ms — User-facing budget for the primary request or read path.
  • Durable boundary: Committed before async — The source of truth is object versions and permissions.
  • Async boundary: At-least-once workers — Keep chunk assembly, virus scanning, indexing, and replication off the synchronous path.

Key entities

  • ResourceSpecresourceId, tenantId, desiredState, version, policyVersion, updatedAt

    Versioned desired state for a distributed file storage managed resource.

  • OperationoperationId, resourceId, requestHash, step, attempt, status

    Durable distributed file storage reconciliation operation with per-step progress.

  • PolicyVersionpolicyId, scope, version, rules, effectiveAt, status

    Auditable distributed file storage policy evaluated before provisioning or mutation.

  • ReconciliationCheckpointresourceId, provider, observedVersion, cursor, lastError, updatedAt

    Provider-specific distributed file storage observation and recovery cursor.

Data flow

  1. 1. Accept a desired-state commandThe distributed file storage control plane authenticates the tenant, validates policy and quotas, checks the expected version, and records the desired state.
  2. 2. Plan a safe operationA planner turns distributed file storage desired state into ordered, bounded steps with dependency checks, blast-radius limits, and rollback metadata.
  3. 3. Reconcile providers asynchronouslyWorkers apply distributed file storage operations through provider adapters, persist checkpoints, rate-limit calls, and treat unknown outcomes as observable state.
  4. 4. Publish observed healthThe serving projection joins desired and observed distributed file storage state with operation status, policy version, freshness, and actionable errors.
  5. 5. Recover and auditRetries, dead letters, drift detection, and operator approvals repair distributed file storage resources without losing the original command or provider evidence.

Deep dives and trade-offs

  • Desired versus observed stateKeep distributed file storage desired state separate from provider-observed state and show both to operators. Make every reconciliation step conditional and resumable so a worker crash does not restart unsafe effects. Version policy and resource state so old operations cannot overwrite newer intent.
  • Provider failures and unknown outcomesUse provider-specific idempotency tokens and query-after-timeout behavior for distributed file storage operations. Bound retries with exponential backoff, circuit breakers, and per-provider quotas. Route irreconcilable drift to an approval or quarantine path instead of retrying forever.
  • Blast radius and operationsPartition distributed file storage work by tenant, region, cluster, or resource class and cap concurrent mutations. Audit who changed desired state, which policy allowed it, and what provider evidence was observed. Alert on drift age, operation backlog, failed steps, policy denials, and stale observations.
  • Push versus pull reconciliationUse event triggers for fast response and periodic scans for missed events, drift, and recovery. A push-only distributed file storage controller silently misses changes when a provider event is lost.
  • Central control plane versus provider-native controllersKeep policy, intent, and audit centralized while isolating provider-specific application logic behind adapters. A monolithic controller becomes hard to scale and couples unrelated provider failure domains.
  • Automatic repair versus approvalAutomate low-risk, reversible distributed file storage changes and require approval for destructive or high-blast-radius operations. Full automation without policy or blast-radius controls can turn a transient signal into a widespread outage.
Diagrammatic — system design practice and architecture review.