Design a Personalized Recommendation Engine — System Design Interview Practice
Design a system that analyzes user behavior, trains ML models, generates recommendations, and serves predictions through an API. Work through the requirements, architecture trade-offs, and an interactive design review.
Concepts and architecture decisions to consider
- gcpConcept to explore
- recommendations aiConcept to explore
- machine learningConcept to explore
Interview prompt
Design behavior collection, feature computation, model training, candidate generation, and serving so users can request personalized recommendations reliably at scale.
- Define the source of truth for events, model versions, and feedback and make retries idempotent.
- Use bounded, partitioned state to meet 100M users and 1M recommendation requests per second and p95 <=100ms prediction.
- Separate the critical request path from feature pipelines, training, evaluation, and batch candidates.
- Explain consistency, failure recovery, authorization, observability, and a degraded mode.
Requirements and scale assumptions
- Support the core workflow to request personalized recommendations.
- Expose status, results, and freshness appropriate to behavior collection, feature computation, model training, candidate generation, and serving.
- Support authorization, validation, updates, deletion, and recovery semantics.
- Meet p95 <=100ms prediction under normal load.
- Scale to 100M users and 1M recommendation requests per second without a single hot key or unbounded synchronous work.
- Do not lose committed state; make retries and duplicate events safe.
- Degrade safely when downstream workers, caches, or external dependencies fail.
- 100M users and 1M recommendation requests per second
- Partition by the primary tenant, user, item, or geographic key and isolate hot partitions.
- Keep serving state bounded; retain raw events or durable records for replay and auditing.
- Peak scale: 100M users — Capacity assumption that drives partitioning and backpressure.
- Latency target: p95 <=100ms prediction — User-facing budget for the primary request or read path.
- Durable boundary: Committed before async — The source of truth is events, model versions, and feedback.
- Async boundary: At-least-once workers — Keep feature pipelines, training, evaluation, and batch candidates off the synchronous path.
Key entities
- DatasetVersiondatasetId, version, schemaHash, qualityStatus, lineage, createdAt
Immutable recommendation engine input version used for reproducible training, evaluation, or replay.
- FeatureSnapshotentityId, featureSetVersion, eventTime, values, sourceWatermarks
Point-in-time recommendation engine features with source watermarks so online and offline values can be compared.
- TrainingRunrunId, datasetVersion, codeVersion, metrics, artifactUri, status
Audited recommendation engine run that records data, code, dependency, and evaluation lineage.
- ModelVersionmodelId, version, stage, schema, qualityGates, endpoint
A promotable recommendation engine model version with rollout state, contract, and rollback metadata.
Data flow
- 1. Register and validate training dataThe recommendation engine gateway records an immutable dataset version, schema, lineage, quality status, and privacy disposition.
- 2. Build point-in-time featuresFeature workers join recommendation engine inputs using event-time watermarks, prevent leakage, and publish the same feature contract for training and serving.
- 3. Train and evaluate asynchronouslyThe orchestrator schedules recommendation engine runs with checkpointed artifacts, reproducible environments, and metrics tied to the exact input versions.
- 4. Gate and serve a model versionA registry compares recommendation engine quality, bias, safety, and compatibility gates before canary or production rollout with an immediate rollback pointer.
- 5. Monitor drift and learn from feedbackOnline inference records latency, errors, drift, and delayed labels so recommendation engine retraining is evidence-driven rather than triggered by guesswork.
Deep dives and trade-offs
- Reproducibility and leakage preventionPin recommendation engine data, feature, code, dependency, and model versions for every run. Use point-in-time joins and quarantine failed quality or privacy checks before training. Keep raw inputs and artifacts immutable so a result can be replayed after a dependency changes.
- Safe promotion and serving contractsSeparate recommendation engine model registration from deployment and require signed artifacts plus schema compatibility. Use shadow traffic, canaries, rollback pointers, and per-version latency/error budgets. Return model version and feature freshness so clients can explain or reproduce a prediction.
- Drift, feedback, and costMeasure feature drift, prediction drift, label delay, and segment-level quality for recommendation engine rather than only aggregate accuracy. Sample expensive inference and cap retraining concurrency with an explicit GPU or compute budget. Keep human corrections and delayed labels linked to the original prediction and model version.
- Batch versus online featuresPrefer a shared feature contract with batch backfills and a low-latency online serving path for decisions that need freshness. Two independently defined transformations create training-serving skew and hard-to-debug regressions.
- Synchronous versus asynchronous inferenceKeep interactive recommendation engine inference synchronous within a strict budget and queue large or expensive jobs. A request path that waits for model loading, enrichment, or retraining turns downstream slowness into an outage.
- Global model versus segment modelsStart with one versioned model and add segment-specific models only when quality or policy evidence justifies the operational cost. Many simultaneously active versions multiply monitoring, rollback, and data-lineage burden.