High Level Design

Design a Video Streaming Service

End-to-end design of a YouTube-like video platform — uploading flow, adaptive streaming, DAG-based transcoding pipeline, CDN cost optimizations, pre-signed URLs, DRM, and fault-tolerant error handling.

August 21, 2026·Updated September 17, 2026

A video streaming service (YouTube, Netflix, Vimeo, Udemy) lets users upload, transcode, and watch videos across devices, at massive scale. The scope is wide — this design focuses on uploading and streaming, the two core workflows.

Before jumping to design, always clarify scope with the interviewer.


Step 1: Requirements#

Clarifying Questions#

QuestionAnswer
Features needed?Upload a video and watch a video
Clients?Mobile apps, web browsers, smart TV
Daily active users?5 million
Daily uploads10% of users
Average daily time spent?30 minutes
International users?Yes, large percentage
Supported video resolutions?Most formats and resolutions
Encryption required?Yes
Max video file size?1 GB
Use existing cloud infrastructure?Yes — strongly recommended

Functional Requirements#

#Requirement
1Users can upload videos (max 1 GB) in most formats and resolutions
2Users can stream/watch videos smoothly across devices
3Users can change video quality during playback
4System supports mobile apps, web browsers, and smart TVs
5Videos are transcoded into multiple formats and resolutions after upload
6International users are supported

Non-Functional Requirements#

#RequirementDetail
1High availabilityVideo streaming must stay available even during partial failures
2ScalabilityHandle 5M DAU; 500K video uploads/day; 150 TB new storage/day
3Low latencySmooth playback with minimal buffering; fast upload starts
4ReliabilityUploads must be resumable; no data loss on failures
5SecurityVideos encrypted; only authorized uploads via pre-signed URLs; DRM for piracy protection
6Low infrastructure costCDN costs ~$150K/day at scale — cost optimization is a first-class concern
7Adaptive streamingServe high resolution on good bandwidth; auto-downgrade on poor connections

Estimations#

MetricCalculationResult
Daily active users—5 million
Daily video uploads5M × 10%500,000 videos/day
Daily storage needed500K × 300 MB150 TB/day
Videos watched per user/day—5 videos
CDN data served/day5M × 5 × 0.3 GB7.5 PB
CDN cost/day (AWS CloudFront @ $0.02/GB)7.5 PB × $0.02~$150,000/day

CDN serving costs are extremely high. A key design goal is to reduce CDN usage without sacrificing performance.


Step 2: APIs#

Video Upload#

StepMethodEndpointDescription
1POST/v1/videosInitiate upload — creates the video record, returns video_id and a pre-signed URL
2PUT{pre_signed_url}Upload the video file directly to blob storage using the pre-signed URL (goes to S3/blob, not the API server)
3PUT/v1/videos/{video_id}/metadataSend title, description, tags, thumbnail — runs in parallel with step 2
4GET/v1/videos/{video_id}/statusPoll transcoding status — uploading / processing / ready / failed

Steps 2 and 3 run in parallel — the client uploads the file and sends metadata simultaneously. Step 2 bypasses the API server entirely; large binary data goes straight to blob storage via the pre-signed URL.


Video Streaming#

MethodEndpointDescription
GET/v1/videos/{video_id}Fetch video metadata — title, description, available resolutions, thumbnail URL
GET/v1/videos/{video_id}/streamReturns the ABR manifest URL (HLS .m3u8 or DASH .mpd)

What /stream actually does: This call goes to the API server, which checks the video's metadata (is it ready? what resolutions exist?) and returns the ABR manifest URL pointing at the CDN. The API server never talks to the CDN — it just tells the client where to go.

What happens next (separate from this API): The client's video player fetches the manifest — a small text file listing all quality variants and the CDN locations of their segments. The player requests video segments directly from CDN edge servers, measuring bandwidth and switching quality chunk by chunk. This is where the CDN cache hit/miss check actually happens — entirely outside the API server.

Interview phrasing: "The /stream endpoint returns the ABR manifest URL. The API server never touches the CDN on the read path — it just tells the client where to go. All subsequent video traffic flows player → CDN, bypassing your infrastructure entirely."


Step 3: High Level Design#

Three main components:

ComponentRole
ClientMobile apps, web browsers, smart TV apps
CDNCaches and serves videos close to users globally
API ServersEverything else — feed, recommendations, user management, upload management

High-level components

Two main workflows: video uploading and video streaming.


1. Video Uploading Flow#

ComponentResponsibility
Load BalancerEvenly distributes requests among API servers
API ServersHandles all user requests except video streaming
Metadata DBStores video metadata; sharded and replicated for performance and HA
Metadata CacheCaches video metadata and user objects for faster reads
Original StorageBlob storage for raw, unprocessed video uploads
Transcoding ServersConvert video into multiple formats/bitrates for device compatibility
Transcoded StorageBlob storage for processed video files
CDNCaches and serves transcoded videos globally
Completion QueueMessage queue for transcoding completion events
Completion HandlerWorkers that consume completion events and update metadata

Video uploading flow

The upload happens in two parallel steps:

A. Upload Actual Video#

  1. Video is uploaded to Original Storage.
  2. Transcoding Servers fetch the video and begin transcoding.
  3. On completion (parallel):
    • Transcoded video pushed to Transcoded Storage → distributed to CDN.
    • transcoding completion event queued in the Completion Queue.
    • Completion Handler workers pull events and update Metadata DB and Cache.
  4. API Server notifies the client that the video is ready for streaming.

Upload video flow

B. Update Video Metadata (Parallel with A)#

While the video file uploads, the client simultaneously sends metadata to the API servers. API servers update the Metadata Cache and Metadata DB.

Update metadata flow


2. Video Streaming Flow#

Downloading = the full video file is downloaded before playback starts. Streaming = playback begins while the file is still downloading — only a small buffer is needed.

  • Videos stream directly from the CDN to clients.
  • The edge server closest to the user serves the video.
  • Cache miss → fetched from Transcoded Storage (origin), cached at edge for future requests.

Streaming Protocols#

ProtocolLatencyCompatibilityBest For
HLSMedium–High (low with LL-HLS ~2s)Excellent (iOS, browsers, TVs)Live & VOD — max device reach
MPEG-DASHMedium–HighExcellent (except iOS)Codec flexibility, adaptive bitrate
WebRTCUltra-low (<0.5s)Excellent (browsers, mobile)Interactive live video, calls
RTMPLow (~2–5s)Strong ingestLive stream ingest to platforms
SRTVery low (~1–2s)GrowingSecure contribution over unreliable networks
RTSPLow (~2–10s)Good (cameras, devices)Surveillance / closed-network streaming

Video streaming flow


Step 4: Detailed Design#

Video Transcoding#

Video transcoding converts a video into multiple formats and bitrates for different devices and network conditions.

Why transcoding is needed:

ReasonExplanation
Storage optimizationRaw videos are huge — transcoding compresses them
Device compatibilityDifferent devices support different formats
Adaptive quality deliveryServe HD to fast networks, lower res to slow networks
Smooth playbackEnable auto/manual quality switching on variable networks

Common formats and codecs:

TypeOptions
Video containersMP4, AVI, MOV, MKV, WebM
Video codecsH.264, H.265 (HEVC), VP9, AV1

DAG-Based Transcoding Pipeline#

Challenges:

  • Transcoding is CPU/GPU intensive and time-consuming.
  • Different creators need different processing (watermarks, thumbnails, HD/SD variants).
  • A single fixed pipeline can't cover all use cases.

Solution: Model transcoding tasks as a Directed Acyclic Graph (DAG) — tasks run sequentially where required and in parallel where possible.

DAG pipeline

Fetch original video
        │
Split into [video stream, audio stream, metadata]
        │
    ┌───┴──────────────────────┐
    ▼                          ▼
Inspection                Audio encoding
    │                          │
    ▼                          │
┌───┼────┬────┬────┐           │
▼   ▼    ▼    ▼    ▼           │
360 480 720 1080 Thumbnail     │
 │   │    │   │                │
 ▼   ▼    ▼   ▼                │
Watermark (on each resolution) │
    │                          │
    └──────────┬───────────────┘
               ▼
        Mux audio + video → final output per resolution

What this structure buys you concretely:

  • Inspection gates everything — if the file is corrupt, the DAG fails fast before wasting CPU on 4 parallel encodes.
  • Encoding and thumbnail generation fan out in parallel from a single inspected source — no artificial serialization between them.
  • Audio encoding runs completely independently, in parallel with the entire video branch. It only needs to rejoin at the very end (muxing).
  • Watermarking is per-resolution — 4 independent parallel tasks, not one sequential bottleneck applied to each resolution one at a time.
  • The final mux is a join point — it waits for both the watermarked video variant AND the encoded audio before it can run. The DAG enforces this dependency explicitly.

DAG stages:

StagePurpose
InspectionValidate video quality, detect corruption or malformed files
Video EncodingTranscode to multiple resolutions: 360p, 480p, 720p, 1080p, 4K
ThumbnailsUse uploader's image or auto-generate from video frames
WatermarkingOverlay branding or ownership identifiers per resolution

Transcoding Architecture#

Transcoding architecture

1. Preprocessor

  • Splits video into GOPs (Groups of Pictures) — independently playable chunks (a few seconds each).
  • Handles backward compatibility for older devices via GOP-based splitting.
  • Generates the DAG from client-defined configuration files.
  • Caches GOPs and metadata for retry/recovery if encoding fails.

2. DAG Scheduler

  • Divides the DAG into execution stages based on task dependencies.
  • Places each stage's tasks into the Resource Manager's task queue.

Example stages:

  • Stage 1: Split original video into video stream, audio stream, metadata.
  • Stage 2 (parallel): Video → encoding + thumbnail | Audio → encoding

DAG scheduler

3. Resource Manager

  • Manages a task queue (priority queue of pending tasks), worker queue (worker availability), and running queue (active task–worker bindings).
  • Workflow: fetch highest-priority task → pick optimal worker → assign → track → remove on completion.

Resource manager

4. Task Workers

  • Execute assigned tasks (encoding, thumbnail generation, etc.).
  • Report completion status back to the Resource Manager.

5. Temporary Storage

Temporary storage is the scratch space between pipeline stages — it holds intermediate data produced during processing, not the original upload and not the final output.

WhatWhereWhy
GOP chunks (video split into segments by Preprocessor)Blob storageLets encoding workers each grab a chunk and encode in parallel instead of one worker processing the whole file
Split streams (video, audio, metadata separated)Blob storageSpecialist workers can process each stream simultaneously
DAG task state and metadataIn-memoryTiny, frequently read by workers — memory keeps access sub-millisecond
Active encoding buffersLocal worker diskHolds intermediate output while a worker is mid-task

Everything here is auto-cleaned once the pipeline completes. Encoded output goes to Transcoded Storage; the original upload stays untouched in Original Storage.


Step 5: System Optimizations#

Speed Optimizations#

1. Parallelize video uploads via GOP-aligned chunking

  • Instead of uploading the full video at once, split it into small GOP-aligned chunks.
  • Enables resumable uploads on failure and faster parallel uploads.
  • GOP splitting can be done client-side to reduce server load.

Speed optimization - chunked upload

2. Upload centers close to users

  • Deploy multiple global upload centers.
  • Users upload to the nearest region (US → NA servers, China → Asia servers).
  • Implemented using CDNs as upload endpoints.

3. Parallelism everywhere with message queues

  • Without queues: encoding waits for full download to finish (sequential bottleneck).
  • With queues: encoding starts as chunks arrive (parallel throughput).

Speed optimization - message queues


Safety Optimizations#

1. Pre-signed upload URLs

Flow:

  1. Client requests a pre-signed URL from the API server.
  2. API server returns a time-bound, permission-scoped URL.
  3. Client uploads video directly using that URL — bypasses the API server for large data.

Safety optimization - pre-signed URLs

2. Protecting videos from piracy

MechanismHow it works
DRM (FairPlay, Widevine, PlayReady)Enforces playback restrictions at the device level
AES encryptionEncrypts stored videos; decrypts only for authorized viewers
Visual watermarkingOverlays logos or identifiers to deter theft and trace leaks

Cost-Saving Optimizations#

Core observation: Video traffic follows a long-tail distribution — a small number of popular videos account for most views; the majority have low or zero traffic.

StrategyDescription
Selective CDN usageServe only popular videos via CDN; deliver long-tail videos from cheaper origin servers
On-demand encodingSkip pre-encoding all variants for low-traffic videos; encode on demand instead — trade-off: slower first-viewer experience
Region-aware distributionPush content only to the regions where it's popular, not globally
Build or partner for CDNAt Netflix-scale, build your own CDN or partner with ISPs (Comcast, AT&T, Verizon) to cut bandwidth costs

Step 6: Error Handling#

Types of Errors#

TypeExampleStrategy
RecoverableVideo segment fails to transcodeRetry a few times; return error code if retries exhaust
Non-recoverableMalformed video formatStop all tasks for that video; return error to client

Component-Level Error Playbook#

ComponentFailure ScenarioHandling Strategy
UploadUpload failsRetry upload a few times
Video splittingClient can't split video by GOPUpload full video; perform GOP splitting server-side
TranscodingVideo/segment fails to transcodeRetry transcoding task
PreprocessorPreprocessor failureRegenerate DAG
DAG SchedulerTask scheduling failureReschedule the task
Resource Manager QueueQueue service downSwitch to replica queue
Task WorkerWorker crashesRetry task on another worker
API ServerServer downRoute request to another stateless API server
Metadata CacheCache node failureRead from replicas; spin up replacement node
Metadata DB – MasterMaster node downPromote a replica to master
Metadata DB – SlaveSlave node downRead from another slave; add new slave node

Step 7: DB Choices#

ComponentDatabaseWhy
Video metadata (title, description, tags, status)MySQL / PostgreSQL (sharded)Structured data; sharded + replicated for performance and HA at scale
User data (profiles, subscriptions)MySQL / PostgreSQLRelational; standard indexed lookups
Metadata cacheRedisFast reads for video info on every page/player load; avoids DB round-trips
Original video storage (raw uploads)S3 / Blob StorageLarge binary files; supports direct upload via pre-signed URLs — no API server in the path
Transcoded video storage (processed variants)S3 / Blob StorageStores multiple encoded versions (360p, 480p, 720p, 1080p) per video; origin for CDN
CDNCloudFront / AkamaiEdge caching; serves transcoded video directly to clients globally
Completion queue (transcoding events)KafkaDurable, high-throughput event stream; decouples transcoding servers from completion handlers
Temporary transcoding storage (GOPs, intermediate chunks)In-memory / local diskShort-lived; auto-cleaned after processing completes

Why blob storage for both video tiers? Binary files don't belong in a database. S3 is cheap, durable, scales infinitely, and integrates natively with CDNs — purpose-built for exactly this.

Why sharded MySQL? Video metadata at scale (800M+ videos on YouTube) can't fit on a single node. Sharding by video_id distributes load while keeping structured query support.

Why Redis over DB for metadata? Every video play starts with a metadata fetch. At 5M DAU watching 5 videos/day, that's 25M reads/day on metadata alone — Redis keeps this sub-millisecond.


FAQs#

If watermarking fails after exhausting retries, should the whole upload be marked as failed?

It depends on why the watermark exists and who owns the content.

For most videos, watermarking is cosmetic branding — an overlay logo or channel name. Shipping without it is a reasonable degradation: mark the video as published with a warning, log the failure, and let the creator re-trigger watermarking later.

But watermarking also serves as a piracy/anti-theft protection mechanism — for premium or licensed content, it makes a leaked copy traceable. For those videos, shipping without a watermark is not safe; a leak becomes untraceable, which may violate licensing agreements or expose the platform to legal liability.

The right model is to treat watermarking as configurable criticality at the video or creator level:

Content typeWatermark criticalityOn failure
Regular user-generated contentNon-criticalPublish without watermark; flag for retry
Premium / licensed contentCriticalBlock publish; surface hard failure to creator
Paid course or DRM-protected contentCriticalBlock publish; do not stream until resolved

The DAG node for watermarking can be tagged with a required: true/false flag per job — the completion handler respects that when deciding whether a partial success counts as a publishable state.


Why a DAG for the transcoding pipeline rather than a flat list of parallel tasks?

A flat parallel task list assumes all tasks are independent — fire everything at once and collect results. A DAG gives you explicit dependencies: some tasks can only start after specific others finish.

For video transcoding, the tasks aren't all independent:

  • You can't encode video until inspection confirms the file isn't corrupted.
  • You can't generate a thumbnail until at least one video frame exists.
  • Audio encoding can only run in parallel with video encoding after the audio stream has been split out from the raw file.
  • Watermarking must happen after encoding, not before.

With a flat parallel list, you'd either artificially sequence everything (losing parallelism) or fire dependent tasks prematurely (causing failures). The DAG gives you both at once: maximum parallelism within each stage, strict ordering across stages.

The DAG also enables granular retries — if thumbnail generation fails, re-run only that node, not the entire pipeline. A flat task list has no way to express that.


How does adaptive bitrate streaming actually work?

The manifest (.m3u8 / .mpd) lists multiple bitrate variants of the same video — rungs on a quality ladder (360p, 480p, 720p, 1080p). Each variant is broken into small chunks, typically a few seconds each.

The player doesn't fetch a whole video file — it fetches one chunk at a time. Before requesting the next chunk, it checks buffer health and current download speed, then picks which rung to use for that chunk. The switch happens chunk by chunk, seamlessly, with no server-side decision-making. The server just serves whichever chunk the player asks for.

This is why the manifest approach makes ABR possible: all the intelligence lives in the client player. The CDN is just a dumb file server — it doesn't know or care which quality the player is on.


If not all videos are cached on CDN, where does the video come from?

The serving path has three tiers depending on the video's popularity:

PathLatencyCostUsed for
CDN edge hitLowestHighestPopular videos already cached
CDN miss → Transcoded StorageMediumMediumLess popular videos; CDN fetches from origin, caches, then serves
On-demand transcode → Transcoded StorageHighestLowestRarely accessed resolutions not yet encoded

The Original Storage (raw uploads) is never in the serving path — it only exists for the transcoding pipeline. "Not caching on CDN" doesn't mean a video is unavailable — it means the request goes directly to Transcoded Storage via an origin server instead of a CDN edge.


Why use cloud services instead of building from scratch?

System design interviews focus on architecture, not low-level implementation. Building from scratch is unrealistic — Netflix uses AWS, Facebook uses Akamai CDN, YouTube uses Google Cloud Storage and Google CDN.


What is blob storage?

A Binary Large Object (BLOB) store holds binary data as a single entity. Examples: AWS S3, Azure Blob Storage, Google Cloud Storage.


How to handle live streaming?

Live streaming shares similarities with VOD (upload → encode → stream) but has key differences:

ConcernVODLive
LatencyFlexibleStrict — specialized protocols (WebRTC, LL-HLS) required
ParallelismHigh — batch encodingLow — small real-time chunks
Error handlingLong retries acceptableMust be lightweight; graceful degradation over recovery

How to handle video takedowns?

Videos violating copyright, policy, or law are removed. Detection happens:

  • At upload — automated scanning during processing.
  • Post-upload — through user flagging and moderation pipelines.

Glossary#

Bitrate — the rate at which bits are processed per unit of time. Higher bitrate = better quality but more bandwidth required.

GOP (Group of Pictures) — a sequence of video frames that can be decoded independently. Splitting videos into GOPs enables parallel encoding and resumable uploads.

Adaptive Bitrate Streaming (ABR) — automatically adjusts video quality based on available network bandwidth. Protocols like HLS and MPEG-DASH implement ABR.