Low Level Design
How to Build a Scalable Music Service
Explore the system design of a music streaming service like Spotify. Learn about positional indexing for playlists, handling concurrency with the Actor model, low-latency playback architecture, and the design patterns that make it scale to millions of users.
Imagine you open Spotify, hit Play, lock your phone, switch to your laptop, and the song continues exactly from where you left off.
Simple for the user.
Insanely complex for the system.
A global music streaming service needs to handle millions of requests per second, maintain playlist ordering across devices, sync playback positions, serve audio with near-zero latency, and make sure your friend doesn’t mess up your collaborative playlist while you’re editing it.
So let’s walk through how to design Spotify—not as a boring system-design checklist, but as an engaging breakdown of why each architectural choice matters.
What Are We Building, Really?#
Spotify isn’t just “play music online.”
It’s:
- A search engine
- A recommendation engine
- A low-latency content delivery network
- A distributed state machine for playback
- A social platform for playlists
- A storage system for millions of songs
So before we architect anything, we need the core behaviors right.
Users can:#
- Search for songs
- Play, pause, skip, or seek
- Create playlists
- Reorder songs
- Share or collaborate on playlists
Artists can:#
- Upload albums and singles
The system must:#
- Personalize recommendations
- Sync playback across devices
- Serve audio instantly via CDN
- Scale to millions of concurrent listeners
Glorious.
Let’s build.
Step 1: The Data Model — Who Lives in Spotify’s World?#
If Spotify were a universe, its “cast” would be:
- Users — listeners (and sometimes playlist curators)
- Artists — the creators
- Albums — collections of songs
- Songs — the heart of the platform
- Playlists — ordered lists of songs, sometimes created together
Relationships are simple:
- One artist → many albums
- One album → many songs
- One playlist → many songs (no duplicates, ordered)
- One song → exactly one album
Think of playlists as a dynamic, shared document—like a Google Doc but for music ordering.
Step 2: The Architecture Behind the Screen#
Spotify’s architecture is built around one rule:
Keep the UI responsive, keep the backend scalable. Here’s the high-level flow:
Every layer exists for a reason.
The Facade#
A thin entry point—no business logic. It just routes requests.
The Service Layer#
This is the brain:
- playlists
- playback
- user auth
- recommendations
- search
Each service is independent, scalable, and focused.
Repositories#
Abstract data access so we can switch between:
- in-memory for testing
- SQL
- NoSQL
- distributed stores
Object Storage + CDN#
Audio files live in S3/GCS.
CDNs cache them in edge locations so playback starts instantly.
Step 3: How Search Really Works#
Search MUST feel magical.
You type “Tay” → you get “Taylor Swift”.
You type “blinding” → "Blinding Lights" pops up before you finish.
This is powered by Elasticsearch with:
- n-grams for partial matches
- fuzzy matching for typos
- synonyms for variations ("weeknd" ≈ "the weekend")
- autocomplete indexes
Search is one of the heaviest services, so it’s completely decoupled.
Step 4: The Hardest Problem — Playlist Ordering#
Most candidates say:
“I’ll store playlists as arrays or linked lists.”
That breaks instantly at scale.
Imagine a playlist with 10,000 songs.
Reordering the 50th song to the 5th position would force shifting thousands of rows.
Not acceptable.
The Real Trick: Positional Indexing#
Songs get assigned positions like: 100, 200, 300, 400,...
- If a user inserts a new song between 200 and 300, it gets: 250
- If a user inserts a new song between 250 and 300, it gets: 275 Only ONE row updates.
If the gaps get too tight, the system runs:
- Re-indexing (expensive but rare)
- Compacting (moving songs to fill gaps)
This is the same idea Notion, Figma, and Google Docs use for collaborative editing.
And yes—it scales beautifully.
Step 5: Playback — The Most Sensitive Flow#
Playback feels simple:
- You hit play
- Music starts
But here’s what REALLY happens:
1. UI sends play(songId) to Spotify#
2. Facade forwards to PlaybackService#
3. PlaybackService:#
- Validates auth
- Loads song metadata from Redis
- Creates/updates a playback session
- Returns a CDN URL
4. UI streams the audio directly from the CDN#
The server never streams audio.
It only coordinates.
Step 6: Handling Chaos — Concurrency#
Here’s a real problem:
What if a user taps: Play → Pause → Seek → Play → Skip within 200ms?
If you run these on shared threads, playback state becomes corrupted.
Spotify solves this using:
Single-Thread Executors (Per User)#
Each user gets their own miniature actor:
- All playback commands execute sequentially
- No races
- No overlapping seeks or plays
It guarantees ordering, even under rapid taps.
Step 7: Collaborative Playlists — Conflicts Everywhere#
Two users editing the same playlist at the same time will ALWAYS cause conflicts.
Spotify solves this with:
Optimistic Concurrency#
- Try update
- If version mismatch → retry
Atomic Updates via Redis Lua#
- Multi-step operations performed as one unit
Last-write-wins (timestamp)#
Good enough for social playlists.
Step 8: Caching — The Secret Behind Instant Playback#
Spotify feels fast because of a layered caching strategy:
| Layer | What It Stores | Why |
|---|---|---|
| CDN | audio chunks | microsecond access |
| Redis | metadata (song, playlist) | low latency |
| App Cache | trending songs | instant recommendations |
| DB | truth | durability |
Metadata is read thousands of times more often than it's written—perfect for Redis.
Step 9: Patterns Used Throughout Spotify#
Command Pattern#
Play, Pause, Seek are commands executed in order.
Strategy Pattern#
Different recommendation algorithms can be swapped dynamically.
Facade Pattern#
The controller layer exposes simple APIs.
Repository Pattern#
Clean persistence separation.
Singleton#
SpotifyController instance.
Actor Model#
Per-user executors = serialized playback actions.
Read-Write Locks#
Thousands of users can read a playlist at the same time while only one can modify.
Final Thoughts#
Spotify is a masterclass in distributed system design.
A simple "play song" action involves:
- CDN fetches
- Redis hits
- playback state synchronization
- metadata caching
- device awareness
- cross-platform consistency
Designing something like Spotify in an interview is NOT about drawing a pretty diagram.
It’s about understanding the trade-offs, concurrency risks, data modeling challenges, and user experience expectations.
If you can explain the reasoning above, you’re already ahead of most candidates interviewing for senior backend roles.
Every layer exists for a reason.