High Level Design
Designing a Unique ID Generator
Designing a distributed unique ID generator — comparing multi-master replication, UUID, ticket servers, and Twitter Snowflake, with a deep dive into the 64-bit Snowflake layout and follow-up considerations.
A unique ID generator creates globally unique identifiers across distributed systems. Properties vary by design: globally unique (no collisions), K-sortable (IDs roughly increase with time), compact, stateless or stateful.
A naive approach is auto-increment in a central database — but that creates a bottleneck and a single point of failure.
Step 1: Requirements#
| Question | Answer |
|---|---|
| Globally unique? | Yes |
| K-sortable (ordered by date)? | Yes — ordered by timestamp, not by increment of 1 |
| ID length | 64-bit integer |
| Numeric only? | Yes |
| Scale | Up to 10,000 ID requests per second |
| Latency | < 1ms per ID generation |
| Availability | Highly available and fault-tolerant |
Challenges in distributed ID generation:
- Collisions — multiple nodes generating IDs independently may overlap.
- Scalability — must support billions of IDs efficiently.
- Ordering — some systems require time-ordered IDs (e.g., Kafka offsets, Twitter posts).
- Availability — must not depend on a single machine or DB auto-increment.
Step 2: Options Comparison#
| Option | Description | Pros | Cons |
|---|---|---|---|
| Multi-Master Replication | Multiple masters generate IDs from guaranteed unique ranges | Scales with number of nodes | Not time-ordered; hard to scale across data centers |
| UUID | 128-bit identifier, no central coordination | Simple; no SPOF; scales with servers | Not K-sortable; 128 bits (large); not numeric |
| Ticket Server | Centralised server generates IDs; clients request them | Numeric; K-sortable; works at small–medium scale | Single point of failure; scalability issues under high load |
| Snowflake (Twitter) | 64-bit ID from timestamp + datacenter ID + machine ID + sequence | Numeric; K-sortable; highly scalable; low latency | Requires unique machine ID configuration; more complex |
Visual comparison of approaches:
Multi-master replication:

UUID:

Ticket server:

Snowflake:

Twitter Snowflake Design#
Snowflake is a 64-bit unique ID algorithm. It provides:
- Global uniqueness
- Time ordering
- High throughput (millions of IDs/sec)
64-bit ID Layout#
| Bits | Field | Notes |
|---|---|---|
| 1 | Sign bit | Always 0 — reserved for future use |
| 41 | Timestamp | Milliseconds since custom epoch — gives ~69 years of IDs |
| 5 | Datacenter ID | Supports up to 32 datacenters |
| 5 | Machine ID | Supports up to 32 machines per datacenter |
| 12 | Sequence Number | 4,096 IDs per machine per millisecond |
How it works#
- Datacenter ID and Machine ID are configured at startup (env vars / config files) and fixed for the lifetime of the instance.
- Timestamp — the part that makes IDs time-sortable. Generated at runtime.
- Sequence Number — incremented by 1 per ID; resets to 0 each millisecond. If the counter exceeds 4,095 within the same millisecond, the generator waits until the next millisecond.
- If a machine restarts, it must use the same Machine ID to avoid collisions.
Throughput: up to 4,096 IDs per machine per millisecond = ~4M IDs/sec per machine. 32 machines per datacenter × 32 datacenters = 1,024 machines globally.
Follow-Ups#
1. Clock Synchronization#
Snowflake assumes all servers have the same clock. With multiple cores or machines, clock drift is real.
Solution: Use NTP for clock synchronisation. If a machine's clock drifts backward, it must wait until the clock catches up before generating new IDs.
2. Section Length Tuning#
| Application type | Trade-off | Approach |
|---|---|---|
| Low-concurrency, long-lived | Don't need millions of IDs/ms | Fewer sequence bits, more timestamp bits → IDs valid for centuries |
| High-concurrency (Twitter, TikTok) | Need thousands of IDs/ms | More sequence bits (12–14) → shorter lifespan (~50–70 years, acceptable) |
3. Machine Failures#
- On restart, the machine must use the same Machine ID.
- If permanently decommissioned, the Machine ID can be reused after a safe period (after the maximum timestamp duration has elapsed).
4. High Availability#
Since ID generation is mission-critical, deploy multiple instances behind a load balancer.
5. Choosing the Right ID Structure#
| Criteria | UUID v4 | Hash (MD5/SHA-1) | Snowflake | ULID | Sharded DB Auto-Increment |
|---|---|---|---|---|---|
| Uniqueness | ✅ Very high (128-bit) | ✅ Deterministic | ✅ Guaranteed | ✅ Very high | ✅ Within shard |
| Time ordering | ❌ None | ❌ None | ✅ Strong | ✅ Strong | ✅ Weak (shard-based) |
| Size | 128-bit | 128–160-bit | 64-bit ✅ | 128-bit | 64-bit ✅ |
| Compact | ❌ Long string | ❌ Long string | ✅ int64 | ❌ 26 chars | ✅ int64 |
| Coordination | None | None | Unique machine IDs | None | Shard offsets |
| Best Use Cases | Session IDs, cache keys, security tokens | File deduplication, content IDs | Tweet IDs, orders, events, logs | DB keys, distributed logs | Legacy/small-scale sharding |