The Odds a UUID Ever Repeats
A version-4 UUID has 128 bits total, but not all of them are random: 6 are fixed by the format itself (4 bits mark it as version 4, 2 more mark the variant), leaving 122 genuinely random bits.[1] Birthday-paradox math — the same math behind "how many people need to be in a room before two share a birthday" — says you'd need to generate roughly 2.71 quintillion (2.71 × 1018) version-4 UUIDs before the odds of any two colliding crossed 50%. Every UUID this site has ever generated, and every one it ever will, is functionally guaranteed unique — not because the math forbids a collision, but because the number of UUIDs that would need to exist for one is bigger than most estimates of the total data ever stored by humanity.
Not all UUIDs are random
"UUID" names a format (128 bits, that familiar 8-4-4-4-12 hyphenated hex layout), not a single method for filling it. The version number — a single hex digit baked into the string itself — says how it was actually generated, and the versions in real use trade off very differently:
- v1 (timestamp + MAC address). Encodes the generating machine's network MAC address and a high-precision timestamp directly into the UUID. Guaranteed unique per machine as long as its clock doesn't move backward, but it leaks two things a random UUID doesn't: roughly when the record was created, and a hardware identifier for the machine that created it.
- v4 (random). What most tools generate by default today. No embedded information beyond "this was made to be a UUID."
- v5 (name-based, SHA-1). Deterministic: the same input name plus the same namespace always produces the same UUID. v3 is the same idea with the older, weaker MD5 instead of SHA-1, kept mainly for backward compatibility.
- v7 (Unix-timestamp-prefixed, random tail). A newer addition to the standard, formalized in RFC 9562 (2024), built specifically to fix a real operational problem with v4: database indexes built on a v4 primary key fragment badly over time because insert order has nothing to do with sort order, whereas a v7 UUID sorts roughly chronologically by creation time the way an auto-incrementing integer does.[2]
The alternative that isn't a UUID at all
Before v7 existed, teams that wanted both "sortable like a timestamp" and "collision-resistant like a UUID" often reached for a non-standard format instead — ULID (Universally Unique Lexicographically sortable ID) and Twitter's Snowflake ID are the two most common. Both encode a timestamp plus randomness or a machine/sequence counter into a fixed-width identifier, sort naturally, and predate v7's standardization.
What actually goes wrong in practice
Real UUID collisions in production systems are, unsurprisingly, essentially never caused by the random-number math failing — they're caused by something upstream of the math: a bad or predictable random number generator, a v1 UUID generated on a virtual machine that cloned its MAC address, or — far more common than either — a caching or retry bug that generates the same UUID twice by accident rather than the generator itself repeating. If two UUIDs in a real system ever collide, the math above is not where to start debugging.
The underlying idea — that a few dozen bits of real entropy is enough to make a collision practically impossible — is the same math Browser Fingerprinting runs in reverse: there, a much smaller entropy budget (tens of bits, not 122) is enough to make a browser practically identifiable instead.