Content-Addressable Storage

9.Content-Addressable Storage

M

In this chapter

We'll meet content-addressable storage — a real addressing model where an object's key is a deterministic hash of its own actual content, meaning identical content always resolves to the identical key — and see exactly how it would have quietly prevented GreenMart's own nine-duplicate-photo discovery from ever costing extra storage at all.

7–9 min

The Problem in Real Life

Sarah runs a real audit on the product-image bucket and finds something odd: the same exact photo, byte for byte, uploaded under nine different keys — once for every color variant of a product that, it turns out, all shipped with an identical stock photo.

Mike winces. "So we're paying to store the same image nine times over?"

M

If it's the exact same file, why are we storing nine separate copies of it?

Mike

A Key GreenMart Chooses vs. A Key the Content Itself Determines

A key derived from content, not chosen

A deterministic hash of the object's actual bytes becomes its address — identical content always produces the identical key.

Deduplication falls out naturally

Since identical content shares one key, the object store only ever needs to keep one real physical copy, with no separate dedup step.

Content-Addressable Storage

Content-addressable storage is a real, genuinely different addressing idea: instead of a key GreenMart chooses freely (like products/electronics/headphones-402.jpg), the key is a real, deterministic hash of the object's own actual content — meaning identical content always produces the exact same real key, no matter how many times, or where, it's uploaded.

  • How the real address gets derived. A cryptographic hash function takes an object's actual bytes and produces a real, fixed-length string — a real fingerprint of that exact content, similar in spirit to the ETag idea from Act 4, but now used as the object's actual address, not just a freshness check. The same real content, hashed the same way, always produces the identical string; even one changed byte produces a completely different one.
  • The real, natural deduplication this creates. If GreenMart's own nine identical product photos were stored content-addressably, all nine uploads would resolve to the exact same real hash-derived key — meaning the object store would only ever need to keep one real physical copy, no matter how many times GreenMart's own team (unknowingly) uploads the identical file. Deduplication isn't a separate real feature bolted on; it falls out naturally from how the addressing itself works.
  • A real, honest trade-off: GreenMart gives up choosing the key. Content-addressable storage means GreenMart can no longer pick a real, human-meaningful key like headphones-402.jpg — the address is whatever the content's own hash happens to be, an unreadable string with no inherent connection to what the object actually is. Real systems that use this approach typically keep a separate, real mapping — a human-meaningful name pointing at the current content hash — layered on top, rather than replacing meaningful keys outright.
  • Where this real idea is genuinely, already familiar. GreenMart has, in fact, already met exactly this idea once before in this course: Act 3's own inode model separates a file's name from its actual content in a related but distinct way. Content-addressable storage takes that separation one real step further — the "address" itself becomes a direct, deterministic function of the content, not an arbitrary number or a freely chosen name at all.

GreenMart now has the real, precise fix for its own nine-duplicate-photo discovery: content-addressable storage doesn't require GreenMart to hunt down and manually deduplicate anything — identical content simply, automatically, only ever occupies real storage once, by the very nature of how its address is computed.

Key Takeaway

Content-addressable storage derives an object's key directly, deterministically, from its own actual content — identical content always produces the identical key, which means genuine, automatic deduplication falls naturally out of the addressing scheme itself, at the real, honest cost of giving up a human-meaningful, freely-chosen key.

Why This Matters

As GreenMart's own catalog grows — more products, more variants, more genuinely repeated assets — content-addressable storage is the real, structural fix for exactly the kind of silent, costly duplication this chapter's own audit just uncovered, without requiring anyone to manually hunt for it.

GreenMart now has a real, second addressing model — content-derived, not freely chosen — and understands exactly how it naturally eliminates duplicate storage. The next chapter shifts from a single copy's own address to a genuinely different real concern: how many real copies of GreenMart's data actually exist, and how durable that makes it.

Next