JSON vs. Binary Formats

6.JSON vs. Binary Formats

M

In this chapter

We'll compare JSON and binary serialization formats — JSON's self-describing, human-readable repetition versus a binary format's compact, schema-aware efficiency — and find the real, correct fix for GreenMart's own growing storage bill: its high-volume, internally-consumed event log is a strong fit for a binary format instead.

7–9 min

The Problem in Real Life

The event log is thriving on Cassandra, but GreenMart's own storage bill for it keeps climbing. Sarah checks the real, actual event records — each one is JSON, readable, and honestly a little wasteful. "Every field name gets spelled out, in full, on every single event," she says. "Millions of times a day."

Mike looks at one real record: {"productId": "402", "quantity": 3, "warehouseId": "WH-12"}. "That's... a lot of repeated text for three real numbers and an ID."

M

Why does every single event repeat the same field names, over and over, millions of times?

Mike

Readable and Repetitive vs. Compact and Schema-Aware

JSON — self-describing, genuinely repetitive

Every record spells out its own field names, which is readable but costs real, repeated bytes at scale.

Binary — compact, but needs a shared schema

A pre-agreed schema removes the need to repeat field names, producing smaller records at the cost of human readability.

JSON vs. Binary Formats

JSON and binary serialization formats represent the same real trade-off from opposite directions: JSON prioritizes being genuinely human-readable and self-describing; binary formats prioritize being genuinely compact and fast to process, at the real cost of no longer being readable by eye at all.

  • Why JSON is genuinely repetitive, by design. JSON is self-describing — every single record spells out its own real field names ("productId", "quantity") alongside its values, which is exactly why it's genuinely human-readable without any outside reference. The real cost: those field names get repeated, in full, in every single record, even across millions of genuinely identical-shaped events — real, wasted bytes, purely for that self-describing convenience.
  • How a binary format removes that real repetition. A binary format like Protobuf, Avro, or MessagePack typically relies on a real, shared schema — agreed on once, in advance, by whatever writes and reads the data — describing exactly which fields exist and in what real order. Because the schema already knows the real shape, individual records don't need to repeat field names at all; they store just the real, raw values, often in compact binary number representations instead of readable text digits, producing a genuinely smaller real result for the exact same information.
  • The real, honest cost of that compactness. A binary record isn't readable by eye — opening one directly shows meaningless raw bytes, not text Mike could glance at and understand. And because meaning depends on a real, shared schema, reading a binary record correctly requires having that exact schema available — a real, added piece of infrastructure JSON's self-describing nature never actually needs.
  • GreenMart's own, real, correct trade-off. For GreenMart's event log — millions of genuinely identical-shaped records a day, read only by GreenMart's own internal systems, never by a person glancing at raw storage — a binary format is a strong, real fit: the schema is known and controlled internally, and the real savings in storage and processing cost, multiplied across millions of records, is genuinely significant. For a public API response a developer might inspect directly in a browser, JSON's own real readability is often worth its real extra bytes.
Table — JSON vs. Binary Formats — The Real Trade-off
PropertyJSONBinary Format
Human-readableYes — plain text, self-describingNo — raw bytes, requires the shared schema to interpret
Field names repeated per record?Yes, in full, every timeNo — schema defines them once, shared in advance
Typical size for the same dataLargerSmaller — often significantly, at scale
Needs a shared schema in advance?NoYes

GreenMart's event log — high-volume, internally-consumed, identically-shaped records — is a strong real fit for a binary format; a public, human-inspected API response often favors JSON's own readability instead.

GreenMart now has the real, precise fix for its own growing storage bill: not a smaller retention window or a cheaper storage tier, but a genuinely more efficient real format for exactly the kind of high-volume, internally-consumed, identically-shaped data the event log actually is.

Key Takeaway

JSON trades real storage and processing efficiency for human readability and needing no shared schema; binary formats trade that readability away for genuinely smaller, faster real records — the right real choice depends on whether the data is read by a person glancing at it, or purely by systems that already agree on its shape.

Why This Matters

As GreenMart's own event volume keeps growing, the serialization format choice compounds — a real, small per-record savings, multiplied across millions of daily events, becomes a genuinely significant real cost difference, exactly the kind of decision this chapter equips GreenMart to make deliberately.

GreenMart now has the real trade-off between JSON and binary formats, and the right real fit for its own high-volume event log. The next chapter covers a related but distinct real lever on the exact same cost: compression.

Next