Tech #020•11 min read•21 September 2026 , Monday

Your Feature Flag Is Off. Your Code Is Still Running.

A flag decides which path executes. It does not decide what got deployed, migrated, scheduled or left behind — and that gap is where production incidents live.

Rajnish Kumar

Rajnish Kumar

Editor-in-Chief & Founder

Your Feature Flag Is Off. Your Code Is Still Running. — Tech dispatch hero image

#01A Flag Makes a Change Conditional, Not Absent

Teams reach for a Feature Flag because it feels like an undo button: ship the code, keep the switch off, flip it on when you're ready, flip it back if anything goes wrong. That mental model is half right, and the wrong half is the expensive one. When a flag is off, the new code is still in the binary. Its imports still load, its module-level initializers still run, its background workers are still registered, and any database migration that shipped with it has already been applied. The only thing the flag controls is one branch in one place. Everything around that branch is live.

The reason this matters is that most flag discipline is built around the moment of turning a flag on: percentage ramps, monitoring, approval. Very little of it looks at the long stretch before that moment, when the change is "off" but fully deployed, or at the stretch after, when the flag is 100% on, nobody is watching it, and it quietly turns into permanent, undocumented branching in the code. The rest of this piece walks through what actually stays active while a flag is off, one real incident where the leftover code did the damage, and what a flag's lifecycle needs to look like to keep it honest.

“A feature flag doesn't make a change disappear. It makes the change conditional.”

#02Deployment Is Not Release

The idea that made flags mainstream is separating two events that used to be the same one. Deploying is putting code on servers. Releasing is exposing behavior to users. Pete Hodgson's widely cited Feature Toggles article on martinfowler.com describes release toggles as the mechanism that lets teams practice trunk-based development by shipping incomplete features as latent code, and notes that this is what lets product owners decide when to release independently of when engineers deploy. A dark launch is the same idea taken further: the new path runs in production, often on real traffic, but its result isn't shown to anyone.

That separation is genuinely valuable, and it comes with a cost that's easy to miss: the moment you split the two, "deployed" stops meaning "safe." A change can be fully deployed and have never been exercised. Whatever runs regardless of the flag's value has to be safe on its own, and the flag can't help with that, because the flag isn't even consulted. Concretely, here is what keeps running while the flag is off:

  • The new code itself, including anything that executes at import, startup or dependency-injection time.
  • Every database migration that was bundled with the change, since schema changes apply to the database and not to a code path.
  • Background job and cron registrations, queue consumers and scheduled tasks that were added alongside the feature.
  • Dependency upgrades that came in with the change, which affect every code path, flagged or not.
  • The flag evaluation call itself, which is a network or cache lookup on the hot path and has its own failure modes.
  • Any new message formats, cache keys or API fields that other services can now observe or depend on.

#03The Incident Where the Old Code Ran: Knight Capital

The clearest real-world case of leftover code being triggered by a flag is Knight Capital on 1 August 2012, and the SEC's own administrative order is unusually specific. Knight was adding support for the New York Stock Exchange's Retail Liquidity Program to its order router, SMARS. The new code was meant to replace an old, unused feature called Power Peg, which Knight had stopped using in 2003. According to the order, the Power Peg code was never deleted: it "remained present and callable," and the new code repurposed a flag that used to activate Power Peg. In 2005, Knight had moved the function that tracked how many shares of an order were already filled to an earlier point in the code, without retesting Power Peg to see whether it still worked when called.

The rollout began on 27 July, in stages across eight servers. One technician failed to copy the new code to one of them, and no second person reviewed the deployment — the SEC notes Knight had no written procedure requiring one. On 1 August, orders carrying the repurposed flag reached that eighth server, which ran the old Power Peg logic. Because the filled-quantity tracking had been moved in 2005, that server never learned an order was complete and kept sending child orders. In roughly 45 minutes, 212 parent orders turned into more than 4 million executions across 154 stocks, and Knight lost over $460 million. The SEC also records that automated emails reading "Power Peg disabled" went out 97 times before the market opened, but they weren't designed as alerts and weren't generally read.

Two details from the order are worth keeping. First, during the incident Knight uninstalled the new code from the seven correctly deployed servers, which the SEC says "worsened the problem," because those servers then ran Power Peg too — the rollback itself activated the dead path. Second, the SEC found Knight had no written protocol for retesting unused code sitting on production servers. This wasn't a modern flag platform, and it would be wrong to blame "feature flags" as a technique. It's a precise demonstration of what this article is about: old code that is "off" is still code, and reusing the switch that once guarded it is a way of turning it back on. The key events, in order, are summarized below.

WhenWhat happenedWhy it mattered
2003Knight stops using Power Peg but leaves the code on production serversThe dead path stayed callable
2005Cumulative-quantity tracking is moved earlier in the code; Power Peg isn't retestedThe old path was silently broken
27 July 2012Staged RLP rollout begins across eight servers; one server is missedSeven servers ran new code, one ran old
1 August 2012Orders with the repurposed flag reach the missed serverThe broken old path ran on live orders
1 August 2012New code is uninstalled from the other seven serversThe rollback spread the failure

#04The Database Doesn't Honor Your Flag

Of everything that stays live when a flag is off, the database is the piece teams most consistently forget. A migration that drops a column, renames one or adds a NOT NULL constraint changes what every version of the code can do, including the old one you'd fall back to if you flipped the flag off. "Turning the flag off" therefore only restores the old behavior if the old code still works against the new schema.

The standard discipline is parallel change, also called expand and contract: Joshua Kerievsky documented it in 2006, and Martin Fowler's write-up describes three phases — expand to support both old and new, migrate clients incrementally, then contract by removing the old version. Fowler calls it a key component of evolutionary database design, and notes it is valuable for continuous delivery precisely because code can be released in any of the three phases. Stripe's description of its online migrations is the same shape at scale: dual-write to old and new tables and backfill, change the read paths, change the write paths, and only then remove the old data, using GitHub's Scientist library to compare old and new results in production while the change is in flight.

The practical rule that falls out of this is to keep the writes unconditional and put the flag only on the reads, so the old column stays correct for as long as anyone might still read from it. Concretely, this looks like:

typescript
// Phase 1 (expand): both columns are written on every save, no flag involved.
async function saveCustomer(c: Customer) {
await db.customers.update(c.id, {
phone: c.phone, // old column: still the source of truth
phone_e164: normalizeE164(c.phone), // new column: kept in sync
});
}
// The flag only chooses which column is READ.
async function getPhone(id: string, user: User): Promise<string> {
const row = await db.customers.get(id);
return flags.enabled("read-phone-e164", user, false)
? row.phone_e164
: row.phone;
}

#05Old and New Code Run Side by Side

Even with a clean schema, the "off" state means two versions of the logic are alive at once, and they don't only meet in the flag check. During a rolling deploy, some instances run the new build while others run the old one. A queue can hold messages produced by one version and consumed by the other. A cron job can start under one flag value and finish under another. A user can be in the new experience on one request and the old one on the next if the flag is evaluated per request against something that changed in between.

Each of these is a compatibility question that the flag itself can't answer. Message payloads need to be readable by both versions for as long as both are running. Long-running work should capture the flag decision once, when the work starts, and carry it forward instead of re-asking halfway through, otherwise a job can begin on the old path and finish on the new one. And any state that the new path writes and the old path doesn't understand needs an explicit plan for what happens to it if the flag is turned off again. None of this shows up in a code review of the flag check, because it lives in the parts of the system the check doesn't touch.

#06Not All Flags Are the Same Kind of Flag

Much of the mess around flags comes from treating them as one thing. Hodgson's article separates them into four categories with very different lifespans, and the category should decide how the flag is managed. A release toggle is meant to live for days or weeks and then be deleted; an experiment toggle lives as long as the experiment needs statistical significance; an ops toggle can be a permanent kill switch; and a permissioning toggle can live for years. The trouble starts when a release toggle is created and then behaves like a permanent one because nobody deleted it, or when an ops toggle is treated as disposable and removed while an operator still depends on it. The categories are compared below.

CategoryTypical lifespanHow dynamicExample
ReleaseDays to weeksStatic, per deployHide an unfinished checkout redesign
ExperimentHours to weeksPer user, consistent cohortsA/B test a new button label
OpsShort-term to permanentFast, operator-controlledKill switch for a costly recommendation call
PermissioningYearsPer request, per user groupBeta features for internal staff

#07Flag Debt Is Real, and It's Measurable

Hodgson's framing is that teams should treat flags as inventory that carries a cost, quoting the line that savvy teams "view their Feature Toggles as inventory which comes with a carrying cost, and work to keep that inventory as low as possible." His practical suggestions are unglamorous and effective: create a removal task at the moment you create the flag, put an expiry date on it, and even write a test that fails once a flag outlives its date, so the debt shows up in the build instead of in someone's memory.

Uber's engineering team measured how large this problem gets. In their ICSE 2020 paper, Piranha: Reducing Feature Flag Debt at Uber, they describe an automated refactoring tool that generates a diff to delete the code behind a stale flag and assigns it to the flag's author. Between December 2017 and May 2019 it generated cleanup diffs for 1,381 flags, about 17% of all flags, and 65% of those diffs landed without any changes. The number to take from that isn't the tool's success rate but the scale: a mature engineering organization found that a meaningful share of its flags had already outlived their purpose and were sitting in production as dead branches. Every one of them is a path someone has to remember could still be reached.

#08The Combinations You'll Never Test

Every flag doubles the number of possible system states, so ten independent flags means 1,024 configurations, and no team tests all of them. Hodgson's practical answer is to test a small number of configurations on purpose: the expected production configuration, the fallback state with the intended flags off, and everything on. He also makes the point that most flags don't interact with each other and most releases change only one flag, which is what makes this tractable. The failure mode to watch for is the flag pair that does interact — two flags that both change the same data path, for instance — because that combination is exactly the one nobody thought to try.

The other question is what happens when the flag system itself misbehaves. The OpenFeature specification, the vendor-neutral standard for flag evaluation, says a client must not throw or terminate abnormally on a failed evaluation and must instead return the caller-supplied default value. That makes the default an engineering decision, not a formality: when the flag service is unreachable, does the code fall back to the old behavior or the new one? For a flag guarding an unfinished feature the answer is almost always "off," but for a kill switch protecting a downstream dependency, "off" might be the dangerous one. It deserves a deliberate choice, made per flag, and written down next to it.

#09Percentage Rollouts and Kill Switches Need Care Too

A percentage rollout is only trustworthy if the same user lands in the same bucket every time. The usual technique is to hash a stable identifier together with the flag's key, so a user doesn't flip between experiences from one request to the next and so different flags bucket independently of each other. A minimal version of that looks like the following, where the flag key acts as a salt:

A ramp is also only as informative as the monitoring behind it: 5% of traffic that nobody is watching tells you nothing, and a ramp on a change that touches shared state, such as a database or a queue, isn't really limited to 5% at all. Kill switches deserve the same suspicion. A kill switch that has never been flipped is untested code, and the Knight Capital order is a sharp reminder: the "off" action they took, removing the new code from seven servers, caused more damage by reactivating the old path. If a switch is meant to be the emergency exit, flip it on purpose in a rehearsal, in both directions, before the emergency.

typescript
import { createHash } from "node:crypto";
// Stable bucket in [0, 100) for this user and this flag.
function bucket(flagKey: string, userId: string): number {
const digest = createHash("sha256").update(`${flagKey}:${userId}`).digest();
return (digest.readUInt32BE(0) % 10_000) / 100;
}
const inRollout = bucket("new-checkout", user.id) < 25; // 25% of users

#010A Flag Lifecycle That Keeps It Honest

None of this argues against flags. Decoupling deploy from release is one of the best ideas in modern delivery, and flags are how most teams do it. What the evidence above argues for is treating a flag as a loan, with a repayment plan attached at the moment it's created. In practice that comes down to a short list that every flag should have written down before it merges:

  • An owner, a category (release, experiment, ops or permission) and an expiry date.
  • A cleanup ticket created at the same time as the flag, not after the rollout.
  • A stated default for when the flag service is unavailable, and a reason for it.
  • Confirmation that the code works against the schema and message formats in both flag states.
  • A tested path for turning it off, exercised at least once before it's needed.
  • A removal order: delete the dead branch first, then the flag configuration, then contract the schema.

#011The Switch Is Not the Safety

The flag is not the safety mechanism; the safety comes from what surrounds it. A switch that's off still leaves a change deployed, migrated and scheduled, and the honest way to describe it is the one this piece started with: it didn't make the change disappear, it made it conditional. The teams that stay out of trouble are the ones who remember what is still running underneath the switch, and who delete the branch once its job is done.

Found this useful? Share it

Have a technical response or architectural perspective to share with the engineering desk?

Submit Engineering Feedback