Feature Flag Kill Switch: Instant Rollback Without a Redeploy
A kill switch is a feature flag that turns a broken path off in seconds. How it differs from a revert, what to wrap, and how to pull it from a dashboard or an agent.
A kill switch is a feature flag you pulled in anger. The new checkout is throwing. You set checkout-v2 to off. The next request renders the old checkout. No revert, no CI, no "who can deploy to prod." Seconds, not a pipeline.
If you have ever waited on a rollback while the site burned, this is the pattern you wanted. It is also the pattern a surprising number of flag vendors put behind an enterprise quote once you ask for the audit log that proves who pulled it.
Kill switch vs revert vs rollback vs "just deploy the old build"
| Move | What it undoes | How fast | What it needs | |---|---|---|---| | Kill switch | Serving the new path | Seconds | The old path still in the bundle, a flag you can flip | | Revert + deploy | The commit | Minutes to an hour | CI, a clean revert, someone with deploy access | | Infra rollback | A release artifact | Minutes | Your host's rollback, and hope the DB still matches | | Scale to zero | The whole app | Fast, and everyone is down | Panic |
Use a revert when the new code should not exist (a secret in the repo, a broken migration you must undo). Use a kill switch when the new code is fine to keep on disk and wrong to keep in the request path. Most product incidents are the second kind.
A kill switch cannot undo a destructive migration. Flags are for behavior, not for schema. Ship migrations expand-then-contract, behind flags, or not at all.
The off path is the product
const live = await flags.isEnabled("checkout-v2", { userId: user.id });return live ? <CheckoutV2 /> : <CheckoutV1 />;
That second branch is the kill switch. If you deleted <CheckoutV1 /> after a 100% rollout, you have a flag that can only go forward. Restore the old path before you need it, or do not call it a kill switch.
For a new path with no predecessor (a brand-new route), the off path is a 404, a redirect, or a maintenance view. That still counts. "Off" has to mean something safe.
Failure mode matters. If the flag service is unreachable, a kill switch should fail closed on the new path: serve old, not new. A marketing banner can fail open. Checkout cannot. Best practices covers fail-open vs fail-closed.
What to wrap
Wrap the blast radius, not every CSS class.
Good kill switches:
- A new checkout, billing, or tax path
- A new auth provider
- A new query or cache that can stampede
- A third-party SDK you can no longer trust this afternoon
- An agent-written feature you have not read
Bad kill switches:
- Every button
- A flag per file
- Nested flags that require a truth table to disable
One flag per risky unit. Percentage-roll it first (percentage rollouts), keep the off path, and you already have the kill switch. The rollout is the rehearsal.
Who is allowed to pull it
During an incident, the person who notices should be able to turn the feature off. That is not "only the platform team" and it is not "anyone with the production AWS account."
Practically:
- Humans: anyone on-call, via the dashboard or a slash command that hits the API.
- Agents: an MCP
kill_flagtool with an agent-scoped key. Every action lands in the audit log as that key, not as a shared admin user. - Production guardrail: require a human confirmation for prod kill switches if you do not want an agent to do it unsupervised. Staging can be autonomous.
The audit log is not optional. "Someone turned checkout off" is not an answer. "Agent key bf_agt_checkout-bot at 02:14 UTC, approved by Mehdi" is an answer. If that log is an enterprise SKU, you are renting the safety mechanism. I wrote up the vendor version of this in LaunchDarkly's enterprise audit log.
Speed is the SLA
A kill switch that propagates in five minutes is a slow deploy with extra steps. You want the next evaluation to see off.
That means:
- SDKs with a short stream or poll interval, or edge evaluation that reads a fresh snapshot.
- No "rebuild the static site" in the path. If your marketing page bakes flags at build time, it is not killable. See feature flags in Next.js.
- A dashboard, API, and MCP that all flip the same object. Three sources of truth is how you kill staging and miss prod.
Betterflag evaluates at the edge, under 100ms. The write (kill_flag) is the same API the dashboard uses. The MCP server has full parity, so Claude Code can pull it without a tab.
A drill, not a document
Once a quarter, kill a harmless flag in production and time it. If it takes more than a minute, or if nobody remembers how, the runbook is fiction. Incident tools that have never been used do not exist.
Write the name of the flag in the runbook, not "the checkout flag, I think it is under Releases."
Kill switch checklist
- Off path still in the code.
- Fail closed on the new path if the service is down.
- One flag per blast radius.
- On-call can flip it. Agents can flip it with a scoped key and an audit line.
- Prod can require a human tap.
- Propagation is seconds, not a rebuild.
- You have drilled it.
Feature flags are how you get here. What they are, how to roll them out, and a product that does not upsell the log: Betterflag, from $9.99/mo, audit log on Scale at $99.99 with no sales call. Alpha waitlist members lock in 50% off for life.
FAQ
- What is a feature flag kill switch?
- A kill switch is a feature flag whose off path is the last known good behavior. When something breaks, you set the flag to off and traffic moves back to the old path in seconds, without reverting a commit or waiting on CI.
- How is a kill switch different from git revert?
- A revert undoes a commit and needs a new deploy. A kill switch leaves the new code on the servers and stops serving it. Reverts are for code that should not exist. Kill switches are for code that should exist but not be reachable right now.
- What features should have a kill switch?
- Anything whose failure takes down money, auth, or a widely used path: checkout, billing, login, a new query, a new dependency. Cosmetic changes can skip it. If you cannot describe the off path, you do not have a kill switch.
- Can an AI coding agent pull a kill switch?
- Yes, if your flag platform exposes it over an API or MCP, with an agent-scoped key and an audit log that names the agent. For production, require a human approval on kill-switch actions. Autonomy is right for staging. Production kill switches should be gated.
