Betterflag
Back to blog

Feature Flag Best Practices That Actually Scale

Practical feature flag best practices: naming, defaults, targeting, percentage rollouts, kill switches, and how to retire flags before they become debt.

Mehdi
August 16, 2026
#feature-flags#best-practices#rollouts#kill-switch#audit-logs

Most feature flag advice is either too cute ("flags are a lifestyle") or too enterprise ("form a Flag Governance Council"). The useful layer is smaller. These are the feature flag best practices I actually follow, and the ones I am building Betterflag around.

If you are new to the idea, start with what feature flags are. This post assumes you already have an isEnabled call and want it not to rot.

1. Name the flag after the feature, not the ticket

checkout-v2 is a good flag key. JIRA-1847 is a crime scene. Six months later nobody remembers what 1847 was, and the flag is still in the payment path.

A flag key should read as a sentence in the code: "is checkout v2 on for this user?" Include the surface (checkout, billing, onboarding) and the change (v2, new-tax-flow). Use kebab-case. Never reuse a key; if you resurrect a feature, make a new flag.

2. Default new flags off

A new feature that defaults to on will ship itself the moment someone creates the flag in production, or the moment the SDK fails open. Default off. Turn it on in development and staging first. Then a named user. Then a percentage.

Kill switches are the exception: they wrap something already live, and they default to on (feature stays up) until you slam them off. Be explicit in the flag description so the next person does not invert it at 2am.

3. Fail closed on money, fail open on chrome

If the flag service is down, your app still has to decide. Pick the failure mode per flag:

  • Fail closed for checkout, auth, billing, permissions, anything that can charge a card or leak data. The old path stays.
  • Fail open for banners, copy tweaks, a new empty state. Users see the new thing; worst case you look a little unfinished.

A global "if SDK errors, return false" is safer than the reverse, and still wrong for a kill switch (a kill switch that fails closed would enable the broken feature when the service is down). Set it on the flag.

4. Never nest flags

if (await flags.isEnabled("checkout-v2")) {
if (await flags.isEnabled("checkout-v2-apple-pay")) {
// you are now in matrix math
}
}

Two booleans is four states. Three is eight. Nobody will test eight. Combine related behavior into one flag, or sequence the rollouts: ship Apple Pay behind checkout-v2 only after checkout-v2 is at 100% and the first flag is gone.

If you need a multivariate, you wanted an experiment platform, not a pile of booleans. Statsig and GrowthBook are honest tools for that job. Betterflag is not.

5. Keep targeting boring

Plan, country, a beta list, a user id. That is enough targeting for most products.

The moment you build a nested AND/OR DSL ("users in DE except plan=free and app version < 3.2 or internal staff"), you will ship a rule nobody can explain during an incident. Put the rare cases in code if you must. Keep the flag configuration something you can read in one glance.

Percentage rollouts should be sticky: hash the user id so the same 10% stay in the 10%. Random-per-request rollouts create flickering UI and unusable metrics. Details in how to do a percentage rollout.

6. One owner, one removal date

Flags without owners become archaeology. When you create a flag, set:

  • Owner: a person or a team, not "eng".
  • Type: release (temporary) or ops (kill switch / config that might stay).
  • Remove after: a date. For a release flag, that is "two weeks after 100%." Put it in the description. Calendar it if you are the kind of team that calendars things.

Review stale flags the way you review flaky tests. A flag that has been 100% on for 30 days is an if you should delete, plus the dead branch.

7. Log every change, including the robots

When a flag flips at 2am, "who did this?" needs a real answer. That is an audit log, not a Slack screenshot.

If your coding agent can create flags and stage rollouts over MCP, the audit trail has to name the agent key, not just the human who issued it. Shared admin tokens make this impossible. Agent-scoped keys are a best practice now, not a novelty.

8. Evaluate as close to the user as you can, cache the rest

Do not fetch the flag service on every React render. Do not hide a flag check behind a waterfall of client components if a Server Component or middleware could have decided already.

Patterns:

  • Server / edge: evaluate once per request, pass the boolean down as a prop.
  • Client: subscribe to a cached snapshot. Do not fetch in useEffect for a boolean.
  • Local evaluation: SDKs that download the config and evaluate in-process avoid a network hop on the hot path.

Feature flags in Next.js and React cover the framework-specific version of this.

9. Test both sides, especially the off path

The off path is what production is running while you roll out. If you only screenshot the new UI, you will discover the old one is broken when you have to kill the flag.

Minimum:

  • Unit test both branches.
  • One integration test where the flag is forced off, one where it is on.
  • A way to override flags in local and CI (BETTERFLAG_OVERRIDES=checkout-v2=on or a test helper). Never point CI at production flags.

10. Do not use flags as a database

A flag is a decision, not a place to store prices, copy decks, or user preferences. Remote config can hold a small JSON blob. The moment you are editing a novel in the dashboard, you wanted a CMS.

Same rule for secrets. Flag attributes and variation values show up in logs, SDKs, and admin UIs. API keys and tokens stay in a secret manager.

A short checklist you can paste in a PR

  • Flag key names the feature, not the ticket.
  • Default is off (unless it is a kill switch).
  • Failure mode is explicit: closed for money, open for chrome.
  • No nested flags.
  • Targeting is plan / country / cohort / id, not a novel.
  • Percentage assignment is sticky on user id.
  • Owner and removal date are on the flag.
  • Audit log will show who (or which agent) flipped it.
  • Both branches have a test.
  • A calendar reminder exists to delete it.

That is the whole practice. The rest is product: percentage rollouts, a kill switch that is actually instant, and a vendor whose pricing does not punish you for hiring or for getting users.

Betterflag is the flags-only version of that list. Join the waitlist if you want to try it; alpha users lock in 50% off for life.

FAQ

What are the most important feature flag best practices?
Name flags after the feature, default new flags off, never nest flags, keep targeting simple, make assignment sticky for percentage rollouts, fail closed on risky paths, and delete flags that have been 100% on. Treat every flag as temporary until you decide it is a permanent kill switch.
How long should you keep a feature flag?
Release flags should die after the rollout finishes, usually days to a few weeks. Kill switches and ops toggles can live longer, but they still need an owner. A flag that has been 100% on for a month is an if-statement you should delete.
Should feature flags default to on or off?
Off, unless the flag is a kill switch for something already live. A new feature that fails open will ship itself if the SDK cannot reach the service. That is rarely what you wanted.
How do you avoid feature flag debt?
Create flags with a removal date, list stale flags in the same review as flaky tests, and refuse to add a new flag in a file that already has two. Flag debt is just conditional debt with a dashboard.