Feature Flags vs Environment Variables: When to Use Each
Environment variables flip at deploy. Feature flags flip per user, at request time. A clear split for secrets, config, remote config, and runtime rollouts.
Teams reach for environment variables because they are already there. NEW_CHECKOUT=true in .env.production, redeploy, everyone gets the new checkout. It works until you need any of: a subset of users, a kill switch faster than a restart, or a staging/prod split that is not "another Vercel project."
That is the whole difference, and it is load-bearing.
Environment variables are process-wide and deploy-scoped. Every request on that instance sees the same value. Changing it means a new deploy, a new pod, or a restart.
Feature flags are per-evaluation and request-scoped. Two users hitting the same server can get different answers. Changing it means flipping a switch. The process does not restart.
If you came here from what are feature flags, this is the companion: when not to use them.
A table you can argue with
| You need | Use |
|---|---|
| Database URL, API secret, NODE_ENV | Environment variable (or a secret manager) |
| The same boolean for the whole process | Environment variable |
| A feature on for you, off for customers | Feature flag (user targeting) |
| A feature on for 10% of traffic | Feature flag (percentage rollout) |
| Turn it off in seconds during an incident | Feature flag (kill switch) |
| Different values per plan or country | Feature flag or remote config |
| A 12kb JSON of marketing copy | A CMS, not a flag |
| A credential | Never a flag |
Why NEXT_PUBLIC_NEW_CHECKOUT stops working
Client-side env vars in Next.js are inlined at build time. Change NEXT_PUBLIC_* and you must rebuild. There is no "flip it for the beta list." There is no "kill it without rolling back the deployment." There is no "10% of sessions."
Server-side env vars are a little better: you can restart to pick up a new value. You still cannot target users. On a platform with many instances, you will restart some before others and live in a split-brain for a few minutes. That is an accidental canary, not a rollout.
// Process-wide. Every user. Needs a restart.const newCheckout = process.env.NEW_CHECKOUT === "true";
// Per user. No restart. Can be 10%, a plan, or off in one click.const newCheckout = await flags.isEnabled("checkout-v2", { userId: user.id });
The first line is the right tool for DATABASE_URL. It is the wrong tool for a checkout rewrite.
Secrets never go in flags
Flag dashboards, SDKs, audit logs, and MCP traces are built to be readable. That is the point. A variation value that says "enabled" is fine. A variation value that says "sk_live_..." will leak into browser snapshots, support tickets, and agent logs.
Secrets belong in a secret manager or in environment variables that never ship to the client. If you would be sad to see it in a screenshot, it is not a flag.
Remote config sits in the middle, on purpose
Remote config is a feature flag that returns a value. Theme, copy, a numeric limit, a JSON list of plans. It is still evaluated per user, still killable without a deploy, still the wrong place for secrets.
Use remote config when:
- Non-engineers should change a value without a PR (a banner, a limit).
- The value should differ by plan or country.
- You want to tune something daily without rebuilding.
Do not use it when:
- The value is a secret.
- The value is a novel. Use a CMS.
- The value is infrastructure. Use env vars / Terraform / your host's config.
The Next.js-specific version of this, including caching and revalidation, is in feature flags in Next.js. The caching story is the hard part, not the boolean.
Staging vs production is not a flag
A common misuse: if (process.env.VERCEL_ENV === "production") wrapped in a flag named enable-in-prod. You wanted environments.
Keep environments (development, staging, production) as separate flag spaces, each with its own values. The flag key is the same (checkout-v2). The value in staging can be on while production is off. That is what environments are for. Collapsing them into one environment plus a targeting rule on hostname will eventually target the wrong host.
Betterflag treats environments as unlimited and included. The flag key stays stable; the values do not leak across.
A decision rule that holds up
Ask three questions:
- Does every instance, every user, need the same value? Yes: env var. No: flag.
- Is a restart an acceptable way to change it? Yes: env var is fine. No: flag.
- Would I be upset to see this value in an admin UI or an MCP log? Yes: secret manager / env var. No: flag or remote config is allowed.
If you answered "per user, no restart, not a secret," you want a feature flag. The remaining choice is whether to build it or buy it. Buying is a one-meter evaluation bill and an SDK. Building is a weekend, then a year of SDKs, audit logs, and edge caches.
Betterflag is the buy option if you want flags-only, predictable pricing, and an MCP server so the agent that wrote the feature can also create the flag. Environment variables stay in .env. That split is the architecture.
FAQ
- Should I use a feature flag or an environment variable?
- Use an environment variable when every instance of the process should share the same value and you are fine waiting for a deploy or restart to change it: secrets, database URLs, NODE_ENV. Use a feature flag when some users should see a feature and others should not, or you need to turn it off without a deploy.
- Are feature flags a replacement for .env files?
- No. .env files and secret managers are for process-wide configuration and credentials. Feature flags are for runtime decisions about user-facing behavior. Putting a database password in a flag dashboard is a security incident, not an architecture.
- What is remote config compared to a feature flag?
- Remote config is a feature flag that returns a value (a string, a number, a JSON blob) instead of just on/off. It is still evaluated per user at runtime. It is not a replacement for environment variables or for a CMS.
- Can I use environment variables for feature flags?
- You can, and plenty of teams start that way. The limit is that an env var cannot target 10% of users, cannot kill a feature without a restart, and cannot differ across users on the same server. The moment you need any of those, you have outgrown NEXT_PUBLIC_NEW_CHECKOUT=true.
