Betterflag
Back to blog

Why I'm Building Betterflag

Agents made shipping faster than releasing is safe. Feature flags are the control plane for that world, and the current tools were built for a different one.

Mehdi
August 19, 2026
#founders#feature-flags#agents#product#story

I used to think feature flags were an enterprise hobby.

Something you adopted after you had a platform team, a compliance checklist, and a LaunchDarkly AE in your inbox. For everyone else: an if, an env var, a comment that said TODO: clean this up. It worked until it didn't, and when it didn't you reverted a commit and promised yourself you'd "do flags properly" next quarter.

That was before my coding agent started shipping faster than I could safely release.

The release math changed

Here is what a normal week looks like now. I describe a feature. Claude Code or Cursor writes it. Tests go green. The PR is open before I have finished my coffee. The code is better than the first draft I would have written myself, and I have not read every line, because the volume does not allow it.

That is the point of the AI era. Generation got cheap. Review did not.

The old contract of shipping was: a human wrote the code, so a human roughly knew what was going into production. The new contract is: a model wrote a lot of code, a human skimmed it, and then someone has to decide whether users should see it. Those are two different decisions. Most teams still treat them as one.

Feature flags are how you split them.

Deploying means the code is on the servers. Releasing means a user can touch it. When those two are the same event, every agent-written feature is a bet you cannot unwind without a revert, a redeploy, and a twenty-minute pipeline while production burns. When they are not, the bet has a kill switch. The code can sit in production, dormant, until you are ready, and if it misbehaves you turn it off in seconds instead of rolling back history.

I did not come to this as a positioning exercise. I came to it because I kept watching agents finish the hard part and then stall on the dangerous part. The agent can write the checkout flow. It cannot, in most flag tools, create the flag, stage the 10% rollout, or pull the kill switch. So the human alt-tabs to a dashboard and does data entry. The most capable operator in the workflow is locked out of the control plane by a UI.

That is a stupid division of labor. It is also how the entire category still works.

Flags matter more now, not less

People keep asking whether feature flags still matter when agents can just... ship. The question has the causality backwards.

Agents made flags more important, because they made the blast radius cheaper to create and more expensive to understand.

More code per day means more surface area in production. More surface area means more ways to break checkout, auth, billing, the one flow that pays the rent. You cannot compensate for that with more careful prompting. You compensate for it with reversibility. A kill switch is the difference between "my agent shipped something weird" being an anecdote and being an incident.

There is a second reason, quieter, that I think the industry has not absorbed yet. Flags are no longer just a product tool. They are the runtime API of the product. The agent that writes your code needs a way to change behavior without waiting for the next deploy. A config file in git ships at deploy speed. A flag at the edge ships in seconds. If you want an agent to roll something to 10%, watch it, and either continue or kill it, the flag is the interface. The dashboard is a human artifact taped on top.

That is why I keep saying the MCP server should be the product, not an integration. Agents do not click. They call tools. A flag platform that only speaks "human with a seat and a browser" is a platform for a workflow that is already shrinking.

Then I opened the existing tools

So flags are more important. The obvious move is to use one of the tools that already exist. I did. Then I bounced.

Not because LaunchDarkly is bad. It is not. It is a serious piece of infrastructure for a serious kind of company: hundreds of engineers, a platform team, change-management, multivariate experiments, a procurement process. If that is you, you should probably keep paying them. They built the product for you.

The problem is that the rest of us were told this is what feature flags are.

A targeting DSL with nested AND/OR, attribute types, and a rule builder that looks like a query engine. Environments as a product surface. Segments, experiments, metrics, holdouts. Approval workflows designed for orgs that have a person whose job is "governance." Pricing that meters seats, service connections, and monthly active users at the same time, so your bill goes up when you hire, when you split a service, and when your product succeeds.

I do not need a physics lab to hang a picture frame. I need to turn a feature on for 10% of users, pin it on for the beta list, and kill it if the error rate jumps. That is the job. That is most people's job. The 90% use case got designed and priced for the 10%.

Look at the rest of the shelf and the pattern repeats, just with different gravity wells. Statsig is an experimentation platform with flags attached. PostHog is an analytics suite with flags attached. GrowthBook wants your data warehouse. The open-source options are real, and they still ask you to become the platform team. Somewhere along the way, an if-statement acquired a sales motion and a forty-page getting-started guide.

I keep meeting founders who say they do not need flags yet. What they mean is they do not want to adopt a platform. So they hardcode the boolean, ship, incident, revert, and tell themselves the same story next month. Complexity is not a feature they are declining. It is the tax that keeps them on the unsafe path.

Complexity is a product decision

I want to be precise about this, because "too complex" is easy to say and easy to dismiss as "you just have not grown into it."

Most teams, including teams that will one day be large, spend most of their flag life on four primitives:

  1. On or off.
  2. A percentage rollout.
  3. A small amount of targeting. Plan, country, a beta cohort. Not a rules engine.
  4. A kill switch that is faster than a deploy.

Everything else is real, and some of it is even good. It is also not why you adopt flags on day one, and it is not what an agent needs to call at 11pm when a canary looks wrong. The category optimized for the power-user ceiling and left the floor unusable. Setup that takes a week. Concepts you have to learn before your first toggle. A dashboard that assumes the operator is a human who has been trained on the dashboard.

Agents made that mismatch obvious. A coding agent can hold about four tools in its head and use them well: create the flag, set the rollout, target a cohort, kill it. Give it a 47-field targeting model and it will either hallucinate a rule or stop and ask you to click. The simple interface is not a beginner mode. It is the one that matches how the work actually happens now.

I think a lot of flag tools know this and cannot do anything about it. You cannot unship a decade of enterprise surface area. You cannot make the dashboard the optional view when the dashboard is the product, and the API is the thing you document in a sidebar. The incentives run the other way: more features, more seats, more "contact sales."

So I stopped trying to configure my way out of it.

What I actually built

Betterflag is feature flags and nothing else. No analytics suite. No experimentation platform. No CMS bolted on from a previous life. One category, on purpose, because the tools that tried to be everything are the ones that made flags an afterthought.

The primitives are the four above. The write path is the API and the MCP server, with full parity, so the agent that wrote the code can create the flag, turn it on in staging, roll it to 10% in production, and kill it without opening a tab. Every agent gets its own scoped key. Every action lands in the audit log attributed to that key. You can require a human tap for the scary stuff. This is not "trust the AI." It is narrower and more accountable than handing a human the admin dashboard.

The dashboard still exists. It is becoming the observation layer: rollout curves, audit trail, the place you look, not the only place you operate.

Pricing is one meter: evaluations. Unlimited flags, seats, environments. Starting at $9.99 a month. The short version is that we meter the thing that costs money to serve, not your headcount or your MAUs. LaunchDarkly's meters are the contrast.

Evaluations are served from the edge. Plumbing that is slow is not plumbing.

The bet

I am not building this because the world needs another boolean lookup. I am building it because the world started generating software faster than it can safely release it, and the tools that were supposed to make releasing safe got too heavy for the people who now need them most.

Feature flags are the control plane of the AI era. They should not require a platform team, a sales call, or a week of setup. They should be a sentence you say to your agent: gate this, roll it to ten percent, kill it if it breaks.

That is the product. Everything else is what we refused to add.

Betterflag is in private alpha. If that sentence matches how you actually ship, come try it. If you think agent-controlled flags are reckless, I want that argument too. I have guardrails for a reason, and I would rather hear where they should sit from people who have been burned than from people who have only written the blog post.

I am building this in public. The boring version, including the parts that go badly, is the only version worth writing.

FAQ

Why was Betterflag built?
Coding agents now write features faster than a human can safely release them. Feature flags split deploy from release: the code can sit dormant, then roll to 10%, then die in seconds if it breaks. Existing flag tools were built for dashboards and enterprise procurement, not for agents that need to call create_flag and kill_flag.
Is Betterflag an experimentation platform?
No. Flags only: on/off, percentage rollouts, a little targeting, and a kill switch. For a stats engine, Statsig or GrowthBook is the better pick.
How is Betterflag different from LaunchDarkly?
LaunchDarkly is the right tool for a large enterprise with a platform team. Betterflag is flags-only, priced on one meter (evaluations), with an MCP server as the primary write path. No MAU bill, no service-connection bill, audit log on a published Scale price.