Betterflag
Back to blog

Shipping a Feature From Claude Code: Flag, Rollout, Kill Switch

A hands-on walkthrough of a full agent-driven release: Claude Code (or Cursor) writes the code, creates the flag over MCP, stages a 10% production rollout, watches evaluation stats, and rolls forward. No dashboard required.

Mehdi
August 21, 2026
#claude-code#cursor#feature-flags#rollout#tutorial#mcp

This is the demo I show people when they ask what "agentic-first feature flags" means in practice. One session in Claude Code, one feature, zero dashboard visits. Everything below works the same in Cursor or any MCP-capable client.

The setup (once)

Connect the Betterflag MCP server. Easiest path is OAuth: add https://mcp.betterflag.app/mcp as a remote server, click Connect, pick your org. Or with an agent key from the dashboard's Keys page:

claude mcp add --transport http betterflag https://mcp.betterflag.app/mcp \
--header "Authorization: Bearer bf_agt_..."

The agent now has tools like create_flag, toggle_flag, set_rollout, set_targeting, get_evaluation_stats, and kill_flag. Its key is scoped, and everything it does lands in the audit log attributed to that key, not to you.

Step 1: build behind a flag

The prompt:

Add the new pricing table variant to the billing page, behind a feature flag called pricing-table-v2. Create the flag: off everywhere.

Claude Code writes the component, wraps the render path in a flag check, and then (this is the part that used to be a human's job) calls create_flag over MCP. The flag exists in every environment, disabled, before the PR is even open.

const showV2 = await flags.isEnabled("pricing-table-v2", { user });
return showV2 ? <PricingTableV2 /> : <PricingTable />;

Step 2: verify in staging

Enable pricing-table-v2 in staging and give me the URL to check.

One toggle_flag call, scoped to staging. Production is untouched. You click around staging like a normal person. This is the moment to catch the embarrassing stuff, while the blast radius is zero.

Step 3: stage the production rollout

Looks good. Roll it out to 10% in production and keep it pinned on for the beta cohort.

Two calls: set_rollout (10%, production) and set_targeting (beta segment always on). Evaluation is deterministic per user: the same visitor stays in the same bucket, so nobody flickers between the old and new table on refresh.

Step 4: watch it, then roll forward

How is pricing-table-v2 evaluating in prod over the last hour?

get_evaluation_stats answers with real numbers: evaluations, distribution across variants. Pair it with your error tracker: if checkout conversion holds and errors are flat, keep going.

Take it to 50%. …Take it to 100%.

Each step is a one-line instruction and an audited API call that propagates to the edge in seconds, not a deploy.

Step 0, really: the kill switch

The reason this whole workflow is safe to hand to an agent:

Kill pricing-table-v2.

kill_flag turns it off everywhere, instantly. Under 100ms evaluation at the edge means the bad path stops being served before you've finished typing the Slack apology. And because enabling was staged, disabling is boring: no revert commit, no redeploy, no 20-minute pipeline while production burns.

The guardrail question

"So an agent can just flip things in production?" Only if you configure it that way. Guardrails let you require human confirmation for prod-affecting actions: the agent proposes the change, you approve it, both are audited. For staging, let it run free; for kill_flag on prod, maybe you want autonomy (a fast kill beats a slow approval); for a 100% prod rollout, maybe you want the confirmation click. The point is that you draw the line, per action, instead of the tool assuming every operator is a human with a seat.

Why I built it this way

Every step above used to be a context switch: IDE → dashboard → IDE → dashboard. Agents made the pattern unbearable, because the agent finishes the code in minutes and then waits on a human to do clicks. Making the flag platform speak MCP with full API parity (same capabilities as the dashboard, priced by evaluations) removes the wait entirely.

The dashboard's still there. I just haven't needed it to ship in weeks. Setup and the "why" live in Feature Flags Over MCP.

FAQ

Can Claude Code create and roll out feature flags?
Yes, if your flag platform speaks MCP. With Betterflag connected, Claude Code can create_flag, toggle_flag in staging, set_rollout in production, and kill_flag, with every action attributed to an agent-scoped key.
How do you do a 10% rollout from Cursor or Claude Code?
After the flag exists and staging looks good, ask the agent to set the production rollout to 10% and pin it on for a beta cohort. Assignment is sticky on user id so the same visitor stays in the same bucket.
What if the agent-shipped feature breaks?
Ask it to kill the flag, or hit the kill switch yourself. The off path has to still exist in the code. Evaluation at the edge means the bad path stops being served in seconds, not after a revert and a pipeline.
Should production kill switches require a human?
You draw the line. Staging can be free. A production 100% ramp might want a confirmation click. A production kill switch might want autonomy, because a fast kill beats a slow approval. Both sides of the decision are audited.