---
title: "Shipping a Feature From Claude Code: Flag, Rollout, Kill Switch | Betterflag"
description: "A hands-on walkthrough of a full agent-driven release: Claude Code (or Cursor) writes the code, creates the flag over MCP, stages a 10% production rollout, watches evaluation stats, and rolls forward. No dashboard required."
canonical: https://betterflag.app/blog/agent-driven-rollouts-claude-code-cursor
source: https://betterflag.app/blog/agent-driven-rollouts-claude-code-cursor
---
# Shipping a Feature From Claude Code: Flag, Rollout, Kill Switch

**Published:** 2026-08-21  
**Author:** Mehdi  
**Tags:** claude-code, cursor, feature-flags, rollout, tutorial, mcp

This is the demo I show people when they ask what "agentic-first feature flags" means in practice. One session in Claude Code, one feature, zero dashboard visits. Everything below works the same in Cursor or any MCP-capable client.

## The setup (once)

Connect the Betterflag MCP server. Easiest path is OAuth: add `https://mcp.betterflag.app/mcp` as a remote server, click Connect, pick your org. Or with an agent key from the dashboard's Keys page:

```sh
claude mcp add --transport http betterflag https://mcp.betterflag.app/mcp \
  --header "Authorization: Bearer bf_agt_..."
```

The agent now has tools like `create_flag`, `toggle_flag`, `set_rollout`, `set_targeting`, `get_evaluation_stats`, and `kill_flag`. Its key is scoped, and everything it does lands in the audit log attributed to that key, not to you.

## Step 1: build behind a flag

The prompt:

> Add the new pricing table variant to the billing page, behind a feature flag called `pricing-table-v2`. Create the flag: off everywhere.

Claude Code writes the component, wraps the render path in a flag check, and then (this is the part that used to be a human's job) calls `create_flag` over MCP. The flag exists in every environment, disabled, before the PR is even open.

```tsx
const showV2 = await flags.isEnabled("pricing-table-v2", { user });
return showV2 ? <PricingTableV2 /> : <PricingTable />;
```

## Step 2: verify in staging

> Enable pricing-table-v2 in staging and give me the URL to check.

One `toggle_flag` call, scoped to staging. Production is untouched. You click around staging like a normal person. This is the moment to catch the embarrassing stuff, while the blast radius is zero.

## Step 3: stage the production rollout

> Looks good. Roll it out to 10% in production and keep it pinned on for the beta cohort.

Two calls: `set_rollout` (10%, production) and `set_targeting` (beta segment always on). Evaluation is deterministic per user: the same visitor stays in the same bucket, so nobody flickers between the old and new table on refresh.

## Step 4: watch it, then roll forward

> How is pricing-table-v2 evaluating in prod over the last hour?

`get_evaluation_stats` answers with real numbers: evaluations, distribution across variants. Pair it with your error tracker: if checkout conversion holds and errors are flat, keep going.

> Take it to 50%. …Take it to 100%.

Each step is a one-line instruction and an audited API call that propagates to the edge in seconds, not a deploy.

## Step 0, really: the kill switch

The reason this whole workflow is safe to hand to an agent:

> Kill pricing-table-v2.

`kill_flag` turns it off everywhere, instantly. Under 100ms evaluation at the edge means the bad path stops being served before you've finished typing the Slack apology. And because *enabling* was staged, disabling is boring: no revert commit, no redeploy, no 20-minute pipeline while production burns.

## The guardrail question

"So an agent can just flip things in production?" Only if you configure it that way. Guardrails let you require human confirmation for prod-affecting actions: the agent proposes the change, you approve it, both are audited. For staging, let it run free; for `kill_flag` on prod, maybe you want autonomy (a fast kill beats a slow approval); for a 100% prod rollout, maybe you want the confirmation click. The point is that *you* draw the line, per action, instead of the tool assuming every operator is a human with a seat.

## Why I built it this way

Every step above used to be a context switch: IDE → dashboard → IDE → dashboard. Agents made the pattern unbearable, because the agent finishes the code in minutes and then waits on a human to do clicks. Making the flag platform speak MCP with full API parity (same capabilities as the dashboard, [priced by evaluations](/pricing)) removes the wait entirely.

The dashboard's still there. I just haven't needed it to ship in weeks. Setup and the "why" live in [Feature Flags Over MCP](/blog/feature-flags-mcp).

## FAQ

### Can Claude Code create and roll out feature flags?

Yes, if your flag platform speaks MCP. With Betterflag connected, Claude Code can create_flag, toggle_flag in staging, set_rollout in production, and kill_flag, with every action attributed to an agent-scoped key.

### How do you do a 10% rollout from Cursor or Claude Code?

After the flag exists and staging looks good, ask the agent to set the production rollout to 10% and pin it on for a beta cohort. Assignment is sticky on user id so the same visitor stays in the same bucket.

### What if the agent-shipped feature breaks?

Ask it to kill the flag, or hit the kill switch yourself. The off path has to still exist in the code. Evaluation at the edge means the bad path stops being served in seconds, not after a revert and a pipeline.

### Should production kill switches require a human?

You draw the line. Staging can be free. A production 100% ramp might want a confirmation click. A production kill switch might want autonomy, because a fast kill beats a slow approval. Both sides of the decision are audited.
