Spaces, Tabs, Panes, and Precise Words: Why I Break Agent Work Out the Way I Do

In the previous post I toured Herdr and how I lay it out: one space per folder, one tab per kind of work, agents in tall columns, shells in short stacks, all of it spread across a 5120×1440 ultrawide. That’s the what. This one is the why, and then the part that actually decides whether any of it pays off: how I talk to the agents once they’re in their panes.

Short version: the layout isn’t decoration. It’s context management for two kinds of brains, mine and the models’, and the prompts are the same idea at a smaller scale. Every tool in this stack, Claude Code, Codex, Cursor, and the human, gets worse the more unrelated junk is in its head. So I keep the junk out, structurally first and then verbally.

Part 1: Why I break it out like this

Spaces are folders because the folder is the agent’s whole world

An agent’s working directory isn’t a detail. It’s the edge of its universe. Whatever it can read, grep, build, and break starts at that folder. So when I make one space per folder, I’m not organizing windows. I’m drawing the same boundary the agent already lives inside, and making it visible to me.

That buys three things:

  1. The sidebar is a readout of my working memory. Ten spaces means ten things in flight. If a folder isn’t in the sidebar, I’m not working on it. When I’m done, the space goes away. There’s no “misc” space and no “scratch” space, because “misc” is where context goes to die.
  2. Blast radius is legible. If something is going sideways in ledger-service, the dot next to ledger-service says so. Nothing in that space can be quietly editing storefront-web, because it doesn’t live there.
  3. It pairs with worktrees. When I need two implementation agents on the same repo at once, each gets its own worktree and its own branch, and in Herdr that’s just another folder, so another space (herdr worktree create does this directly). I made the full case in Giving Every Agent Its Own Branch: Git Worktrees for Parallel AI Work: “Run two coding agents against a single working tree and they will fight.” Spaces-as-folders keeps the referee’s view honest.

Tabs are modes because each mode is a different verb

I name tabs domain-modeling, implementation, and pr-reviews because those are three different jobs. Each one comes with a different posture from me and a different register of prompt to the agent.

TabThe verbMy postureWhat the agent is allowed to do
domain-modelingthinkSkeptical, SocraticRead, reason, write a model. No code. Plan mode.
implementationchangeDirective, fencedEdit inside an explicit scope, after an approved plan
pr-reviewsjudgeAdversarial, evidentiaryRead the diff, report findings. No fixing.

The failure mode I’m designing against is verb drift. You start a conversation reviewing a PR, you say “huh, good catch,” and three messages later the “reviewer” is rewriting the module, now with the confidence of something that just graded its own homework. Separate tabs, and separate agent sessions in them, make the verb change explicit. If I want a fix, I go to the implementation tab and ask for one there, with a scope. It’s the same rule I put on the read-only reviewer agents in Ask, Plan, Confirm: Making Agents Stop Before They Start: the reviewer doesn’t get to “slide from ‘here is a finding’ into ‘and I fixed it’ without crossing the same line everyone else crosses.” The tab is that line, drawn on screen.

There’s a human benefit too. When I switch to a tab named pr-reviews, my brain switches hats before I’ve read a word. Mode labels are cheap and they work.

A pane is one agent, one context window, one job

Every agent pane gets one job for the life of that conversation. When the job is done, the conversation ends, and the next job gets a fresh agent. Context windows are big now, but they aren’t free and they aren’t clean. Every stale plan, abandoned approach, and “actually, ignore that” sits in there and pulls on the next answer. A pane per job is the cheapest context hygiene there is.

The panes next to an agent are there for verification, not decoration:

  • In domain-modeling, the glossary and domain notes sit beside the agent, so I can check that it’s using the words we agreed on.
  • In implementation, a diffstat and a shell sit beside the agents, so I can see what actually changed, not what the agent says changed.
  • In pr-reviews, the raw diff sits beside the reviewer, so every file:line it cites can be checked by moving my eyes and nothing else.

Trust, but keep the evidence in the same field of view.

Tall for agents, short for shells, and why 32:9 matters

Agents emit long, structured output (plans, findings, tables), so they get full-height vertical splits. Shells and watchers are glanceable, so they get short horizontal stacks in a utility column. That isn’t a style choice. It’s matching pane shape to how much reading each pane needs.

The ultrawide makes it work without compromise. At 5120×1440 I get three or four full-height agent columns and the utility column and the sidebar, with nothing hidden behind a zoom or another tab. That matters for one reason above all: a blocked agent is a cost that accrues silently. Herdr’s sidebar tells me that something is blocked. Having it physically on screen tells me why, in peripheral vision, without moving. The less effort it takes to notice, the shorter the stall.

Part 2: Keeping Claude, Codex, and Cursor effective and accurate

All of the above is plumbing. The water is language.

I’ve been on this soapbox a while. In Precision in Words, Precision in Code: The Power of Writing in Modern Development I argued that “Writing isn’t just ‘extra work’—it’s the work that clarifies, simplifies, and accelerates everything else,” and in AI Prompt Engineering: Mastering Language Constructs I went into how the actual constructs, imperatives, conditionals, and contextual markers, shape what comes back. With agents that act, not just answer, the stakes go up. A vague word in a chat gets you a vague paragraph. A vague word in an agent prompt gets you a nine-file diff.

Here’s what I actually do. Every example below is a real prompt and a real response, captured for the screenshots in this repo with Claude Code in plan mode against throwaway demo projects.

1. Use the domain’s words, and ban the overloaded ones

Accurate language starts with the same word meaning the same thing every time: in the docs, in the code, in the prompt, in the review. That’s just ubiquitous language from domain-driven design, and agents need it more than humans do, because they fill every gap with the most statistically likely meaning, which is rarely your meaning.

So I keep a glossary in the notes repo (more on that in Notes Repos as Shared Project Context, where the glossary is described as “cheap insurance against confident mistakes”), and I put the words in the prompt, including the ones that are off limits:

Read docs/domain.md and internal/ledger/entry.go. Using only the domain terms Account, Entry, Journal, and Posting (never ‘transaction’), describe a Posting aggregate: its invariants, what it owns, and what it must reject. Do not write code. Output a short markdown model.

“Transaction” is banned because in a ledger it’s hopelessly overloaded: a database transaction, a business event, a bank-statement line. Pick one meaning, give it its own word, and forbid the ambiguous one. Here’s what came back:

Two things to notice. The model sticks to Account, Entry, Journal, Posting the whole way through. And it marks which rules it inferred from general double-entry practice versus which came from the repo: “Rules marked (inferred) come from standard double-entry practice, not the repo.” That distinction is gold in domain modeling, and it showed up because the prompt was explicit about where the truth lives (those two files).

2. Lead with scope, as the first word

My implementation prompts literally start with the word Scope:.

Scope: cmd/main.go only. Plan a minimal CLI entrypoint that loads a Journal from a JSON file and prints the balance per Account. Do not touch internal/ledger. Keep the planned diff under 50 lines.

Scope first, because everything after it gets read in its light. Then a positive fence (only this file), a negative fence (not that package), and a size limit stated as a number, not a vibe. “Keep it small” means nothing. “Under 50 lines” is checkable. I went deeper on why scope and diff size are the main levers in Hurting or Helping Devs?.

The payoff:

The repo had no go.mod, so cmd/main.go couldn’t import internal/ledger. An unfenced agent “fixes” that by adding a module file, rewiring imports, and maybe stubbing a Journal type in the package I’d fenced off. This one stopped and asked, offered the in-scope option as the recommendation, and named the out-of-scope option as out of scope: “This touches a second file, which goes beyond the cmd/main.go-only scope.” Herdr flagged it blocked, and I could see it on the ultrawide without switching anything. That’s the whole system working: precise words create the stop, and the layout makes the stop cheap.

3. Pick verbs that mean exactly one thing

plan, implement, review, and fix are different verbs, and I never let one stand in for another.

  • Plan means produce a plan and change nothing. I back it with tooling: agents start in plan mode (claude --permission-mode plan; Codex and Cursor have their own read-only and ask-first modes), so the verb and the permission agree.
  • Implement means change things inside the approved plan’s scope, and only after an explicit yes. Approval is scoped to the plan that was approved. This is the three-beat gate from Ask, Plan, Confirm.
  • Review means read and report. It doesn’t mean fix.
  • Fix means a new implementation task, with its own scope, in the implementation tab.

When the verb is precise, the agent’s job is precise, and so is my evaluation of whether it did the job.

4. Specify the shape of the answer

If I can’t check an answer quickly, I’ll check it badly. So I specify the output format in a way that’s easy to verify against the pane next door:

Review the diff main…feature/posting-aggregate. Classify every change as additive, mutative, or destructive. Flag any naming that drifts from docs/domain.md. Report findings as a numbered list with file:line. Do not propose rewrites outside the diff.

“Additive, mutative, or destructive” isn’t decoration either. It’s the vocabulary from Additive vs. Mutative vs. Destructive Code Changes (and Why AI Agents Love the Wrong One at 2:13AM), and giving the reviewer that vocabulary makes it classify instead of just commenting.

And because the question was precise, the reviewer had attention left over for what mattered. Under “side notes” it flagged that any Direction other than the exact string "debit", including "" or "Debit", gets counted as a credit, and that an empty Posting passes validation as balanced. Neither is a naming issue. Both are exactly the kind of quiet behavioral drift that post warned about: code that looks right and behaves like an alien artifact. A vague “please review this PR” tends to bury that under eleven style nits.

The same pattern works for any review lens. On the storefront demo, “integer cents only, no floating point until display formatting. Numbered findings with file:line, severity first” produced a short, ranked list I could verify in about thirty seconds.

5. Front-load the context instead of looping for it

A thin prompt plus five rounds of correction is worse than one well-fed prompt, and more expensive. I made that argument at length in “Loop Engineering” Is Mostly Just Broken SDLC Wearing a Costume: “A well-fed single pass beats a starved five-pass loop most of the time.” In practice that means every prompt names where the truth lives: the files to read, the glossary, the plan in the notes repo, the decision that was already made and shouldn’t be re-litigated.

It also means the environment has to be ready. An agent that can’t build, test, or reach the systems it needs will improvise, and improvisation is the opposite of accuracy. The checklist I use for that is in Agent-Ready Local Environments and MCP: “An agent with broad access and no rules will eventually do something technically impressive and socially awful.”

6. Same template, any agent

Claude Code, Codex, and Cursor have different personalities and different defaults, but the prompt shape carries across them. Herdr doesn’t care which one is in a pane. It detects them all, so I can put a different tool in the reviewer pane than the one that wrote the code. A second model has different blind spots, and that’s the point of a second opinion. The template:

Scope: <path(s)> only. <Do not touch: path(s)>.
Context: Read <files / glossary / plan>. Use the terms <A, B, C>; never <overloaded term>.
Task (<verb: plan | implement | review>): <one sentence, one job>.
Constraints: <additive only | under N lines | no new dependencies | tests first>.
Output: <numbered list | file:line | severity first | markdown model | test names first>.
If anything above is ambiguous or blocks you, stop and ask.

That last line matters most. It turns “the agent is stuck” from a silent failure into a blocked state, which Herdr turns into a dot I can see from across the room.

7. When an agent goes blocked or done

The loop on my side is short and boring, on purpose:

  1. Blocked: read the question in place (it’s already on screen), answer it, and if the answer is a decision, write it into the notes repo’s decisions file so the next session doesn’t ask again.
  2. Done: read the output against the evidence pane next to it. If it’s a plan, approve it or correct it. If it’s findings, open a scoped implementation task in the other tab for anything worth fixing.
  3. Finished job: end that agent’s conversation. New job, new agent, clean context.

The point of all of it

The spaces, tabs, panes, and ultrawide exist to keep context small, boundaries visible, and stalls cheap. The prompting exists to keep meaning precise. Neither works without the other. A gorgeous Herdr layout full of agents with vague instructions is just a very wide way to watch things go wrong in parallel, and the most precise prompt in the world doesn’t help if the agent has been blocked on it since lunch and you never noticed.

Draw the boundaries on screen. Draw them again in words. Then let the agents work.

References

Herdr

Previously on Composite Code

Elsewhere

Herding Agents: How I Actually Use Herdr (on a Ridiculously Wide Screen)

One coding agent is a tool. Two is a pair. Six is a zoo, and somebody has to know which animal is chewing on the fence.

For most of this year that somebody was me, armed with a pile of terminal windows, a couple of tmux sessions I half remembered setting up, and the nagging feeling that some agent somewhere had been politely waiting on a yes/no question for twenty minutes. Agents don’t tap you on the shoulder. They sit there, cursor blinking, and burn the most expensive thing in the whole setup, which is wall-clock time.

Herdr fixed most of that for me. This post covers what it is, how it works, and how I lay it out day to day. There’s a companion post (TBD) on why I split things up the way I do and how I prompt inside the panes. This one is the tour.

A coding interface displaying a script on the left side and an output console on the right, with a dark theme and multiple lines of code written in a programming language.

What Herdr is

Herdr bills itself as “the runtime your coding agents live on.” Less marketing, more mechanics: it’s a terminal workspace manager built for AI coding agents. If you’ve used tmux, you already have about 70% of the mental model. It’s:

  • A single Rust binary. No Electron and no browser tab pretending to be a terminal. brew install herdr, or curl -fsSL https://herdr.dev/install.sh | sh, and you’re done.
  • Something that runs inside your terminal. iTerm, Ghostty, WezTerm, kitty, Alacritty, even inside tmux if you enjoy nesting dolls. It doesn’t replace your terminal. It organizes what’s in it.
  • A client/server pair. A background server owns the terminal sessions. The thing you look at is just a client. Close the window, detach, drop SSH, and the agents keep working. Reattach and they’re right where you left them. (I learned this the fun way while building the screenshots for this post: the demo window got closed out from under me mid-run, and every agent was still sitting there when I reattached.)
  • Agent-aware. This is the actual point. Herdr detects Claude Code, Codex, Cursor’s agent, Copilot CLI, OpenCode, Devin and a pile of others, then classifies each one’s state so you don’t have to go look.

The model: session › space › tab › pane

Herdr’s concepts page lays out four nouns. The precision matters, so here they are the way I think about them.

NounWhat Herdr says it isWhat I use it for
SessionA persistent server namespace. herdr attaches to the default one; herdr --session <name> gives you an isolated one.My one long-lived “everything” session, plus throwaway named sessions when I’m experimenting (like the demo session for this post’s screenshots).
Workspace (the sidebar calls them spaces)“The top-level project container. Use one workspace per repo, task, or investigation.”One space per folder I’m working in. Full stop.
Tab“A layout inside a workspace.”One tab per kind of work: pr-reviews, domain-modeling, implementation.
Pane“A real terminal.”One agent, or one shell, or one thing I’m watching. Arranged in vertical or horizontal splits.

Then there are the agent states, which make the whole thing worth it:

  • Blocked: the agent needs input, approval, or a decision. This is the one you care about.
  • Working: it’s chewing.
  • Done: it finished and you haven’t looked yet.
  • Idle: it finished and you have looked.
  • Unknown: Herdr can’t tell, usually because it’s a plain shell.

Those states roll up into the sidebar. Each space gets a dot, and there’s an agents list underneath that shows every agent across every space. One glance and I know who’s blocked, who’s working, and who’s done and quietly waiting for me to read their output.

That’s the shoulder tap agents never gave me.

How it works (enough to be dangerous)

Keyboard

The default prefix is ctrl+b, deliberately tmux-flavored, and the keymap is “prefix-first” so Herdr doesn’t steal keystrokes from your shell, editor, or the agent. The ones I actually use:

KeysDoes
prefix vSplit vertical: new pane to the right, side by side
prefix -Split horizontal: new pane below, stacked
prefix h / j / k / lMove focus left / down / up / right
prefix zZoom the current pane (and again to unzoom)
prefix rResize mode
prefix cNew tab
prefix shift+tRename tab
prefix 1…9Jump to tab N
prefix shift+nNew workspace (space)
prefix wWorkspace picker
prefix bToggle the sidebar
prefix qDetach (everything keeps running)
prefix ?Help, for when you forget all of the above

Run herdr --default-config to see the full annotated config, including how to remap any of it. Mouse works too: click, drag, split. I won’t judge.

The CLI and socket API

Here’s where it stops being “tmux with a status light.” Everything you can do with a keystroke you can also do from the command line over Herdr’s socket API. That includes the agents themselves, since they’re running in panes Herdr owns. Every screenshot in this post came from a layout I scripted like this, in an isolated session so my real one never got touched:

# a space per folder
herdr workspace create --cwd ~/Codez/ledger-service --label ledger-service

# name the first tab for the kind of work happening in it
herdr tab rename w3:t1 domain-modeling

# agent on the left, reference material on the right, split again downward
herdr pane split w3:p1 --direction right --ratio 0.5
herdr pane split w3:p2 --direction down  --ratio 0.5

# start an agent and hand it a scoped prompt
herdr pane run w3:p1 "claude --permission-mode plan"
herdr agent prompt w3:p1 "Read docs/domain.md and internal/ledger/entry.go. Using only the domain terms Account, Entry, Journal, and Posting (never 'transaction'), describe a Posting aggregate..."

# and then, the good part
herdr agent list      # who is working / blocked / done
herdr agent wait w3:p1 --until blocked --until done   # block until it needs me
herdr pane read w3:p1 # read what it said, without switching to it

That’s enough rope for agents to spawn panes, prompt each other, and notice when a sibling is blocked. I’ll admit that’s both very cool and the setup for a future post titled something like “How My Agents Unionized.”

Integrations

herdr integration install claude (or codex, copilot, opencode, and so on) drops a small hook into each agent so it reports its lifecycle to Herdr directly, rather than Herdr inferring state by watching the screen. herdr integration status shows what’s wired up. Do this. The state dots get noticeably more accurate.

The rest of the toolbox

  • Worktrees: herdr worktree create makes a git worktree and opens it as its own space. It fits the one-branch-per-agent pattern I wrote about in Giving Every Agent Its Own Branch: Git Worktrees for Parallel AI Work.
  • Remote machines: herdr --remote <ssh-target> puts panes on other boxes in the same sidebar, and they reconnect independently. See Connecting machines.
  • Session state: layouts persist and restore. Processes don’t survive a server restart (nothing does, that’s physics), but the shape of your workspace does. See Session state.

How I actually use it

That’s the brochure. Here’s what my screen looks like.

One space per folder

Every folder I’m actively working in gets its own space, named after the folder. As I type this, my sidebar has ten of them: interlinedlist, then interlinedlist-ios, interlinedlist-macos-native, interlinedlist-windows-app, interlinedlist-android, interlinedlist-architecture-migration, a couple of client and side projects, and the repo this post lives in.

Is that a lot? Sure. Is it more than what’s actually in my head? No. It’s exactly what’s in my head, and that’s the point. The sidebar is a direct readout of my working memory. If it isn’t a folder I’m working in, it doesn’t get a space. When I stop working in a folder, the space goes away. My config.toml even sorts the agent list by space (agent_panel_sort = "spaces") so the agents group the same way my brain does.

It’s mostly for my own mental organization, and I’m fine with that. The agents don’t care what the sidebar looks like. I do.

One tab per kind of work

Inside a space, tabs aren’t “more room.” Tabs are modes. I name them for the kind of work being done:

  • domain-modeling: one agent, plus the glossary, domain notes, and recent history open next to it. The agent’s job is to think, not type.
  • implementation: one or more agents, each scoped to a slice of the code, with a diffstat or test runner and a plain shell beside them.
  • pr-reviews: a reviewing agent next to the actual diff, so I can check its claims against the code without switching windows.

Most spaces have just one tab, whichever mode that folder is in today. The busy ones have all three. Switching tabs is switching hats, and the tab label tells me which hat is on before I’ve read a single line.

One or more terminals, split vertically or horizontally

Within a tab I keep a simple rule: agents get tall columns, shells get short stacks.

Agents produce long, scrolling output: plans, numbered findings, diffs they want to show me. That wants height. So agents get vertical splits (prefix v), side by side, full height. Shells, test runners, git log, diffstats and log tails are glanceable. They get horizontal splits (prefix -) stacked in the right-hand column.

The ultrawide changes the math

Here’s the cheat. I run all of this on a 5120×1440, 32:9 ultrawide: two 2560×1440 monitors fused into one with no bezel down the middle. At normal terminal font sizes, that’s room for three or four full-height agent columns plus a stacked utility column and the sidebar, all visible at once, nobody squished.

On a regular 16:9 screen, that same implementation tab looks like this:

Workable. Tight. Add a third agent and you’re zooming panes in and out all day. On the ultrawide it’s this:

The difference isn’t really “more stuff.” It’s that nothing is hidden. When an agent goes blocked, I don’t need to find it. It’s already on screen, in my peripheral vision, with a red dot in the sidebar to back it up. Herdr tells me who needs attention. The ultrawide means I can see why without changing anything.

That screenshot is an agent I’d told to touch cmd/main.go only. It found the repo had no go.mod, which meant staying in scope required a decision. So it stopped, went blocked, and asked. Herdr flagged it, and I could see the question without switching anything. Total cost: one keystroke. The alternative, an agent “helpfully” adding a module file and rewiring imports across a package I’d explicitly fenced off, costs a lot more than one keystroke. Why that worked, and how I write prompts so it keeps working, is the next post (TBD).

The short version

  • Herdr is a terminal workspace manager for coding agents: one Rust binary, runs in your terminal, background server, agent-state detection.
  • Spaces are folders. One per folder you’re working in. It’s a mirror of your head, not a filing cabinet.
  • Tabs are modes of work. pr-reviews, domain-modeling, implementation.
  • Panes: agents tall, shells short.
  • The sidebar tells you who’s blocked. The ultrawide lets you see why without moving.

Install it, run herdr, hit ctrl+b ?, and give it a day. Worst case, you’ve got a nice tmux. Best case, you stop finding agents that have been patiently blocked since lunch.

References

Notes Repos as Shared Project Context

The first thing most people put in a new repository is code. That is fine. It is also incomplete.

When I start a real project these days, something I expect Cursor, Claude, Codex, and I to grind on together for more than an afternoon, I create a notes repo early. Sometimes it sits beside the product repo. Sometimes it lives inside the product repo as a deliberate notes/ or docs/project/ tree. Either way, the job is the same: give the humans and the agents one durable place to put the stuff that does not belong in source files but absolutely belongs in the work.

This is not documentation for show. It is shared working memory.

Why a notes repo, not a pile of chat

Chat is ephemeral. Agent transcripts are useful, but they are not a system of record. Plans get buried. Decisions get restated wrong three sessions later. Somebody new (human or model) shows up and has to rediscover who owns what, what “done” means, and which external systems are in play.

A notes repo fixes that by making the context something you can open and search. Cursor can open it. Claude can read it. Codex can search it. I can edit it when the plan changes. Everyone is looking at the same markdown, not reconstructing the project from memory and half-remembered Slack threads.

I treat it as the base layer under the code. The product repo holds the build. The notes repo holds the thinking that keeps the build pointed in the right direction.

What actually goes in it

I keep the structure boring on purpose. Fancy wikis rot. Flat, named markdown files get used.

Project brief. What we are building, who it is for, what it is not. One page. If you cannot say it in a page, you do not have a project yet. You have a mood.

Current plan. The active plan, not a museum of every plan that ever existed. Short enough to approve in one breath. Specific enough that “yes” means something. When the plan changes, rewrite it. Do not append a novel of abandoned approaches unless those approaches teach a real constraint.

Decisions. Lightweight ADRs, or just dated notes: we chose X because Y, and Z is explicitly out of scope. Agents love to re-litigate settled choices. Write the settlement down.

Team map. Who is who. Humans, roles, ownership. Which agent lanes exist if you have specialized agents. Who reviews security. Who can approve schema changes. “The team” includes the tools now, so name them and their jobs.

Glossary. Product words, domain words, the three acronyms everybody uses differently. Cheap insurance against confident mistakes.

Working notes. Scratch space for the current thread of work: open questions, links to tickets, sketches, “we tried this and it failed because…”. This is the joint notebook Cursor, Claude, Codex, and I write into while we work.

Pointers out. Links to the product repos, design files, staging URLs, runbooks, and the environment/setup notes. The notes repo should know where the rest of the world lives without trying to duplicate it.

I do not put secrets here. I put the names of secrets and where they are supposed to live. That distinction matters.

How I use it with agents

The pattern is simple: before an agent starts implementing, it reads the brief, the current plan, and the relevant decision notes. After a meaningful chunk of work, it updates the working notes or the plan so the next session does not start from zero.

That pairs cleanly with the ask-plan-confirm habit I already use. The notes repo is where the approved plan lives between sessions. The agent does not get to invent a new goal because the chat scrolled away.

A few practical rules that hold up:

  • One current plan file. Archive old plans if you must, but do not make the agent guess which file is live.
  • Prefer short files with clear names over one giant NOTES.md.
  • Write for the next reader who was not in the room. That reader might be you in two weeks, or an agent with a fresh context window.
  • Update notes as part of the work, not as a cleanup chore after merge. If it is optional, it will not happen.

What this saves

It saves re-explaining the project every morning. It saves the “wait, who owns auth?” loop. It saves agents from optimizing the wrong goal because the real goal lived in a Slack thread from Thursday.

It also saves me from being the only continuity process on the team. Continuity is a file. Version it. Diff it. Argue with it in a pull request if the decision is big enough.

Start smaller than you think

You do not need a knowledge management platform. You need a repo, a README that says what the repo is for, and a handful of markdown files that stay honest.

Create it when the project becomes real. Keep it next to the code in your mental model. Make every agent treat it as required reading before they touch the product tree.

Code is the artifact. Notes are the shared brain that keeps the artifact from wandering off.

Toggle Switches When the Thing Behind the Switch Is a Whole System

There is a version of a feature toggle that looks wonderfully simple in a pull request:

if (features.newThing) {
doNewThing();
}

Then newThing turns out to be a data pipeline, a shipping promise, or a screen people have already started using. Now the switch has an owner, a scope, a failure mode, and a date when we ought to remove it. The if statement is the least interesting part.

I want to work through three concrete examples. The first moves a flat-file order feed through Databricks, then cuts a tenant over to normalized PostgreSQL behind a data API. The second turns on expedited shipping, with the same decision implemented in TypeScript and Go. The third releases a saved-filters interface in React and SwiftUI. Each is a different kind of switch, and each can surprise you if you treat it as a Boolean sprinkled through the codebase.

The examples are deliberately small enough to read, but the boundaries are real: immutable inputs, idempotency, tenant scoping, stable rollout decisions, and the difference between hiding a button and actually controlling a capability. Let’s get into it.

1. The order feed: Databricks today, PostgreSQL data API tomorrow

Imagine a partner dropping one immutable CSV file per batch into object storage. Each row is an item on an order:

order_id,customer_id,customer_name,sku,quantity,event_at
o-100,c-7,Ada,rail-pass,2,2026-09-24T10:00:00Z
o-100,c-7,Ada,seat-upgrade,1,2026-09-24T10:00:00Z
o-101,c-8,Grace,rail-pass,1,2026-09-24T10:01:00Z

The existing path loads the file into a Databricks Delta bronze table, then produces a current-order table for queries. The proposed path reads that same object, validates it, and sends a complete batch to a private data API. The API writes three normalized PostgreSQL tables: customers, orders, and order_items.

immutable CSV in object storage
|
batch worker
|
tenant route snapshot
/ \
Databricks PostgreSQL data API
COPY INTO validate + transaction
bronze/curated customers/orders/items
\ /
read adapter for the tenant

The switch is a tenant route, not a random choice made separately for each row. For a given batch, resolve the route once and put it in the batch log. Changing the route while a file is half processed would create an impressively confusing incident.

The route and the file contract

I keep the control-plane data separate from the pipeline code. In this example the configuration is a checked, versioned snapshot supplied to the worker; in production it might come from a flag service. The important bit is that the worker receives one decision and uses it for the whole operation.

// pipeline/route.ts
export type PipelinePath = "databricks" | "postgres";
export interface RouteSnapshot {
version: number;
defaultPath: PipelinePath;
tenants: Record<string, PipelinePath>;
}
export interface BatchRef {
tenantId: string;
batchId: string;
fileName: string; // a basename under the tenant's immutable S3 prefix
sha256: string; // digest of the file's bytes, recorded when uploaded
}
export function choosePath(
config: RouteSnapshot,
tenantId: string,
): { path: PipelinePath; configVersion: number } {
return {
path: config.tenants[tenantId] ?? config.defaultPath,
configVersion: config.version,
};
}
export function checkBatchRef(batch: BatchRef): void {
if (!/^[a-z0-9-]{1,64}$/.test(batch.tenantId) ||
!/^[a-zA-Z0-9-]{1,100}$/.test(batch.batchId) ||
!/^[a-zA-Z0-9-]+\.csv$/.test(batch.fileName) ||
!/^[a-f0-9]{64}$/.test(batch.sha256)) {
throw new Error("Invalid batch reference");
}
}

The file key is a basename by design. The worker constructs the object path from a fixed bucket and tenant prefix. That keeps a caller from turning a batch submission into “please read whatever URL I hand you.” The SHA-256 is about replay identity: batch-42 with new bytes is an error, not a cute way to overwrite history.

The existing Databricks path

Here is the setup on the lakehouse side. The external location and warehouse already have access to the bucket. I’m showing Databricks SQL because COPY INTO is a good fit for an incremental feed of files, and the loaded-file tracking makes retries tractable. If your feed is millions of files, Databricks points you toward Auto Loader instead.

CREATE TABLE IF NOT EXISTS main.orders.order_rows_bronze (
order_id STRING,
customer_id STRING,
customer_name STRING,
sku STRING,
quantity STRING,
event_at STRING
) USING DELTA;
CREATE TABLE IF NOT EXISTS main.orders.order_rows_current (
order_id STRING,
customer_id STRING,
customer_name STRING,
sku STRING,
quantity INT,
event_at TIMESTAMP
) USING DELTA;

The TypeScript adapter submits a SQL statement to a warehouse and waits for completion. For this example every query returns either no rows or a small lookup result; large query results need the API’s external-links disposition and a separate paging design.

// pipeline/databricks.ts
type StatementResult = {
statement_id?: string;
status: { state: string; error?: { message: string } };
result?: { data_array?: string[][] };
};
export class DatabricksSql {
constructor(
private readonly host: string,
private readonly token: string,
private readonly warehouseId: string,
) {}
private async request(path: string, init?: RequestInit): Promise<StatementResult> {
const response = await fetch(`${this.host}${path}`, {
...init,
headers: {
Authorization: `Bearer ${this.token}`,
"Content-Type": "application/json",
...init?.headers,
},
});
if (!response.ok) throw new Error(`Databricks HTTP ${response.status}`);
return response.json() as Promise<StatementResult>;
}
async execute(statement: string, parameters: { name: string; value: string }[] = []) {
let result = await this.request("/api/2.0/sql/statements", {
method: "POST",
body: JSON.stringify({
warehouse_id: this.warehouseId,
statement,
parameters,
wait_timeout: "10s",
disposition: "INLINE",
format: "JSON_ARRAY",
}),
});
const deadline = Date.now() + 120_000;
while (result.status.state === "PENDING" || result.status.state === "RUNNING") {
if (!result.statement_id || Date.now() > deadline) {
throw new Error("Databricks statement exceeded worker deadline");
}
await new Promise(resolve => setTimeout(resolve, 1_000));
result = await this.request(`/api/2.0/sql/statements/${result.statement_id}`);
}
if (result.status.state !== "SUCCEEDED") {
throw new Error(result.status.error?.message ?? `Statement ${result.status.state}`);
}
return result.result?.data_array ?? [];
}
}
export async function loadDatabricksBatch(
sql: DatabricksSql,
batch: BatchRef,
): Promise<void> {
checkBatchRef(batch);
const prefix = `s3://example-order-feed/${batch.tenantId}/`;
// Only the basename is interpolated; checkBatchRef restricts its alphabet.
await sql.execute(`
COPY INTO main.orders.order_rows_bronze
FROM '${prefix}'
FILEFORMAT = CSV
FILES = ('${batch.fileName}')
FORMAT_OPTIONS ('header' = 'true')
`);
// The source file is a complete order snapshot. Deduplicate source rows
// before MERGE: two matching source rows for one target key are ambiguous.
await sql.execute(`
MERGE INTO main.orders.order_rows_current AS target
USING (
SELECT order_id, customer_id, customer_name, sku,
CAST(quantity AS INT) AS quantity,
CAST(event_at AS TIMESTAMP) AS event_at
FROM main.orders.order_rows_bronze
QUALIFY ROW_NUMBER() OVER (
PARTITION BY order_id, sku ORDER BY CAST(event_at AS TIMESTAMP) DESC
) = 1
) AS source
ON target.order_id = source.order_id AND target.sku = source.sku
WHEN MATCHED AND source.event_at >= target.event_at THEN UPDATE SET
customer_id = source.customer_id,
customer_name = source.customer_name,
quantity = source.quantity,
event_at = source.event_at
WHEN NOT MATCHED THEN INSERT
(order_id, customer_id, customer_name, sku, quantity, event_at)
VALUES (source.order_id, source.customer_id, source.customer_name,
source.sku, source.quantity, source.event_at)
`);
}

There is a deliberate simplification here: the merge scans bronze and models upserts, not item deletion. If a later snapshot can remove an item, the source contract needs tombstones or a replace-whole-order operation. No toggle solves an undefined deletion contract. Also, in a real lakehouse I would key these tables by tenant, or put each tenant in its own governed schema; the sample’s order_id must be globally unique for the shown merge.

The PostgreSQL path and its data API

On the new path, the database is an implementation detail of the data API. The worker never gets a PostgreSQL connection string. A private, authenticated service receives a complete batch; the service owns validation, idempotency, and the transaction.

CREATE TABLE customers (
tenant_id text NOT NULL,
customer_id text NOT NULL,
customer_name text NOT NULL,
PRIMARY KEY (tenant_id, customer_id)
);
CREATE TABLE orders (
tenant_id text NOT NULL,
order_id text NOT NULL,
customer_id text NOT NULL,
event_at timestamptz NOT NULL,
PRIMARY KEY (tenant_id, order_id),
FOREIGN KEY (tenant_id, customer_id)
REFERENCES customers (tenant_id, customer_id)
);
CREATE TABLE order_items (
tenant_id text NOT NULL,
order_id text NOT NULL,
sku text NOT NULL,
quantity integer NOT NULL CHECK (quantity > 0),
PRIMARY KEY (tenant_id, order_id, sku),
FOREIGN KEY (tenant_id, order_id)
REFERENCES orders (tenant_id, order_id) ON DELETE CASCADE
);
CREATE TABLE ingestion_batches (
tenant_id text NOT NULL,
batch_id text NOT NULL,
sha256 char(64) NOT NULL,
committed_at timestamptz NOT NULL DEFAULT now(),
PRIMARY KEY (tenant_id, batch_id)
);

The flat file repeats customer names and order IDs. The relational model stores each customer and order once, then each item under that order. This normalization is useful for operational queries and constraints; it is not a claim that PostgreSQL is automatically the better analytics engine. The switch is about the workload we actually need to serve.

The worker downloads the immutable object and parses it. This uses @aws-sdk/client-s3 and csv-parse/sync. I cap the file at 10,000 rows here so the data API can handle one transaction; larger feeds should use bounded chunks with a manifest and a stronger commit protocol.

// pipeline/postgres-path.ts
import { S3Client, GetObjectCommand } from "@aws-sdk/client-s3";
import { parse } from "csv-parse/sync";
import { createHash } from "node:crypto";
export type OrderRow = {
order_id: string;
customer_id: string;
customer_name: string;
sku: string;
quantity: number;
event_at: string;
};
const s3 = new S3Client({});
export async function loadPostgresBatch(
batch: BatchRef,
apiBase: string,
serviceToken: string,
): Promise<void> {
checkBatchRef(batch);
const key = `${batch.tenantId}/${batch.fileName}`;
const object = await s3.send(new GetObjectCommand({
Bucket: "example-order-feed", Key: key,
}));
if (!object.Body) throw new Error("Empty object body");
const bytes = await object.Body.transformToByteArray();
const digest = createHash("sha256").update(bytes).digest("hex");
if (digest !== batch.sha256) throw new Error("Batch content changed");
const records = parse(Buffer.from(bytes), {
columns: true, skip_empty_lines: true, bom: true,
}) as Record<string, string>[];
if (records.length === 0 || records.length > 10_000) {
throw new Error("Batch size outside accepted range");
}
const rows: OrderRow[] = records.map(row => ({
order_id: row.order_id,
customer_id: row.customer_id,
customer_name: row.customer_name,
sku: row.sku,
quantity: Number(row.quantity),
event_at: row.event_at,
}));
const response = await fetch(`${apiBase}/v1/batches`, {
method: "POST",
headers: {
Authorization: `Bearer ${serviceToken}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ ...batch, rows }),
});
if (!response.ok) throw new Error(`Data API rejected batch: ${response.status}`);
}

The API side is where the transaction belongs. zod checks the wire shape; PostgreSQL constraints still carry the final integrity guarantee. The gateway authenticates the service token and supplies the tenant identity; it must verify that the tenant in the request body matches that identity. I’m leaving gateway wiring out of this excerpt, but I would not expose this route to the public internet as an unauthenticated import endpoint.

// data-api/batches.ts (Fastify + pg + zod)
import Fastify from "fastify";
import pg from "pg";
import { z } from "zod";
const pool = new pg.Pool({ connectionString: process.env.DATABASE_URL });
const app = Fastify();
const rowSchema = z.object({
order_id: z.string().min(1).max(100),
customer_id: z.string().min(1).max(100),
customer_name: z.string().min(1).max(200),
sku: z.string().min(1).max(100),
quantity: z.number().int().positive(),
event_at: z.string().datetime({ offset: true }),
});
const batchSchema = z.object({
tenantId: z.string().regex(/^[a-z0-9-]{1,64}$/),
batchId: z.string().min(1).max(100),
fileName: z.string().endsWith(".csv"),
sha256: z.string().regex(/^[a-f0-9]{64}$/),
rows: z.array(rowSchema).min(1).max(10_000),
});
app.post("/v1/batches", async (request, reply) => {
const parsed = batchSchema.safeParse(request.body);
if (!parsed.success) {
return reply.code(400).send({ error: "Invalid batch payload" });
}
const input = parsed.data;
// Gateway auth must bind this tenantId to the authenticated caller.
const byOrder = new Map<string, typeof input.rows>();
for (const row of input.rows) {
const group = byOrder.get(row.order_id) ?? [];
group.push(row);
byOrder.set(row.order_id, group);
}
for (const rows of byOrder.values()) {
const first = rows[0];
const skus = new Set<string>();
for (const row of rows) {
if (row.customer_id !== first.customer_id ||
row.customer_name !== first.customer_name ||
row.event_at !== first.event_at || skus.has(row.sku)) {
return reply.code(400).send({ error: "Inconsistent order snapshot" });
}
skus.add(row.sku);
}
}
const client = await pool.connect();
try {
await client.query("BEGIN");
const inserted = await client.query(
`INSERT INTO ingestion_batches (tenant_id, batch_id, sha256)
VALUES ($1, $2, $3) ON CONFLICT DO NOTHING RETURNING batch_id`,
[input.tenantId, input.batchId, input.sha256],
);
if (inserted.rowCount === 0) {
const existing = await client.query(
`SELECT sha256 FROM ingestion_batches
WHERE tenant_id = $1 AND batch_id = $2`,
[input.tenantId, input.batchId],
);
await client.query("COMMIT");
if (existing.rows[0]?.sha256 !== input.sha256) {
return reply.code(409).send({ error: "Batch ID reused with new bytes" });
}
return reply.send({ status: "already-committed" });
}
for (const [orderId, rows] of byOrder) {
const first = rows[0];
await client.query(
`INSERT INTO customers (tenant_id, customer_id, customer_name)
VALUES ($1, $2, $3)
ON CONFLICT (tenant_id, customer_id)
DO UPDATE SET customer_name = EXCLUDED.customer_name`,
[input.tenantId, first.customer_id, first.customer_name],
);
await client.query(
`INSERT INTO orders (tenant_id, order_id, customer_id, event_at)
VALUES ($1, $2, $3, $4)
ON CONFLICT (tenant_id, order_id) DO UPDATE SET
customer_id = EXCLUDED.customer_id,
event_at = EXCLUDED.event_at
WHERE orders.event_at <= EXCLUDED.event_at`,
[input.tenantId, orderId, first.customer_id, first.event_at],
);
const current = await client.query(
`SELECT event_at FROM orders WHERE tenant_id = $1 AND order_id = $2
FOR UPDATE`,
[input.tenantId, orderId],
);
if (new Date(current.rows[0].event_at).toISOString() !==
new Date(first.event_at).toISOString()) continue; // stale snapshot
await client.query(
`DELETE FROM order_items WHERE tenant_id = $1 AND order_id = $2`,
[input.tenantId, orderId],
);
for (const row of rows) {
await client.query(
`INSERT INTO order_items (tenant_id, order_id, sku, quantity)
VALUES ($1, $2, $3, $4)`,
[input.tenantId, orderId, row.sku, row.quantity],
);
}
}
await client.query("COMMIT");
return reply.send({ status: "committed", orders: byOrder.size });
} catch (error) {
await client.query("ROLLBACK");
throw error;
} finally {
client.release();
}
});

Notice the batch marker is inserted inside the same transaction as the rows. If the transaction fails, the marker disappears too. Retrying the file then does useful work instead of being mistaken for a success. The FOR UPDATE also serializes replacement of one order’s items. For exact ties in event_at, the upstream contract should supply a monotonically increasing revision; timestamps alone cannot tell which different snapshot wins.

The read half of the data API keeps PostgreSQL behind the boundary too:

app.get<{ Params: { orderId: string } }>(
"/v1/orders/:orderId",
async (request, reply) => {
// tenantId comes from authenticated gateway context, never a query param.
const tenantId = request.headers["x-verified-tenant-id"] as string;
const order = await pool.query(
`SELECT o.order_id, o.event_at, c.customer_id, c.customer_name
FROM orders o JOIN customers c
ON c.tenant_id = o.tenant_id AND c.customer_id = o.customer_id
WHERE o.tenant_id = $1 AND o.order_id = $2`,
[tenantId, request.params.orderId],
);
if (order.rowCount === 0) return reply.code(404).send();
const items = await pool.query(
`SELECT sku, quantity FROM order_items
WHERE tenant_id = $1 AND order_id = $2 ORDER BY sku`,
[tenantId, request.params.orderId],
);
return { ...order.rows[0], items: items.rows };
},
);

That x-verified-tenant-id header is only safe if a trusted gateway strips any client-supplied copy and injects its own value. In a direct deployment, put authentication and tenant extraction in a Fastify hook instead. The point is to show where tenant ownership is enforced, because “the new API is private” is not an authorization strategy.

The consumer needs the same route decision. Here the Databricks query uses a named parameter and the PostgreSQL side uses the data API. The adapter presents one order shape to its caller. This example assumes the order ID is globally unique in Databricks; if it is only unique per tenant, add tenant_id to the Delta tables and both predicates.

// pipeline/read-order.ts
export type OrderView = {
orderId: string;
customerId: string;
customerName: string;
eventAt: string;
items: { sku: string; quantity: number }[];
};
export async function getOrder(
config: RouteSnapshot,
tenantId: string,
orderId: string,
deps: { databricks: DatabricksSql; apiBase: string; serviceToken: string },
): Promise<OrderView | null> {
if (choosePath(config, tenantId).path === "postgres") {
const response = await fetch(
`${deps.apiBase}/v1/orders/${encodeURIComponent(orderId)}`,
{ headers: { Authorization: `Bearer ${deps.serviceToken}` } },
);
if (response.status === 404) return null;
if (!response.ok) throw new Error(`Data API HTTP ${response.status}`);
const row = await response.json() as {
order_id: string; customer_id: string; customer_name: string;
event_at: string; items: { sku: string; quantity: number }[];
};
return {
orderId: row.order_id, customerId: row.customer_id,
customerName: row.customer_name, eventAt: row.event_at,
items: row.items,
};
}
const rows = await deps.databricks.execute(`
SELECT order_id, customer_id, customer_name, sku,
CAST(quantity AS STRING), CAST(event_at AS STRING)
FROM main.orders.order_rows_current WHERE order_id = :order_id
ORDER BY sku
`, [{ name: "order_id", value: orderId }]);
if (rows.length === 0) return null;
return {
orderId: rows[0][0], customerId: rows[0][1],
customerName: rows[0][2], eventAt: rows[0][5],
items: rows.map(row => ({ sku: row[3], quantity: Number(row[4]) })),
};
}

Finally, the worker uses its pinned route. Read traffic should use the tenant route only after the PostgreSQL side has been backfilled and checked.

// pipeline/worker.ts
export async function processBatch(
config: RouteSnapshot,
batch: BatchRef,
deps: { databricks: DatabricksSql; apiBase: string; serviceToken: string },
) {
checkBatchRef(batch);
const decision = choosePath(config, batch.tenantId);
// Persist {batchId, sha256, path, configVersion} in your job ledger.
if (decision.path === "databricks") {
await loadDatabricksBatch(deps.databricks, batch);
} else {
await loadPostgresBatch(batch, deps.apiBase, deps.serviceToken);
}
return decision;
}

My cutover order would be: add the new API and schema, backfill it from immutable files, compare order counts and sampled contents, run both paths in a controlled shadow period, then route one tenant’s writes and reads to PostgreSQL. Measure lag, API errors, rejected rows, and mismatches by tenant. Keep the original feed while rollback is needed. If writes have happened only on the new path, flipping the read switch back to Databricks can show stale data; rollback needs replay to the old path or a freeze until it catches up. That’s the part the tidy if statement never tells you.

2. Expedited shipping: a feature flag with a bill attached

Here’s a second example: turn on an expedited-shipping offer for a percentage of accounts. The new path does not merely paint a badge. It changes the checkout quote, so we need one stable decision per order and a record of the rule that produced the price.

The rule is: the account is included in the rollout, the destination is in the supported region, the cart subtotal is at least 7,500 cents, and the feature has not been killed globally. I use a stable FNV-1a hash of accountId for the rollout. It is small and portable across TypeScript and Go; it is not a security primitive. A production flag service can own the assignment instead, provided both stacks read the same assignment.

TypeScript implementation

// checkout/expedited.ts
export type ShippingFlag = {
enabled: boolean;
killSwitch: boolean;
rolloutPercent: number; // integer 0..100
revision: string;
};
export type QuoteInput = {
accountId: string;
subtotalCents: number;
region: "US-LOWER-48" | "US-OTHER" | "INTERNATIONAL";
};
export type ShippingQuote = {
standardCents: number;
expeditedCents: number | null;
decision: {
offered: boolean;
flagRevision: string;
reason: string;
};
};
export function bucket(accountId: string): number {
let hash = 0x811c9dc5;
for (const byte of new TextEncoder().encode(accountId)) {
hash ^= byte;
hash = Math.imul(hash, 0x01000193) >>> 0;
}
return hash % 100;
}
export function quoteShipping(input: QuoteInput, flag: ShippingFlag): ShippingQuote {
if (!Number.isSafeInteger(input.subtotalCents) || input.subtotalCents < 0 ||
!Number.isInteger(flag.rolloutPercent) ||
flag.rolloutPercent < 0 || flag.rolloutPercent > 100) {
throw new Error("Invalid quote input or flag configuration");
}
const reason = !flag.enabled || flag.killSwitch ? "disabled"
: bucket(input.accountId) >= flag.rolloutPercent ? "outside-rollout"
: input.region !== "US-LOWER-48" ? "unsupported-region"
: input.subtotalCents < 7_500 ? "below-minimum"
: "eligible";
return {
standardCents: 799,
expeditedCents: reason === "eligible" ? 1499 : null,
decision: {
offered: reason === "eligible",
flagRevision: flag.revision,
reason,
},
};
}
// At checkout creation, persist the quoted cents and decision with the order.
// At payment confirmation, charge that stored quote after its normal expiry
// and inventory checks. Do not ask the flag again halfway through checkout.

This is where teams often reach for Math.random() and accidentally give the same customer a different offer on every refresh. The bucket makes the rollout stable. The saved decision makes an in-progress checkout stable even if the operator moves from 10% to 20% while a customer is entering their card details.

The same boundary in Go

If another service computes a quote in Go, it must agree on byte encoding, hash, threshold, region names, and cents. This uses only the standard library.

package shipping
import (
"errors"
"hash/fnv"
)
type Flag struct {
Enabled bool
KillSwitch bool
RolloutPercent int
Revision string
}
type Input struct {
AccountID string
SubtotalCents int64
Region string
}
type Decision struct {
Offered bool
FlagRevision string
Reason string
}
type Quote struct {
StandardCents int64
ExpeditedCents *int64
Decision Decision
}
func Bucket(accountID string) int {
h := fnv.New32a()
_, _ = h.Write([]byte(accountID)) // UTF-8 bytes, as in TextEncoder
return int(h.Sum32() % 100)
}
func QuoteShipping(in Input, flag Flag) (Quote, error) {
if in.SubtotalCents < 0 || flag.RolloutPercent < 0 ||
flag.RolloutPercent > 100 {
return Quote{}, errors.New("invalid quote input or flag configuration")
}
reason := "eligible"
switch {
case !flag.Enabled || flag.KillSwitch:
reason = "disabled"
case Bucket(in.AccountID) >= flag.RolloutPercent:
reason = "outside-rollout"
case in.Region != "US-LOWER-48":
reason = "unsupported-region"
case in.SubtotalCents < 7500:
reason = "below-minimum"
}
result := Quote{
StandardCents: 799,
Decision: Decision{
Offered: reason == "eligible", FlagRevision: flag.Revision,
Reason: reason,
},
}
if result.Decision.Offered {
expedited := int64(1499)
result.ExpeditedCents = &expedited
}
return result, nil
}

I would put a shared set of fixture inputs in both test suites: account IDs with ASCII and non-ASCII characters, rollout at 0 and 100, subtotal at 7,499 and 7,500 cents, each region, and the kill switch. The tests should assert that TypeScript and Go assign the same account to the same bucket. This is one of those boring cross-language details that becomes a very exciting checkout bug if you skip it.

The rollout sequence is straightforward: ship both implementations dark, compare decisions against fixtures, enable internal accounts, then 1%, 10%, and onward while watching quote errors, conversion, fulfillment capacity, and customer support reports. The kill switch stops new offers. Existing orders retain their stored price and promise; changing those would be a different business operation. When the feature is permanent, remove the rollout branch and leave the eligibility rule as normal checkout code.

3. Saved filters: the interface switch and the user’s own toggle

The third example is a saved-filters panel for an order list. There are two switches here that people tend to conflate. The feature flag says whether this account has access to saved filters. The user preference says whether the available panel is currently shown. If the feature flag is off, the preference cannot manufacture access.

The server returns a capability document after authentication:

{
"savedFilters": true,
"revision": "ui-2026-09-24-3"
}

The endpoint that creates or lists saved filters must evaluate the account’s capability again. A React component that omits a button is a nicer screen, not an access control check. The same applies on iOS.

TypeScript and React

The browser fetches the capability, treats unknown as off, and lets the person hide or show the panel with a visible toggle. The preference is local to the browser in this version; if we want it to follow a user across devices, that becomes a server-side preference with its own API and migration.

// SavedFiltersFeature.tsx
import { useEffect, useState } from "react";
type Capabilities = { savedFilters: boolean; revision: string };
type Filter = { id: string; name: string; query: string };
export function SavedFiltersFeature() {
const [capability, setCapability] = useState<Capabilities | null>(null);
const [filters, setFilters] = useState<Filter[]>([]);
const [open, setOpen] = useState(
() => localStorage.getItem("saved-filters-open") === "true",
);
const [error, setError] = useState<string | null>(null);
useEffect(() => {
const controller = new AbortController();
fetch("/v1/me/features", { signal: controller.signal, credentials: "include" })
.then(response => {
if (!response.ok) throw new Error("Could not load features");
return response.json() as Promise<Capabilities>;
})
.then(setCapability)
.catch(err => {
if (err.name !== "AbortError") setError(err.message);
});
return () => controller.abort();
}, []);
useEffect(() => {
if (!capability?.savedFilters || !open) return;
const controller = new AbortController();
fetch("/v1/saved-filters", { signal: controller.signal, credentials: "include" })
.then(response => {
if (!response.ok) throw new Error("Could not load saved filters");
return response.json() as Promise<Filter[]>;
})
.then(setFilters)
.catch(err => {
if (err.name !== "AbortError") setError(err.message);
});
return () => controller.abort();
}, [capability?.savedFilters, open]);
if (!capability?.savedFilters) {
return error ? <p role="alert">{error}</p> : null;
}
return <section aria-label="Saved filters">
<label>
<input type="checkbox" checked={open} onChange={event => {
const next = event.target.checked;
setOpen(next);
localStorage.setItem("saved-filters-open", String(next));
}} />
Show saved filters
</label>
{error && <p role="alert">{error}</p>}
{open && <ul>{filters.map(filter =>
<li key={filter.id}><button type="button" onClick={() => {
// The order list owns applying filter.query, after validating its DSL.
window.dispatchEvent(new CustomEvent("apply-saved-filter", {
detail: { id: filter.id },
}));
}}>{filter.name}</button></li>,
)}</ul>}
</section>;
}

The button passes a filter ID, not arbitrary SQL from storage. The order list asks its own API to apply that ID under the current account. If the flag goes off while this page is open, a fresh capability fetch on navigation or a short-lived cache will remove the panel; the API check shuts off server access immediately.

The same feature in SwiftUI

On iOS, @AppStorage is the user preference. The capability still comes from the server. A small model loads it and, only when allowed and opened, loads the filters.

import SwiftUI
struct Capabilities: Decodable {
let savedFilters: Bool
let revision: String
}
struct SavedFilter: Decodable, Identifiable {
let id: String
let name: String
let query: String
}
@MainActor
final class SavedFiltersModel: ObservableObject {
@Published private(set) var capability: Capabilities?
@Published private(set) var filters: [SavedFilter] = []
@Published private(set) var errorMessage: String?
// apiBase and URLSession are injected so previews/tests can use fixtures.
private let apiBase: URL
private let session: URLSession
init(apiBase: URL, session: URLSession = .shared) {
self.apiBase = apiBase
self.session = session
}
func loadCapability() async {
do {
let (data, response) = try await session.data(
from: apiBase.appending(path: "v1/me/features"))
guard (response as? HTTPURLResponse)?.statusCode == 200 else {
throw URLError(.badServerResponse)
}
capability = try JSONDecoder().decode(Capabilities.self, from: data)
if capability?.savedFilters != true { filters = [] }
errorMessage = nil
} catch {
capability = nil // unknown is off
filters = []
errorMessage = "Features could not be loaded."
}
}
func loadFilters() async {
guard capability?.savedFilters == true else { return }
do {
let (data, response) = try await session.data(
from: apiBase.appending(path: "v1/saved-filters"))
guard (response as? HTTPURLResponse)?.statusCode == 200 else {
throw URLError(.badServerResponse)
}
filters = try JSONDecoder().decode([SavedFilter].self, from: data)
errorMessage = nil
} catch {
filters = []
errorMessage = "Saved filters could not be loaded."
}
}
}
struct SavedFiltersView: View {
@StateObject private var model: SavedFiltersModel
@AppStorage("savedFiltersOpen") private var isOpen = false
let applyFilter: (String) -> Void
init(apiBase: URL, applyFilter: @escaping (String) -> Void) {
_model = StateObject(wrappedValue: SavedFiltersModel(apiBase: apiBase))
self.applyFilter = applyFilter
}
var body: some View {
Group {
if model.capability?.savedFilters == true {
Section("Saved filters") {
Toggle("Show saved filters", isOn: $isOpen)
if isOpen {
ForEach(model.filters) { filter in
Button(filter.name) { applyFilter(filter.id) }
}
}
}
}
if let error = model.errorMessage {
Text(error).foregroundStyle(.red)
}
}
.task { await model.loadCapability() }
.task(id: isOpen && model.capability?.savedFilters == true) {
if isOpen { await model.loadFilters() }
}
}
}

The app’s authenticated URLSession would carry the user’s credentials; the sample leaves that wiring to the host app. When someone taps a filter, the host sends the ID to the order-list API rather than trusting the locally decoded query. I would test both clients with the same capability responses: on, off, failed request, and revocation after the screen has loaded. I would also test the API endpoint directly with the feature disabled. The server is the actual gate.

The switch has a lifecycle

Across these examples, I would keep a little record for every flag: owner, default, scope, rollout plan, telemetry, rollback behavior, and removal date. The pipeline route is scoped to a tenant and batch. The shipping offer is scoped to a stable account cohort and then frozen into an order quote. The interface flag is scoped to an authenticated account, while the visible on/off control is a separate user preference.

That distinction is the useful mental model. A toggle switch is an operational decision point, and the rest of the system has to agree on what was decided. Give it one boundary, persist decisions when they affect money or durable data, measure the new path, and remove the temporary branch after the migration is over. Otherwise the switch becomes one more permanent mystery in the codebase, which is a lousy reward for trying to ship safely.

References

Agent-Ready Local Environments and MCP

Getting a coding agent productive is less about the prompt and more about whether the machine in front of it can actually build, run, and reach the systems the work depends on.

Cursor, Claude, Codex, and the rest are fast when the local world is already honest: dependencies install, the app boots, tests can run, and the external surfaces (GitHub, GitLab, Slack, Outlook, Teams, whatever is in play) are reachable through clear, permissioned paths. When those pieces are missing, the agent spends its tokens rediscovering your laptop instead of doing the job.

This post is about making that setup cheap and repeatable.

The real prerequisite is a bootable story

Every repo should answer, in files the agent can read, a small set of questions:

  1. What do I need installed on this machine?
  2. How do I get from clone to a running local system?
  3. Which services does this project talk to, and which of those am I expected to use locally?
  4. Where do secrets live, and what is not allowed?
  5. How do I know the setup worked?

If those answers only exist in your head, the agent will invent something almost right. Almost right is expensive.

I keep this in the product repo as first-class docs and scripts: README, CONTRIBUTING, docs/setup, scripts/bootstrap, .env.example. When the work spans more than one codebase, I also mirror the “where the rest of the world lives” map in the project notes repo.

Make setup mechanical

Agents are excellent at following a checklist. They are mediocre at inferring your undocumented brew taps and tribal knowledge.

What works well right now:

A single bootstrap path. One script or documented sequence: install language runtimes, install dependencies, copy env templates, start local dependencies, run a smoke check. Prefer boring and explicit over clever.

Pinned, discoverable tooling. Version files (.nvmrc, .node-version, mise.toml, asdf configs, rust-toolchain, etc.) beat README poetry. The agent should be able to detect the toolchain without asking you which Node major you meant last quarter.

.env.example that matches reality. Every required variable named. No secret values. Short comments for where to get each one. If a variable is only for production, say so.

Smoke tests for the environment. A command that proves the local stack is alive: npm run doctor, make verify, a compose healthcheck, a migration status call. Green means “safe to start feature work.” Red means “fix the machine, not the feature.”

Dev containers or explicit host setup. Pick one and document it. Either the agent works inside a known container, or it works on the host with a known bootstrap. Mixing both without saying which is canonical creates two slightly broken worlds.

This is the same instinct as my old dev-setup-osx habit, updated for a world where the “new developer” might be an agent that showed up thirty seconds ago.

Tell the agent where the working systems are

Local build is only half the map. Most real work touches other systems: source hosts, chat, mail, issue trackers, cloud consoles, internal APIs.

Write down the topology.

  • Product repos and their remotes (GitHub, GitLab, both, mirrors).
  • Which environments exist: local, staging, production, preview apps.
  • Which human collaboration surfaces matter: Slack channels, Teams teams, Outlook lists, Linear/Jira projects.
  • Which of those the agent is allowed to read or write.
  • How authentication is supposed to happen for each.

Do not make the agent guess that “the deploy” means GitHub Actions in one repo and a GitLab pipeline in another. Put the pointers in markdown next to the code or in the notes repo. Ambiguity here turns into the wrong PR opened against the wrong remote, or a “fix” that never lands where humans look.

MCP is how agents reach the rest of the desk

Model Context Protocol servers are the practical bridge between the coding agent and the tools already on your desk. Used well, they turn “go check the thread / issue / inbox / pipeline” into a first-class action instead of a copy-paste scavenger hunt.

The useful pattern is not “connect everything.” It is “connect the systems this project actually depends on, with the least privilege that still helps.”

A sane MCP layout for project work often includes:

  • Source control: GitHub and/or GitLab for issues, PRs, checks, and releases.
  • Chat: Slack or Teams for the channels where decisions and unblockers live.
  • Mail/calendar: Outlook or similar when the work is genuinely gated on threads or meetings, not because it is fun to give an agent your inbox.
  • Project tracking: whatever holds the tickets, if that is not already the git host.
  • Docs and runbooks: if they live outside the repo and agents keep asking for them.

Wire these at the user or project level in Cursor (and equivalents elsewhere), then document in the repo which MCP servers are expected for this project and what they are for. An agent that knows “Teams is available for the eng channel, GitHub is available for PRs, Outlook is not in scope for this repo” wastes less time and takes fewer weird actions.

Authentication matters. Prefer the product’s normal OAuth/device flows. Keep tokens out of the notes repo and out of prompts. If a server needs auth, say so in setup docs and stop there. Do not paste credentials into markdown “for convenience.”

Boundaries beat cleverness

An agent with broad access and no rules will eventually do something technically impressive and socially awful: comment in the wrong channel, open a PR against the wrong fork, or dig through mail that was never part of the task.

Give it clear limits:

  • Read-only by default where write access is not required.
  • Explicit allow-lists of repos, channels, and projects.
  • Repo instructions that say when to use MCP versus when to stay local.
  • The same ask-plan-confirm gate you use for code changes when the action leaves the laptop (posting, labeling, merging, emailing).

Local environment setup gets the agent building. MCP gets the agent collaborating. Boundaries keep both from becoming a mess you have to unwind.

A minimal checklist I actually use

When I stand up a project for agent-assisted work, I want at least this:

  1. Clone, bootstrap, and smoke test documented and scripted.
  2. Toolchain versions pinned in-repo.
  3. .env.example complete; real secrets elsewhere.
  4. Notes/brief that name the remotes, environments, and collaboration surfaces.
  5. MCP servers configured for the systems in play, with purpose notes in the project docs.
  6. Clear write/read expectations for each integration.
  7. A first prompt that points the agent at setup docs before feature work.

None of that is fancy. All of it is what makes Cursor, Claude, Codex, and friends look brilliant on day one instead of lost in your PATH.

The goal is simple: when an agent sits down at the project, the local world boots, the surrounding systems are findable, and the rules of engagement are already written down.