Spaces, Tabs, Panes, and Precise Words: Why I Break Agent Work Out the Way I Do

In the previous post I toured Herdr and how I lay it out: one space per folder, one tab per kind of work, agents in tall columns, shells in short stacks, all of it spread across a 5120×1440 ultrawide. That’s the what. This one is the why, and then the part that actually decides whether any of it pays off: how I talk to the agents once they’re in their panes.

Short version: the layout isn’t decoration. It’s context management for two kinds of brains, mine and the models’, and the prompts are the same idea at a smaller scale. Every tool in this stack, Claude Code, Codex, Cursor, and the human, gets worse the more unrelated junk is in its head. So I keep the junk out, structurally first and then verbally.

Part 1: Why I break it out like this

Spaces are folders because the folder is the agent’s whole world

An agent’s working directory isn’t a detail. It’s the edge of its universe. Whatever it can read, grep, build, and break starts at that folder. So when I make one space per folder, I’m not organizing windows. I’m drawing the same boundary the agent already lives inside, and making it visible to me.

That buys three things:

  1. The sidebar is a readout of my working memory. Ten spaces means ten things in flight. If a folder isn’t in the sidebar, I’m not working on it. When I’m done, the space goes away. There’s no “misc” space and no “scratch” space, because “misc” is where context goes to die.
  2. Blast radius is legible. If something is going sideways in ledger-service, the dot next to ledger-service says so. Nothing in that space can be quietly editing storefront-web, because it doesn’t live there.
  3. It pairs with worktrees. When I need two implementation agents on the same repo at once, each gets its own worktree and its own branch, and in Herdr that’s just another folder, so another space (herdr worktree create does this directly). I made the full case in Giving Every Agent Its Own Branch: Git Worktrees for Parallel AI Work: “Run two coding agents against a single working tree and they will fight.” Spaces-as-folders keeps the referee’s view honest.

Tabs are modes because each mode is a different verb

I name tabs domain-modeling, implementation, and pr-reviews because those are three different jobs. Each one comes with a different posture from me and a different register of prompt to the agent.

TabThe verbMy postureWhat the agent is allowed to do
domain-modelingthinkSkeptical, SocraticRead, reason, write a model. No code. Plan mode.
implementationchangeDirective, fencedEdit inside an explicit scope, after an approved plan
pr-reviewsjudgeAdversarial, evidentiaryRead the diff, report findings. No fixing.

The failure mode I’m designing against is verb drift. You start a conversation reviewing a PR, you say “huh, good catch,” and three messages later the “reviewer” is rewriting the module, now with the confidence of something that just graded its own homework. Separate tabs, and separate agent sessions in them, make the verb change explicit. If I want a fix, I go to the implementation tab and ask for one there, with a scope. It’s the same rule I put on the read-only reviewer agents in Ask, Plan, Confirm: Making Agents Stop Before They Start: the reviewer doesn’t get to “slide from ‘here is a finding’ into ‘and I fixed it’ without crossing the same line everyone else crosses.” The tab is that line, drawn on screen.

There’s a human benefit too. When I switch to a tab named pr-reviews, my brain switches hats before I’ve read a word. Mode labels are cheap and they work.

A pane is one agent, one context window, one job

Every agent pane gets one job for the life of that conversation. When the job is done, the conversation ends, and the next job gets a fresh agent. Context windows are big now, but they aren’t free and they aren’t clean. Every stale plan, abandoned approach, and “actually, ignore that” sits in there and pulls on the next answer. A pane per job is the cheapest context hygiene there is.

The panes next to an agent are there for verification, not decoration:

  • In domain-modeling, the glossary and domain notes sit beside the agent, so I can check that it’s using the words we agreed on.
  • In implementation, a diffstat and a shell sit beside the agents, so I can see what actually changed, not what the agent says changed.
  • In pr-reviews, the raw diff sits beside the reviewer, so every file:line it cites can be checked by moving my eyes and nothing else.

Trust, but keep the evidence in the same field of view.

Tall for agents, short for shells, and why 32:9 matters

Agents emit long, structured output (plans, findings, tables), so they get full-height vertical splits. Shells and watchers are glanceable, so they get short horizontal stacks in a utility column. That isn’t a style choice. It’s matching pane shape to how much reading each pane needs.

The ultrawide makes it work without compromise. At 5120×1440 I get three or four full-height agent columns and the utility column and the sidebar, with nothing hidden behind a zoom or another tab. That matters for one reason above all: a blocked agent is a cost that accrues silently. Herdr’s sidebar tells me that something is blocked. Having it physically on screen tells me why, in peripheral vision, without moving. The less effort it takes to notice, the shorter the stall.

Part 2: Keeping Claude, Codex, and Cursor effective and accurate

All of the above is plumbing. The water is language.

I’ve been on this soapbox a while. In Precision in Words, Precision in Code: The Power of Writing in Modern Development I argued that “Writing isn’t just ‘extra work’—it’s the work that clarifies, simplifies, and accelerates everything else,” and in AI Prompt Engineering: Mastering Language Constructs I went into how the actual constructs, imperatives, conditionals, and contextual markers, shape what comes back. With agents that act, not just answer, the stakes go up. A vague word in a chat gets you a vague paragraph. A vague word in an agent prompt gets you a nine-file diff.

Here’s what I actually do. Every example below is a real prompt and a real response, captured for the screenshots in this repo with Claude Code in plan mode against throwaway demo projects.

1. Use the domain’s words, and ban the overloaded ones

Accurate language starts with the same word meaning the same thing every time: in the docs, in the code, in the prompt, in the review. That’s just ubiquitous language from domain-driven design, and agents need it more than humans do, because they fill every gap with the most statistically likely meaning, which is rarely your meaning.

So I keep a glossary in the notes repo (more on that in Notes Repos as Shared Project Context, where the glossary is described as “cheap insurance against confident mistakes”), and I put the words in the prompt, including the ones that are off limits:

Read docs/domain.md and internal/ledger/entry.go. Using only the domain terms Account, Entry, Journal, and Posting (never ‘transaction’), describe a Posting aggregate: its invariants, what it owns, and what it must reject. Do not write code. Output a short markdown model.

“Transaction” is banned because in a ledger it’s hopelessly overloaded: a database transaction, a business event, a bank-statement line. Pick one meaning, give it its own word, and forbid the ambiguous one. Here’s what came back:

Two things to notice. The model sticks to Account, Entry, Journal, Posting the whole way through. And it marks which rules it inferred from general double-entry practice versus which came from the repo: “Rules marked (inferred) come from standard double-entry practice, not the repo.” That distinction is gold in domain modeling, and it showed up because the prompt was explicit about where the truth lives (those two files).

2. Lead with scope, as the first word

My implementation prompts literally start with the word Scope:.

Scope: cmd/main.go only. Plan a minimal CLI entrypoint that loads a Journal from a JSON file and prints the balance per Account. Do not touch internal/ledger. Keep the planned diff under 50 lines.

Scope first, because everything after it gets read in its light. Then a positive fence (only this file), a negative fence (not that package), and a size limit stated as a number, not a vibe. “Keep it small” means nothing. “Under 50 lines” is checkable. I went deeper on why scope and diff size are the main levers in Hurting or Helping Devs?.

The payoff:

The repo had no go.mod, so cmd/main.go couldn’t import internal/ledger. An unfenced agent “fixes” that by adding a module file, rewiring imports, and maybe stubbing a Journal type in the package I’d fenced off. This one stopped and asked, offered the in-scope option as the recommendation, and named the out-of-scope option as out of scope: “This touches a second file, which goes beyond the cmd/main.go-only scope.” Herdr flagged it blocked, and I could see it on the ultrawide without switching anything. That’s the whole system working: precise words create the stop, and the layout makes the stop cheap.

3. Pick verbs that mean exactly one thing

plan, implement, review, and fix are different verbs, and I never let one stand in for another.

  • Plan means produce a plan and change nothing. I back it with tooling: agents start in plan mode (claude --permission-mode plan; Codex and Cursor have their own read-only and ask-first modes), so the verb and the permission agree.
  • Implement means change things inside the approved plan’s scope, and only after an explicit yes. Approval is scoped to the plan that was approved. This is the three-beat gate from Ask, Plan, Confirm.
  • Review means read and report. It doesn’t mean fix.
  • Fix means a new implementation task, with its own scope, in the implementation tab.

When the verb is precise, the agent’s job is precise, and so is my evaluation of whether it did the job.

4. Specify the shape of the answer

If I can’t check an answer quickly, I’ll check it badly. So I specify the output format in a way that’s easy to verify against the pane next door:

Review the diff main…feature/posting-aggregate. Classify every change as additive, mutative, or destructive. Flag any naming that drifts from docs/domain.md. Report findings as a numbered list with file:line. Do not propose rewrites outside the diff.

“Additive, mutative, or destructive” isn’t decoration either. It’s the vocabulary from Additive vs. Mutative vs. Destructive Code Changes (and Why AI Agents Love the Wrong One at 2:13AM), and giving the reviewer that vocabulary makes it classify instead of just commenting.

And because the question was precise, the reviewer had attention left over for what mattered. Under “side notes” it flagged that any Direction other than the exact string "debit", including "" or "Debit", gets counted as a credit, and that an empty Posting passes validation as balanced. Neither is a naming issue. Both are exactly the kind of quiet behavioral drift that post warned about: code that looks right and behaves like an alien artifact. A vague “please review this PR” tends to bury that under eleven style nits.

The same pattern works for any review lens. On the storefront demo, “integer cents only, no floating point until display formatting. Numbered findings with file:line, severity first” produced a short, ranked list I could verify in about thirty seconds.

5. Front-load the context instead of looping for it

A thin prompt plus five rounds of correction is worse than one well-fed prompt, and more expensive. I made that argument at length in “Loop Engineering” Is Mostly Just Broken SDLC Wearing a Costume: “A well-fed single pass beats a starved five-pass loop most of the time.” In practice that means every prompt names where the truth lives: the files to read, the glossary, the plan in the notes repo, the decision that was already made and shouldn’t be re-litigated.

It also means the environment has to be ready. An agent that can’t build, test, or reach the systems it needs will improvise, and improvisation is the opposite of accuracy. The checklist I use for that is in Agent-Ready Local Environments and MCP: “An agent with broad access and no rules will eventually do something technically impressive and socially awful.”

6. Same template, any agent

Claude Code, Codex, and Cursor have different personalities and different defaults, but the prompt shape carries across them. Herdr doesn’t care which one is in a pane. It detects them all, so I can put a different tool in the reviewer pane than the one that wrote the code. A second model has different blind spots, and that’s the point of a second opinion. The template:

Scope: <path(s)> only. <Do not touch: path(s)>.
Context: Read <files / glossary / plan>. Use the terms <A, B, C>; never <overloaded term>.
Task (<verb: plan | implement | review>): <one sentence, one job>.
Constraints: <additive only | under N lines | no new dependencies | tests first>.
Output: <numbered list | file:line | severity first | markdown model | test names first>.
If anything above is ambiguous or blocks you, stop and ask.

That last line matters most. It turns “the agent is stuck” from a silent failure into a blocked state, which Herdr turns into a dot I can see from across the room.

7. When an agent goes blocked or done

The loop on my side is short and boring, on purpose:

  1. Blocked: read the question in place (it’s already on screen), answer it, and if the answer is a decision, write it into the notes repo’s decisions file so the next session doesn’t ask again.
  2. Done: read the output against the evidence pane next to it. If it’s a plan, approve it or correct it. If it’s findings, open a scoped implementation task in the other tab for anything worth fixing.
  3. Finished job: end that agent’s conversation. New job, new agent, clean context.

The point of all of it

The spaces, tabs, panes, and ultrawide exist to keep context small, boundaries visible, and stalls cheap. The prompting exists to keep meaning precise. Neither works without the other. A gorgeous Herdr layout full of agents with vague instructions is just a very wide way to watch things go wrong in parallel, and the most precise prompt in the world doesn’t help if the agent has been blocked on it since lunch and you never noticed.

Draw the boundaries on screen. Draw them again in words. Then let the agents work.

References

Herdr

Previously on Composite Code

Elsewhere

Herding Agents: How I Actually Use Herdr (on a Ridiculously Wide Screen)

One coding agent is a tool. Two is a pair. Six is a zoo, and somebody has to know which animal is chewing on the fence.

For most of this year that somebody was me, armed with a pile of terminal windows, a couple of tmux sessions I half remembered setting up, and the nagging feeling that some agent somewhere had been politely waiting on a yes/no question for twenty minutes. Agents don’t tap you on the shoulder. They sit there, cursor blinking, and burn the most expensive thing in the whole setup, which is wall-clock time.

Herdr fixed most of that for me. This post covers what it is, how it works, and how I lay it out day to day. There’s a companion post (TBD) on why I split things up the way I do and how I prompt inside the panes. This one is the tour.

A coding interface displaying a script on the left side and an output console on the right, with a dark theme and multiple lines of code written in a programming language.

What Herdr is

Herdr bills itself as “the runtime your coding agents live on.” Less marketing, more mechanics: it’s a terminal workspace manager built for AI coding agents. If you’ve used tmux, you already have about 70% of the mental model. It’s:

  • A single Rust binary. No Electron and no browser tab pretending to be a terminal. brew install herdr, or curl -fsSL https://herdr.dev/install.sh | sh, and you’re done.
  • Something that runs inside your terminal. iTerm, Ghostty, WezTerm, kitty, Alacritty, even inside tmux if you enjoy nesting dolls. It doesn’t replace your terminal. It organizes what’s in it.
  • A client/server pair. A background server owns the terminal sessions. The thing you look at is just a client. Close the window, detach, drop SSH, and the agents keep working. Reattach and they’re right where you left them. (I learned this the fun way while building the screenshots for this post: the demo window got closed out from under me mid-run, and every agent was still sitting there when I reattached.)
  • Agent-aware. This is the actual point. Herdr detects Claude Code, Codex, Cursor’s agent, Copilot CLI, OpenCode, Devin and a pile of others, then classifies each one’s state so you don’t have to go look.

The model: session › space › tab › pane

Herdr’s concepts page lays out four nouns. The precision matters, so here they are the way I think about them.

NounWhat Herdr says it isWhat I use it for
SessionA persistent server namespace. herdr attaches to the default one; herdr --session <name> gives you an isolated one.My one long-lived “everything” session, plus throwaway named sessions when I’m experimenting (like the demo session for this post’s screenshots).
Workspace (the sidebar calls them spaces)“The top-level project container. Use one workspace per repo, task, or investigation.”One space per folder I’m working in. Full stop.
Tab“A layout inside a workspace.”One tab per kind of work: pr-reviews, domain-modeling, implementation.
Pane“A real terminal.”One agent, or one shell, or one thing I’m watching. Arranged in vertical or horizontal splits.

Then there are the agent states, which make the whole thing worth it:

  • Blocked: the agent needs input, approval, or a decision. This is the one you care about.
  • Working: it’s chewing.
  • Done: it finished and you haven’t looked yet.
  • Idle: it finished and you have looked.
  • Unknown: Herdr can’t tell, usually because it’s a plain shell.

Those states roll up into the sidebar. Each space gets a dot, and there’s an agents list underneath that shows every agent across every space. One glance and I know who’s blocked, who’s working, and who’s done and quietly waiting for me to read their output.

That’s the shoulder tap agents never gave me.

How it works (enough to be dangerous)

Keyboard

The default prefix is ctrl+b, deliberately tmux-flavored, and the keymap is “prefix-first” so Herdr doesn’t steal keystrokes from your shell, editor, or the agent. The ones I actually use:

KeysDoes
prefix vSplit vertical: new pane to the right, side by side
prefix -Split horizontal: new pane below, stacked
prefix h / j / k / lMove focus left / down / up / right
prefix zZoom the current pane (and again to unzoom)
prefix rResize mode
prefix cNew tab
prefix shift+tRename tab
prefix 1…9Jump to tab N
prefix shift+nNew workspace (space)
prefix wWorkspace picker
prefix bToggle the sidebar
prefix qDetach (everything keeps running)
prefix ?Help, for when you forget all of the above

Run herdr --default-config to see the full annotated config, including how to remap any of it. Mouse works too: click, drag, split. I won’t judge.

The CLI and socket API

Here’s where it stops being “tmux with a status light.” Everything you can do with a keystroke you can also do from the command line over Herdr’s socket API. That includes the agents themselves, since they’re running in panes Herdr owns. Every screenshot in this post came from a layout I scripted like this, in an isolated session so my real one never got touched:

# a space per folder
herdr workspace create --cwd ~/Codez/ledger-service --label ledger-service

# name the first tab for the kind of work happening in it
herdr tab rename w3:t1 domain-modeling

# agent on the left, reference material on the right, split again downward
herdr pane split w3:p1 --direction right --ratio 0.5
herdr pane split w3:p2 --direction down  --ratio 0.5

# start an agent and hand it a scoped prompt
herdr pane run w3:p1 "claude --permission-mode plan"
herdr agent prompt w3:p1 "Read docs/domain.md and internal/ledger/entry.go. Using only the domain terms Account, Entry, Journal, and Posting (never 'transaction'), describe a Posting aggregate..."

# and then, the good part
herdr agent list      # who is working / blocked / done
herdr agent wait w3:p1 --until blocked --until done   # block until it needs me
herdr pane read w3:p1 # read what it said, without switching to it

That’s enough rope for agents to spawn panes, prompt each other, and notice when a sibling is blocked. I’ll admit that’s both very cool and the setup for a future post titled something like “How My Agents Unionized.”

Integrations

herdr integration install claude (or codex, copilot, opencode, and so on) drops a small hook into each agent so it reports its lifecycle to Herdr directly, rather than Herdr inferring state by watching the screen. herdr integration status shows what’s wired up. Do this. The state dots get noticeably more accurate.

The rest of the toolbox

  • Worktrees: herdr worktree create makes a git worktree and opens it as its own space. It fits the one-branch-per-agent pattern I wrote about in Giving Every Agent Its Own Branch: Git Worktrees for Parallel AI Work.
  • Remote machines: herdr --remote <ssh-target> puts panes on other boxes in the same sidebar, and they reconnect independently. See Connecting machines.
  • Session state: layouts persist and restore. Processes don’t survive a server restart (nothing does, that’s physics), but the shape of your workspace does. See Session state.

How I actually use it

That’s the brochure. Here’s what my screen looks like.

One space per folder

Every folder I’m actively working in gets its own space, named after the folder. As I type this, my sidebar has ten of them: interlinedlist, then interlinedlist-ios, interlinedlist-macos-native, interlinedlist-windows-app, interlinedlist-android, interlinedlist-architecture-migration, a couple of client and side projects, and the repo this post lives in.

Is that a lot? Sure. Is it more than what’s actually in my head? No. It’s exactly what’s in my head, and that’s the point. The sidebar is a direct readout of my working memory. If it isn’t a folder I’m working in, it doesn’t get a space. When I stop working in a folder, the space goes away. My config.toml even sorts the agent list by space (agent_panel_sort = "spaces") so the agents group the same way my brain does.

It’s mostly for my own mental organization, and I’m fine with that. The agents don’t care what the sidebar looks like. I do.

One tab per kind of work

Inside a space, tabs aren’t “more room.” Tabs are modes. I name them for the kind of work being done:

  • domain-modeling: one agent, plus the glossary, domain notes, and recent history open next to it. The agent’s job is to think, not type.
  • implementation: one or more agents, each scoped to a slice of the code, with a diffstat or test runner and a plain shell beside them.
  • pr-reviews: a reviewing agent next to the actual diff, so I can check its claims against the code without switching windows.

Most spaces have just one tab, whichever mode that folder is in today. The busy ones have all three. Switching tabs is switching hats, and the tab label tells me which hat is on before I’ve read a single line.

One or more terminals, split vertically or horizontally

Within a tab I keep a simple rule: agents get tall columns, shells get short stacks.

Agents produce long, scrolling output: plans, numbered findings, diffs they want to show me. That wants height. So agents get vertical splits (prefix v), side by side, full height. Shells, test runners, git log, diffstats and log tails are glanceable. They get horizontal splits (prefix -) stacked in the right-hand column.

The ultrawide changes the math

Here’s the cheat. I run all of this on a 5120×1440, 32:9 ultrawide: two 2560×1440 monitors fused into one with no bezel down the middle. At normal terminal font sizes, that’s room for three or four full-height agent columns plus a stacked utility column and the sidebar, all visible at once, nobody squished.

On a regular 16:9 screen, that same implementation tab looks like this:

Workable. Tight. Add a third agent and you’re zooming panes in and out all day. On the ultrawide it’s this:

The difference isn’t really “more stuff.” It’s that nothing is hidden. When an agent goes blocked, I don’t need to find it. It’s already on screen, in my peripheral vision, with a red dot in the sidebar to back it up. Herdr tells me who needs attention. The ultrawide means I can see why without changing anything.

That screenshot is an agent I’d told to touch cmd/main.go only. It found the repo had no go.mod, which meant staying in scope required a decision. So it stopped, went blocked, and asked. Herdr flagged it, and I could see the question without switching anything. Total cost: one keystroke. The alternative, an agent “helpfully” adding a module file and rewiring imports across a package I’d explicitly fenced off, costs a lot more than one keystroke. Why that worked, and how I write prompts so it keeps working, is the next post (TBD).

The short version

  • Herdr is a terminal workspace manager for coding agents: one Rust binary, runs in your terminal, background server, agent-state detection.
  • Spaces are folders. One per folder you’re working in. It’s a mirror of your head, not a filing cabinet.
  • Tabs are modes of work. pr-reviews, domain-modeling, implementation.
  • Panes: agents tall, shells short.
  • The sidebar tells you who’s blocked. The ultrawide lets you see why without moving.

Install it, run herdr, hit ctrl+b ?, and give it a day. Worst case, you’ve got a nice tmux. Best case, you stop finding agents that have been patiently blocked since lunch.

References