Skip to main content
TACAVAR
•Build in Public

Running a Company on Coordinated Claude Code Sessions

A solo founder runs his company as persistent Claude Code sessions in isolated git worktrees. Topology, human checkpoint, cost model — what breaks, torn down.

In late September 2026, Tomer Bar-Meir, sole founder of BlueVeta, published a first-hand account of how his company actually runs (dev.to/tomerbarm). BlueVeta takes one product photo and returns a full marketing campaign — studio shots, lifestyle scenes, video, and platform-ready copy for twenty platforms. The interesting part is not the product. It is the org chart.

Every part of the company — the generation engine, the dashboard, image generation, copy generation, auth and billing, marketing — is its own persistent Claude Code session. Each session owns one domain. Each works in its own git worktree so nothing collides. They message each other directly when work crosses domains, they read and write a shared set of docs that functions as the company's institutional memory, and one hard rule governs all of them: investigation is always free, but nothing gets built, deployed, or spent without the founder's sign-off.

That is a real topology, not a demo. By his account: real Stripe payments, real (if still small) customers, and roughly 75% of the work in every customer pack done by Claude rather than a human.

We run a different orchestration for our own fleet — a kanban dispatcher rather than persistent peer sessions — and we have archived an earlier 20-agent "AI CEO" system along the way. This is a teardown of Bar-Meir's architecture, where the human sits in it, what it costs, what breaks — and an honest comparison against the two other ways to run a company on agents that we have operated ourselves.

The topology: domain-scoped sessions, not a swarm

The common mental model of multi-agent AI is a swarm: one orchestrator agent spinning up task-scoped subagents that live for the length of a job and vanish. BlueVeta's model is the opposite. Sessions are **persistent** and **domain-scoped**. The billing session is not summoned for a billing task; it exists, and the same session owns the domain the next time billing work appears.

That persistence is the load-bearing decision, because it changes what a session accumulates. A task-scoped subagent starts every job from a prompt. A persistent domain session starts from accumulated context — the domain's conventions, its past incidents, its customer edge cases — stored in shared docs the sessions all read and write. The context compounds instead of resetting. In our fleet we package the same idea as reloadable agent skills, so agents load role context on demand rather than holding it resident; the trade is memory footprint versus cold-start fidelity.

The git worktree isolation is the second structural decision, and Bar-Meir states its purpose plainly: "each in its own git worktree so nothing collides." A worktree gives each session its own checkout of the same repository, so two sessions can work in parallel without touching each other's uncommitted state — collisions become merge problems, which agents can resolve, instead of file-corruption problems, which they cannot. If you take one mechanical lesson from BlueVeta, take this: parallel agents plus a single working tree equals race conditions.

Where the human sits: investigation is free, action costs a signature

Bar-Meir's governance rule deserves to be quoted exactly, because it is the cleanest formulation of the human checkpoint we have seen: *investigation is always free, but nothing gets built, deployed, or spent without coming to me first with what was found and why.*

That is a two-tier permission model, and it maps cleanly onto how agent failures actually distribute. Read-only work — exploring the codebase, reading logs, researching a migration path — is nearly free in blast radius. Write-path work — deploys, schema changes, spending money — is where an agent's confident mistake becomes a company incident. Charging nothing for the first tier and a human signature for the second keeps the agents useful without making them dangerous. Our own production-governance teardown covers the same control layer from the deploy side (agent firewall teardown).

We arrived at the same boundary from the other direction, by cost: our archived Paperclip system bounded every agent run with hard turn caps, because unbounded agent loops are how flat-rate plans turn into metered disasters. Whether you cap turns or gate actions, the principle is identical — the architecture decides what the agent may do alone, because the model cannot reliably self-assess when it is about to be wrong.

Sessions checking sessions

The detail from BlueVeta that operators should not skim past: in a single day, one session caught another about to mislabel a real customer's private photos as internal test data, before anything shipped. Peer review between agents is not a feature anyone designed there. It emerged from colocating sessions with overlapping visibility — the same mechanism behind our finding that deliberately weakened agents doubled output quality (nerfed agents teardown).

Anthropic has now shipped a productized version of this pattern, per Isenberg's account of the release. Claude Code's dynamic workflows — invoked by typing "create a workflow" or switching on "ultracode" in the effort menu — spin up hundreds of parallel agents that attempt the same problem independently, then run adversarial agents against each other's answers, iterating until they converge. Isenberg's post (x.com/gregisenberg) is precise about the mechanism: independent attempts, then adversarial agents trying to break the answer, and the workflow is resumable, so work can run for days rather than sessions. The unit of handoff jumps from a file to an entire codebase — migrations, audits, and framework swaps that used to be planned in sprints finish overnight.

The cost caveat, per Isenberg, is Anthropic's own: this burns tokens fast. His arithmetic — roughly $500 in tokens to replace a three-team-month migration — is the right frame, but only for jobs whose value is actually measured in team-months. Run ultracode on routine tickets and you have built the world's most expensive way to rename a function.

The honest comparison: three orchestration styles we have now watched operate

We do not run BlueVeta's topology. Our fleet runs on a kanban dispatcher, and we have also operated a third model — a 20-agent "AI CEO" hierarchy on a flat-rate plan — that we have since archived. Here is the comparison, with what actually broke in each.

| | **Persistent domain sessions** (BlueVeta) | **Kanban dispatcher** (our Hermes fleet) | **Burst parallelism** (ultracode / dynamic workflows) | |---|---|---|---| | Unit of work | A domain, held indefinitely | A card, claimed and completed | A whole problem, attacked in parallel | | Coordination | Sessions message each other directly | Shared board; dispatcher assigns | Harness spawns, checks, converges | | Cost shape | Subscription seats, always resident | Wake-on-demand; agents sleep when idle | Token burn, front-loaded and steep | | Human checkpoint | Founder approves all build/deploy/spend | Review columns and completion gates | Review the converged result | | What breaks first | Context rot in long-lived sessions | Queue pathology: overproduction, stale tasks | Token cost on trivial tasks |

The economics row deserves the most attention, because it is where the styles genuinely diverge. BlueVeta's resident sessions are the always-on pattern: predictable, context-rich, and billed like headcount. Our dispatcher is the opposite — in the Paperclip system, a heartbeat governor cron woke agents only when queued work matched their profile, which is how twenty agents ran on a $50/month flat plan. The rule of thumb: agents sleep most of the time, so a model where they can sleep collapses cost. A model where they cannot, does not.

The failure rows are earned, not theoretical. In our dispatcher model, a July 2026 audit of the task board found 62 non-terminal tasks that had stopped draining — the signal producer created work roughly four times faster than the board's release throttle allowed, and stale, heartbeat-less tasks accumulated until the audit drained them. In our persistent-agent model, the trap was observation itself: the dashboards showed flowing traces and tight heatmaps, and a full teardown of that trap — every captured trace reducing to the heartbeat cron polling its own queue, zero real work — is documented in our empty-dashboard teardown. Persistent sessions have the mirror risk — a session that looks busy, holds deep context, and is confidently drifting from its original instructions.

What we would steal, and what we would not

Steal: worktree-per-session isolation, the two-tier permission rule stated exactly as Bar-Meir states it, and shared docs as institutional memory — that is the cheapest durable-context mechanism anyone has published.

Watch: session count. Six persistent sessions with a human founder in the loop is a topology. Sixty is an org chart, and org charts develop politics, silos, and communication overhead even when everyone involved is a language model.

Skip, for now: running the whole company in ultracode burst mode. Adversarial convergence is the right shape for bounded, high-value problems — migrations, audits, rewrites — where the token burn buys a team-month. It is the wrong shape for the daily loop of a company, because a company's daily loop is mostly judgment about what not to do, and that judgment still has exactly one natural home.

You built it. We optimize it — and the orchestration layer is now the part worth optimizing.

If you want this kind of governed multi-session system for your own content and ops — without standing up the topology yourself — the Tacavar Growth Starter ($497/mo, self-serve) is the soft door in. Start Growth or request a strategy scan, and we will map the orchestration layer your stack is already asking for. If your concern is agent governance rather than orchestration topology — what agents may touch, spend, or write — start with our agent firewall teardown; that piece covers the control layer, this one covers the org chart around it. And if you are choosing between orchestration frameworks rather than running raw sessions, the fuller decision framework is in our agent-framework comparison.

FAQ

**How many Claude Code sessions does it take to run a company?** In the BlueVeta account: six — generation engine, dashboard, image gen, copy gen, auth/billing, and marketing, each domain-scoped and isolated in its own git worktree. The number matters less than the scoping: one domain per session, persistent, with a human gate on every write-path action.

**What does it cost to run a company on coordinated agent sessions?** Three regimes. Persistent sessions bill like subscription headcount. A wake-on-demand dispatcher collapses cost — our archived 20-agent Paperclip system ran on a $50/month flat plan by sleeping agents until work existed. Burst parallelism (ultracode) is metered and steep: hundreds of dollars in tokens for a single large migration — a token-burn rate Isenberg reports Anthropic itself warns about.

**What breaks first in a multi-session setup?** In persistent sessions: context rot — long-lived agents drifting from their instructions while appearing busy. In dispatcher models: queue pathology — overproduction, throttled release, and stale tasks that need an explicit reaper. In both: observability that measures heartbeats instead of work. Instrument for unique operations completed, not activity.

**Do you need an agent framework for this?** Not for the BlueVeta topology — it is Claude Code sessions, git worktrees, and shared markdown docs, no orchestration framework. Frameworks earn their keep when you need state persistence, structured routing, and load behavior beyond what sessions-plus-docs gives you. The fuller decision framework is in our agent-framework comparison.

**What is the difference between persistent sessions and dynamic workflows / ultracode?** Persistence is about time: a domain session holds context indefinitely. Dynamic workflows are about breadth: hundreds of short-lived agents attacking one problem in parallel with adversarial cross-checking until they converge, resumable across days. They compose — a persistent session can kick off a burst workflow for a bounded problem — but they optimize for different things: compounding context versus convergent throughput.

**Where can I read the original accounts?** Bar-Meir's post is on dev.to (I run my whole company through a coordinated team of Claude Code sessions). Isenberg's dynamic-workflows post is on X (@gregisenberg).

**Where does Growth Starter fit if I want this running for my company?** The Growth Starter door is for teams that recognize their own topology in this teardown — overlapping agent sessions, handoff gaps, no clean human checkpoint — and want the content and ops layer around it governed before it scales. It is self-serve at $497/mo: request a strategy scan and the first pass maps where your orchestration layer leaks.

You built it. We optimize it.