Span of control = how many agents you can effectively supervise

How many agents one developer can actually supervise at once. The binding limit is the orchestrator's context, not the token bill, which caps a batch at two to four.

L4 · GOVERNEDWhat this level takes
MUSTNot met, not at this level
  • Span of control is measured: how many parallel agents each developer effectively supervises
  • Performance evaluation includes agent supervision effectiveness (not just personal code output)
  • Developer role is formally defined as "manager of agent fleet"
SHOULDExpected in practice, not required
  • A span-of-control limit is defined per role and derived from what the orchestrator can actually keep in context
  • Agent supervision training is part of standard developer onboarding
EVIDENCEHow you would check
  • Updated role descriptions defining developer as agent supervisor
  • Span of control metrics dashboard
  • Performance review criteria including agent supervision effectiveness
DEPENDS ON
  • Organization L3 (Team Structure & Roles) - context engineer and platform engineer roles must be established
  • Development L4 (Coding Agent Usage) - parallel agents per developer must be operational

What It Is

Span of control is a management concept from organizational science: how many direct reports can one manager effectively supervise? The answer for human teams is typically 5-9, with 7 as a common optimum. The same concept applies to agent fleet management: how many AI agents can one developer effectively supervise simultaneously? The answer turns out to be in the same range - typically 3-7 - and for the same underlying reasons.

The parallel between human team management and agent fleet management is not superficial. In both cases, the limiting factor is the supervisor's attention and cognitive bandwidth. A manager with 12 direct reports cannot give each person the direction, feedback, and support they need; quality degrades and people go off track. A developer with 8 parallel agents cannot review each agent's output carefully, catch errors promptly, and course-correct quickly enough; quality degrades and agents go off track. The cognitive load of maintaining a mental model of "where is each unit, what do they need, what decision needs to be made next" scales with the number of units, and eventually exceeds the supervisor's capacity.

The 3-7 range for agents specifically reflects two constraints. The lower bound (3) represents the minimum parallelism that actually delivers a throughput multiplier worth the coordination overhead. Running 2 agents has limited advantage over sequential work once you account for the attention cost. The upper bound (7) reflects the maximum number of simultaneous agent states that most developers can hold in working memory while maintaining adequate supervision quality. Some developers with highly developed task management systems and high-context, well-defined tasks can operate at 7-10; most developers working on moderately complex tasks with moderate context quality are most effective at 4-6.

The practical span of control for a given developer on a given day depends on several variables: task complexity (simple, well-defined tasks support higher spans), context quality (agents with excellent context make fewer disruptive decisions), review cycle discipline (structured 15-minute cycles support higher spans than ad-hoc monitoring), and the developer's experience with fleet management (experienced fleet managers can handle more agents than those new to the practice).

There is, however, a harder constraint underneath all of that, and 2026 located it precisely. Rahul Garg's "The Orchestrator's Tax" (martinfowler.com, July 28) argues that the value of a multi-agent setup comes from protecting the orchestrator's working memory, not from parallelism, and that what degrades decisions is context pollution rather than token cost. His line is the one to remember: "Tokens are spent once. Context shapes every decision that follows." This reframes the whole question. The reason the fifth concurrent agent hurts is not that it costs more or that the human gets tired; it is that every status update, every partial result, every interruption it generates lands in the orchestrator's context and degrades every subsequent decision made there. A team optimizing its token bill and a team optimizing its orchestrator's context will make opposite choices, and only one of them is optimizing the thing that binds.

Garg's four standing rules follow directly, and they are the operative numbers for this practice: cap batches at two to four agents, restrict status polling, allow no concurrent repo-wide git operations, and treat overlapping file ownership between agents as a signal to consolidate rather than a coordination problem to manage. The batch cap is much tighter than the personal-span figures above, and the two are not in conflict once you separate them. The batch is what shares one orchestrator's context; a developer may have more work in flight than that across separate, isolated contexts. What they cannot do is fold six agents' chatter into one working memory and expect the decisions made there to hold up. Addy Osmani's "Practical Loop Engineering" (August 14) reports running five to ten parallel agents and maxing out at around five, alongside a rule that belongs in every fleet setup: never let a drafting agent approve its own work, use a separate verifier. And Rachel Laycock's "The Conductor Developer" (July 31) names what is actually being rationed here - human attention, not coding capacity - along with the awkward fact that attention management, energy management and deciding under incomplete information are precisely the skills no career ladder currently measures.

Why It Matters

Understanding span of control provides a principled framework for structuring AI-augmented work:

  • Sets realistic throughput expectations - teams that expect developers to run unlimited parallel agents will be disappointed; teams that plan around 4-6 agents per developer get accurate velocity projections
  • Identifies where productivity gains plateau - the throughput benefit of adding agents is not linear; the first 3 provide high additional value, the next 2-3 provide moderate additional value, beyond 7 the marginal value typically falls below the marginal supervision cost
  • Guides task design decisions - if you want to increase a developer's effective span, the most direct intervention is increasing context quality and task specificity so that agents need less supervision per task; this is a better investment than trying to increase raw supervision capacity
  • Informs hiring and team structure - an organization of 10 developers with spans of 5 agents per developer can sustain 50 simultaneous agent workstreams; this changes how you think about the relationship between headcount and development capacity
  • Creates a growth metric for developers - increasing your effective span of control is a concrete, measurable developer skill; a developer who moves from a span of 3 to a span of 6 has demonstrably become a more effective fleet manager; this is a new dimension of developer growth that the traditional role structure doesn't capture
  • Points the optimization effort at the right resource - teams that treat parallelism as a cost problem tune their token spend and keep hitting the same quality wall; the resource actually being consumed is the orchestrator's context, and once you know that, the interventions change: smaller batches, less polling, cleaner ownership boundaries, none of which show up on an invoice
TIP

Track your actual span of control over a two-week period. Log how many agents you started, how many you were supervising simultaneously at peak, and how often you had to clean up agent work that went significantly wrong. This is your real span of control, not your theoretical maximum. Start from this baseline rather than aspirational targets.

Getting Started

  1. Establish your current effective span - run at your current natural comfort level for two weeks and observe: at what number of parallel agents does your review quality decline? At what number do you start missing agent problems at check-in? This is your current span ceiling. Most developers discover their effective span is 3-4 before they develop fleet management discipline.
  2. Identify your span-limiting factor - is your limit cognitive (too many agent states to hold in working memory)? Logistical (your terminal setup makes it hard to see multiple agents at once)? Or review bottleneck (you can manage the agents but reviewing their outputs takes too long)? Different limiting factors have different solutions.
  3. Address logistics before cognitive load - the easiest span expansion is tooling. A well-designed terminal layout showing all agents simultaneously (tmux split panes, multiple iTerm2 windows, a status dashboard) reduces the cognitive load of maintaining mental state about where each agent is. Set up your environment before trying to push your span higher.
  4. Increase context quality to expand span - the biggest multiplier for span is agent context quality. Agents that are better contextualized make better decisions, need less supervision, and produce cleaner outputs. A 20% improvement in context quality can increase your effective span by 1-2 agents because you're spending less supervision time per agent. Invest in context before pushing span higher.
  5. Expand span incrementally - don't jump from 3 to 6. Add one agent at a time and observe the effect on quality and review burden. If quality holds when you add a fourth agent, keep the four running for a week and observe again. If review burden becomes unsustainable, go back to three and identify the bottleneck before expanding.
  6. Use task homogeneity to support higher spans - running 6 agents on similar, well-understood task types is easier than running 4 agents on highly varied complex tasks. When you want to operate at higher spans, batch similar work. When you have heterogeneous, complex tasks, operate at lower spans with more supervision time per agent.
  7. Cap the batch at two to four and hold the line - whatever your personal span, the number of agents sharing one orchestrator's context should stay between two and four. This is the rule to adopt first, because it is the one that protects the resource that actually binds. If you need more work in flight, run additional batches in separate contexts rather than widening one.
  8. Restrict status polling and forbid concurrent repo-wide git operations - polling feels like diligence and behaves like pollution: every status check writes into the orchestrator's context and degrades the decisions made afterwards. Set a check-in rhythm and stick to it instead of watching. Concurrent repo-wide git operations are the other reliable way to corrupt shared state; serialize them.
  9. Read overlapping file ownership as a signal to consolidate - when two agents in a batch keep touching the same files, that is not a coordination problem to solve with more supervision. It is evidence that the work should have been one task. Merge them rather than arbitrating between them.
  10. Put a separate verifier behind the fleet - never let a drafting agent approve its own work. A dedicated verification step, agent or human, is what keeps a wider batch from converting straight into rework, and it is the difference between running more agents and running more agents usefully.

Common Pitfalls

Treating span as a competition. Some developers try to maximize span as a status signal ("I run 10 agents at once"). This is the wrong mental model. The goal is effective span - the maximum number of agents you can supervise while maintaining quality and not creating a cleanup workload that exceeds the throughput benefit. An effective span of 5 is better than a nominal span of 10 with poor quality.

Ignoring the review bottleneck. Many developers find that their span is not limited by supervision capacity but by review capacity. They can monitor 8 agents but can't review 8 completed PRs per day at the quality standard the team requires. Span of control must account for the full developer bandwidth cost: supervision plus review. If review is the bottleneck, increasing span creates a PR backlog that eliminates the throughput benefit.

Not accounting for variation in task complexity. A developer who calibrates span of 6 for simple, well-defined tasks will be overwhelmed if they try to run 6 agents on complex, ambiguous tasks. Task complexity is a first-order variable in effective span. Calibrate span to the tasks at hand, not to a general personal limit.

Failing to decrease span when context quality degrades. If context infrastructure is under-maintained - CLAUDE.md files become stale, MCP servers have outages, documentation falls behind code changes - the effective span will decrease because agents make more errors per task. Organizations that don't maintain context quality will see effective spans decrease as codebases evolve.

Optimizing the token bill instead of the orchestrator's context. This is the most consequential mistake in the practice, because it is the one that feels responsible. A team watching its spend will reach for cheaper models, tighter prompts and shorter runs, and will keep hitting the same wall, because the thing degrading its decisions was never cost. Tokens are spent once; context shapes every decision that follows. The interventions that actually work - smaller batches, less polling, no concurrent repo-wide git operations, consolidating overlapping ownership - are invisible on an invoice, which is exactly why they get skipped.

Missing the organizational span of control question. Individual span of control is one question; the organizational equivalent is another. How many total agents can the organization's infrastructure support simultaneously? How many agents can the review process absorb? These organizational span limits are distinct from individual developer limits and need to be planned for separately.

How Different Roles See It

BobHEAD OF ENGINEERING

Bob is doing sprint planning and trying to estimate velocity for the upcoming quarter. His team has 8 developers operating at various levels of AI fluency. He's trying to figure out how many parallel agent workstreams the team can sustain and what the actual throughput multiplier is compared to pure human development.

What Bob should do: Bob should survey each developer's current effective span and use this to estimate total concurrent agent capacity. If six developers are effectively running 4-5 agents each and two are at 2-3 agents, the team has roughly 28-36 concurrent agent workstreams available. This is a more accurate planning number than "we have 8 developers with AI tools." Bob should also distinguish between developer time (thinking, specifying, reviewing) and agent time (executing tasks). A developer managing 5 agents is still working a full day - they're just doing different work. The throughput multiplier is in the agent execution, not in the developer headcount. Bob should communicate this distinction to stakeholders who ask "how many developers does this replace."

SarahPRODUCTIVITY LEAD

Sarah is designing the organizational metrics dashboard for AI adoption. She wants to include a "productivity" metric but is struggling to define what to measure. She's been asked to demonstrate that the AI tooling investment is delivering value, and stakeholder expectations are high.

What Sarah should do: Sarah should build the metrics around effective span of control because it captures the right thing: how effectively are developers using the AI capacity available to them? She should track, per developer per sprint: number of parallel agents run, number of tasks completed by agents versus directly, and review-to-agent ratio (how many PRs did the developer review versus write). The trend over time should show increasing effective span as developers develop fleet management skills, which translates to increasing throughput per developer. This metric is honest - it measures actual agent utilization, not just tool installation - and it provides actionable data for improving AI practices.

VictorSTAFF ENGINEER - AI CHAMPION

Victor runs at a span of 6-7 agents effectively and has developed the most sophisticated fleet management workflow on the team. He's been asked to help two other developers increase their effective span from 3 to 5. He's not sure how to coach this transition because most of what he does has become intuitive.

What Victor should do: Victor should work backwards from his current practice to make the implicit explicit. He should spend a day narrating his decisions aloud as he manages his agent fleet: "I'm launching three agents now instead of five because these tasks are all in the same service and I want to sequence them to avoid conflicts. I'm checking in on agent two first because it's working in an area with unclear requirements. I'm not interrupting agent four because it's on track with a well-defined task." This narration reveals the judgment calls that experienced fleet management requires. Victor should record this or take notes, then work with the two developers he's coaching to help them make the same judgments consciously before they become automatic. He should also be careful about what he coaches toward. Most of what makes his practice work is the four rules rather than his personal capacity - batches held at two to four, no polling, no concurrent repo-wide git operations, overlapping file ownership merged rather than arbitrated - and those transfer to a developer on their first week. Coaching someone from three agents to five without those rules just moves the failure one agent later.

How This Guide Changed

What each edition changed in this guide, newest first.

  1. V1.6September 2026LATEST

    Relocated the constraint. The guide had framed span as a question of human working memory, which is true but not the thing that binds first: Rahul Garg's "The Orchestrator's Tax" identifies context pollution rather than token cost as what degrades multi-agent decisions, and supplies the operative numbers - cap batches at two to four, restrict status polling, no concurrent repo-wide git operations, and read overlapping file ownership as a signal to consolidate. Those four rules became Getting Started steps, and a new pitfall names the seductive wrong move of optimizing the token bill instead, since the interventions that actually work are invisible on an invoice. Addy Osmani's report of running five to ten agents and maxing around five, plus his rule that a drafting agent never approves its own work, and Rachel Laycock's framing of attention rather than coding capacity as the rationed resource, round out the reframe.

  2. V1.5August 2026

    The opening dropped the management-theory preamble and led with the practical question instead, how many agents one person can actually supervise, so the number arrives before the analogy that produced it.

  3. V1.3June 2026

    The limit and the reasoning behind it were unchanged; three stale references, on human span of control, on cognitive load, and on parallel agent workflows, were re-pointed at their current homes.

  4. V1.0March 2026

    Borrowed a concept from organizational science and asked whether it transfers. If a manager tops out somewhere around seven direct reports, how many agents can one developer supervise at once? The launch answer put agent fleets in the same range and for recognisably the same reasons, which gave fleet management a number to plan against instead of an ambition.

Where does your team actually sit on this?

This guide describes one level of one area. Run the assessment to place your team across all 16 areas, see which gates you have passed, and get a report you can take to your stakeholders.

Start the assessment