SAKIZLI AI
Article15 Sept 2026 · 45 min read24 / 35Free to read · Assets for members

Holacratic AI Teams

Rethinking goals, roles, and autonomy.

AgentsAutonomyGovernanceResponsibility
FFurkan SakızlıAI researcher & tutor · independent
A central control hub connects via glass conduits to six surrounding glass domes, each housing two objects beneath a transparent cover
Each role owns its own domain under the dome – connected to the center, but not dependent on it
Image generated with AI

AI agents are often treated like extremely fast interns: the human decides the route, decomposes the work, formulates every step, and then expects obedient execution. The pattern is understandable. It matches organizational structures built over decades and the way we usually control software. But it leaves part of the value of agentic systems unused: the ability to discover new routes, alternatives, risks, and shortcuts inside a goal framework.

The counter-movement initially sounds radical: think of AI agents not only as tools, but as autonomous roles in a distributed team. Not every step is prescribed. Instead, purpose, responsibility, decision space, and boundaries are defined. Agents may propose options, prioritize, raise objections, and align their work with a goal.

This idea can be described with a concept from organizational design: Holacracy. The term is easy to misunderstand, however. Holacracy does not mean that structure, leadership, and responsibility disappear. Quite the opposite: the formal Holacracy model describes explicit roles, purposes, domains, accountabilities, rules, and governance processes. Its core idea is not structurelessness, but distributed authority[2].

Applied to AI teams, this creates a productive question:

How much autonomy can an AI role receive without responsibility, quality control, and shared purpose becoming blurred?

The answer is neither micromanagement nor unlimited freedom. It is a precise architecture of goals, roles, domains, decision rights, and controllable initiative.

Holacracy does not mean "nobody leads anymore"

The popular shorthand for Holacracy is often: flat hierarchy, everyone equal, nobody is the boss. That is useful as a provocation, but too rough as a description.

The official Holacracy Constitution[1] defines an organization through roles and circles. Roles have a Purpose, Domains that they may control, and Accountabilities that describe ongoing responsibilities. Circles group roles around a common purpose. There are also Policies, governance processes, and functions that can set priorities or strategies.

So Holacracy does not remove structure. It attempts to move authority away from people and titles and toward explicit roles and rules.

For AI teams, that distinction is crucial. An AI should not be "free" because nobody is responsible anymore. It should be able to act autonomously inside a well-defined role because responsibilities and boundaries are clear enough that every trivial action does not require prior approval.

Autonomy then emerges not from the absence of leadership, but from clarity about where leadership is no longer needed.

The real shift: from path to purpose

A conventional prompt often describes the path: analyze A first, then compare B, use method C, create D afterward, and finally give me E.

That can be appropriate when the process itself is fixed or legally prescribed. But it has a cost: the user defines the solution space in advance.

A goal-oriented assignment works differently. It focuses more on: which state should be reached, which quality criteria apply, which boundaries must not be crossed, which resources are available, which decisions the role may make independently, and when uncertainty must be made visible.

The AI is no longer asked to reproduce a known route as obediently as possible. It is asked to find an admissible and evidence-backed route to the goal.

That is a deeper change than "prompt engineering." It is organizational design.

Why too much direction can destroy potential

Language models are strongly optimized for instruction following. That is exactly what makes them useful in everyday work. At the same time, this creates a familiar problem: when the human locks in a direction very early, the model often continues inside that framing even when other solution paths might be better.

This is not a mysterious "obedience instinct." It is a consequence of the interaction. The prompt defines a task space; the narrower that space is, the less exploration happens outside it.

In projects, this produces a paradox: an organization buys a system because of its broad pattern recognition and then forces it to work like an assistant wearing blinders.

A holacratically inspired way of working therefore tries to allow not only answers, but initiative: "Which assumption in our plan looks weak to you?", "Which better route would you propose?", "Which option are we not testing?", and "Which decision would you make differently inside your role?"

These questions change the relationship. The AI is not automatically right. But it receives an institutional space in which deviations can become visible in the first place.

AI as a team member is a metaphor, not a legal status

Precision matters here. An AI is not an employee, a legal person, or a morally accountable member of an organization. When we speak about "equal team members," that is an operational metaphor.

It can usefully mean: the AI may raise professional objections, its proposals are evaluated by evidence rather than rank, it has a defined area of responsibility, within that area it may make local decisions, and it does not need to wait for human micro-approval for every trivial step.

It must not mean: that responsibility is equally distributed, that human accountability disappears, that legal responsibility is transferred to a model, or that an agent may redefine rules or goals on its own.

Operational voice can be distributed. Accountability must remain attributable.

That separation is central, because otherwise a useful organizational principle turns into dangerous anthropomorphism.

The role contract: Purpose, Domain, Accountabilities, Constraints

For holacratically inspired AI teams, naming roles is not enough. "Research agent," "strategy agent," or "reviewer" are labels. Autonomy requires a robust role contract.

A good AI role should include at least four layers.

Purpose – Why does the role exist?

The Purpose does not describe individual tasks. It describes the intended contribution to the larger system.

Example:

"Increase the robustness of strategic decisions by surfacing counter-options, assumptions, and external evidence."

That is stronger than: "Research competitors." The Purpose gives the role a reason to act even when it encounters a new gap that was not explicitly named in the original task.

Domain – What may the role control?

A Domain defines the area in which a role may act independently. It can be a data store, a document, a subprocess, a toolset, or a decision space.

Examples: read access to the entire research pool, but write access only to the analysis area; independent prioritization of open research questions; free tool choice inside an approved tool group; no authority to change the project baseline.

Domains prevent "autonomy" from becoming general access power.

Accountabilities – What does the role owe the team?

Accountabilities are ongoing responsibilities. They turn an abstract Purpose into observable work.

A risk role could, for example, be accountable for: detecting new risks, reassessing existing risks, marking assumptions supported by weak evidence, formulating alternative failure paths, and translating relevant findings into a defined output format.

Constraints – Which boundaries always apply?

Constraints do not define how the work must be done. They define what must not be exceeded.

Examples include: data classes, budget ceilings, allowed tools, prohibited irreversible actions, evidence or testing requirements, and defined escalation cases.

The formula is:

Purpose gives direction. Domain gives space. Accountabilities give responsibility. Constraints give boundaries.

The autonomy envelope: freedom with clear geometry

Autonomy should not be binary. An agent is not simply either "controlled" or "free." A better model is an autonomy envelope: an explicit space of permitted decisions.

Decision typeTypical autonomy level
Search, sort, and cluster informationhigh within approved sources
Determine the sequence of its own subtaskshigh
Make reversible local changesmedium to high
Ask other roles for contributionsdepends on orchestration rules
Change a shared baselinelow / explicitly governed
External publication or irreversible actionstrongly limited
Change goals, governance, or safety rulesoutside normal role autonomy

This envelope matters more than a long prompt. It tells the role where initiative is desired and where it ends.

A professional system therefore does not try to describe every future action. It describes the decision landscape.

"You are allowed to disagree" must be a real function

Many teams say criticism is welcome. In practice, the AI is still prompted to optimize the existing plan. That is not a right to dissent. It is optimization inside a given frame.

An autonomous role needs a defined dissent channel – a mechanism for surfacing relevant deviations.

A robust objection should state: which assumption or decision is affected, which evidence speaks against it, which consequence is likely if nothing changes, which alternative is proposed, and how certain or uncertain the assessment is.

That separates dissent from opinion.

The important cultural rule is not: "The AI should be critical." It is:

"When your role detects a relevant tension between the current state and your Purpose, you must make it visible."

Tensions as a sensing system

In Holacracy, the concept of a "Tension" plays an important role: a perceived gap between present reality and a better possible state.

For AI teams, the idea is surprisingly useful. Agents can act as sensors for tensions: a reviewer detects quality gaps, a market role detects new external signals, a risk role discovers an assumption that is no longer valid, and a user-perspective role detects contradictions between product logic and customer need.

The important point is that not every detected tension automatically triggers a change. A tension is first a signal, not a mandate.

This creates a useful middle ground between two bad extremes: the agent reports nothing because it merely executes orders, or the agent changes the entire project because it found a better idea.

Holacratic AI teams need sensing and proposal rights – not unlimited constitutional change.

Governance and operations are two different levels

One of the most useful lessons from formal Holacracy is the separation between governance and operations[3].

Operations answers: what do we do now? Which next action creates the most value? Which task do we prioritize? Which method do we use inside our role space?

Governance answers: which roles exist? Which Domains do they own? Which Policies apply? Who may decide what? Which structural rules do we change?

For AI systems, this distinction is essential. An agent can have extensive operational freedom while holding no governance authority.

This avoids a common misconception: more autonomy automatically means that an AI may rewrite its own rules, roles, or goals.

It does not. Autonomy can be local and strong while governance remains deliberately stable.

Local authority beats global freedom

"Just let the agent do it" is too vague. A better rule is:

"Let the agent decide independently inside its defined area."

A good distributed structure delegates authority as close as possible to the role that has the relevant context. At the same time, it prevents local decisions from silently creating global consequences.

A content role may change the sequence of its drafts. It should not redefine the target audience of the entire project. A research role may search additional sources. It should not single-handedly overwrite the approved strategic baseline.

Autonomy works best when it is local, reversible, and domain-specific.

Reversibility is a practical delegation filter

A simple heuristic helps answer how much autonomy makes sense:

The easier a decision is to reverse, the more readily it can be delegated to a role.

Reversible decisions: change a search strategy, vary the sequence of internal tasks, test an additional counter-hypothesis, generate a draft variant, adjust internal priorities.

Hard-to-reverse decisions: delete data, transfer money, communicate publicly, change contracts, modify production systems, trigger regulatory approvals.

This heuristic does not replace risk analysis. But it is a good starting point for defining the autonomy envelope.

Strategy should be a heuristic, not a script

Formal Holacracy includes the idea that strategies can guide roles. This is particularly interesting for agentic systems.

A strategy does not need to prescribe every step. It can be expressed as a heuristic: "Prefer robust evidence over fast completeness.", "Prefer reversible experiments over large up-front investment.", "If two options appear equivalent, choose the one with lower lock-in.", and "Prefer the customer problem over the feature request."

Such heuristics give direction without fixing the solution path.

This is the difference between leadership through principles and leadership through microsteps.

Equality does not mean uniformity

A holacratic AI team does not improve when all roles can do everything and know the same things. Value comes from functional asymmetry.

A reviewer should think differently from a generator. A risk role should have different success criteria from a delivery role. A user-perspective role may challenge a technically perfect plan if it misses the actual need.

"Equal footing" can usefully mean: findings are not ignored because of rank, every role has a legitimate area of responsibility, any role may raise criticism when it concerns its Purpose, and decisions follow rules, evidence, and Domain rather than loudness.

It does not mean every role has the same authority.

AI needs institutional function, not a personality

It is tempting to humanize roles: "You are Lea, the critical strategist," "You are Max, the creative visionary." Such personas can influence tone and perspective, but they do not solve the organizational problem.

For a robust AI team, more important questions are: what Purpose does the role pursue? Which data may it see? Which decisions may it make? Which artifacts may it change? Which type of objection must it report? How is its performance measured?

A role is not a character profile. It is an institutional function inside the system.

Shared purpose prevents local brilliance without global value

Autonomous roles can perform brilliantly locally and still make the overall project worse. That happens when they optimize their own targets without considering the higher-level purpose.

Every role therefore needs two levels: its own Purpose, and an explicit relationship to the project's overall Purpose.

A risk role must not simply maximize the number of risks found. Otherwise every project becomes impossible. Its Purpose might instead be: "Improve decision quality by making material risks proportionally visible."

A quality role must not optimize forever. Otherwise nothing ships. Its Purpose must be compatible with delivery, time, and value.

Autonomy without shared purpose creates local optimization. Holacratic structures only work when roles use their freedom in service of a larger system.

Goal conflicts must become visible, not disappear

Distributed authority does not eliminate conflict. It often makes conflict more visible.

Typical conflicts include: quality versus speed, safety versus convenience, innovation versus standardization, cost versus redundancy, and short-term delivery versus long-term maintainability.

A weak agent setup hides those conflicts inside a prompt. A good setup encodes them as explicit tensions.

Then an agent can state:

"My role would prefer option A because it strengthens the quality objective. It adds two days to delivery, however, which conflicts with the delivery target."

This turns disagreement into a manageable organizational signal.

Consensus is not the goal

Holacratically inspired collaboration must not be confused with grassroots voting. More agent votes do not create truth.

Three models can share the same error. Five agents can repeat the same weak source. A majority vote can overrule a well-supported minority finding.

A distributed AI team should therefore optimize not primarily for consensus, but for clear decision logic: who owns the Domain? What evidence exists? Is the decision reversible? Which role has professional responsibility for the issue? Which finding affects the overall Purpose?

The organization does not always need agreement. It needs traceable responsibility.

What current multi-agent research pushes back on

The holacratic metaphor is attractive, but it should not be romanticized.

Anthropic notes in a recent study of multi-agent systems that today's agents can coordinate efficiently when other agents look like well-defined tool calls with clear inputs and outputs. Coordination becomes harder when agents must treat one another as long-lived peers with their own goals and no clear hierarchy. That is exactly where new systemic risks can emerge.[6]

This observation matters: holacratic AI teams are not proof that hierarchy is technically obsolete. They are a design space in which distributed roles can create benefits when purpose, boundaries, and coordination remain explicit.

OpenAI likewise describes both centrally managed and decentralized multi-agent patterns.[4] In the decentralized pattern, specialized agents can hand off work to one another on equal footing. At the same time, the recommendation is to build agent systems incrementally and to add complexity only when it is truly needed.

Current practice therefore argues neither for "everything centralized" nor for "everything free." It argues for context-appropriate distribution of authority.

Holacracy as inspiration, not a copy-paste org chart

It would be a mistake to copy the Holacracy Constitution directly onto AI agents. Humans and models differ fundamentally.

Humans have: social experience, implicit organizational knowledge, accountability, motivation, personal interests, and long-term identity.

Agents, depending on the system, have: limited or constructed context, model-dependent capabilities, tool access, variable reliability, no legal accountability, and no stable organizational identity in the human sense.

Holacracy should therefore be used here as design vocabulary: Purpose, Role, Domain, Accountabilities, Tensions, Policies, distributed authority. Not as a claim that AI agents are digital employees with the same properties as humans.

A simple maturity model for autonomy

Instead of "manual or autonomous," a staged development is more useful.

LevelWorking modeTypical characteristic
1 · ToolHuman gives every taskno independent initiative
2 · Delegated roleRole handles a clear assignmentlocal method choice
3 · Goal-oriented rolePurpose + Domain guide workdetects new subtasks independently
4 · Distributed teamseveral roles act in their Domainsobjections, handoffs, local decisions
5 · Adaptive governancerole structure can evolve through rulesonly through explicit governance process

The last level is not automatically "better." It is simply more complex.

Many real projects will work best at level 2 or 3. Organizations should not seek the maximum possible autonomy, but the minimum necessary control with the maximum useful initiative.

The autonomy control map

For every role, six questions should be answered:

DimensionGuiding question
GoalWhich state should the role improve?
DomainWhat may it control independently?
DecisionWhich decisions may it make without asking?
ResourcesWhich tools, data, time, and budgets may it use?
DissentWhen must it report an objection or alternative?
BoundaryWhich action is explicitly outside its authority?

These six fields form the core of a holacratically inspired AI organization. They create more freedom than micromanagement, but much more clarity than "just do it."

Practical example: a strategy project without super-interns

Suppose a company wants to develop a new service for small and midsize businesses.

A conventional setup might look like this: the human gives a chatbot one task after another – research the market, define the audience, compare competitors, draft the offer, check risks. The AI follows the sequence.

A holacratically inspired setup might instead contain four roles:

Market role: Purpose – make relevant market movement and real demand visible. Domain – market sources, competitor data, trend analyses. Autonomy – search strategy, prioritization, counterexamples.

Customer role: Purpose – ensure that the offer solves an actual problem. Domain – interviews, personas, jobs-to-be-done, feedback. Autonomy – object to feature ideas without a demand signal.

Economics role: Purpose – secure a viable business logic. Domain – cost assumptions, pricing logic, scenarios. Autonomy – develop alternative models and challenge assumptions.

Quality/risk role: Purpose – expose blind spots without blocking innovation. Domain – assumptions, risks, evidence quality. Autonomy – flag critical tensions and request counter-checks.

The human defines overall Purpose, strategic boundaries, and non-delegable decisions. The roles may act inside their Domains and challenge one another professionally.

The decisive point is not that four agents are involved. It is that each role has a legitimate perspective of its own and does not merely reproduce the same boss prompt.

Six anti-patterns in holacratic AI teams

Anti-patternWhy it fails
Pseudo-freedomThe agent is told to “think freely” but may only produce desired results.
Roles without DomainsEveryone may do everything; responsibility blurs.
Democracy illusionMajority voting replaces evidence.
Autonomy without shared PurposeRoles optimize locally against one another.
Personas instead of functionsCharacter descriptions replace decision rights and outputs.
Governance inside the live taskAgent changes rules, goals, or structure during execution.

The common denominator is missing institutional clarity.

Autonomy creates a new management job

The more operational freedom roles receive, the less the human must steer individual steps. At the same time, another kind of work becomes more important: clarifying goals, cutting roles cleanly, defining Domains, exposing conflicts, assigning decision rights, setting performance criteria, and maintaining boundaries.

The human role shifts from task distributor to system designer.

That is the actual organizational shift. Agentic AI does not necessarily reduce management. It changes what management means.

How much freedom is enough?

Three tests help.

1. Insight gain

Does additional autonomy create valuable new options, risks, or perspectives that would not have emerged under micromanagement?

2. Controllability

Is it still understandable why the role acted and whether the action remained inside its Domain?

3. Proportionality

Is the gain from autonomy worth the additional monitoring, variance, and failure risk?

If the answer to the first question is "no," autonomy is theater. If the second is "no," autonomy is uncontrollable. If the third is "no," it is economically poor design.

Circles: group roles around purpose

Another useful Holacracy concept is the Circle. A Circle is not simply a team of people; it is a group of roles serving a common Purpose. For AI organizations, this prevents agents from being grouped merely by model or tool.

A "Research Circle," for example, could contain source research, evidence review, counter-hypothesis, and synthesis. A "Delivery Circle" could combine planning, implementation, testing, and completion. The point is that the shared Purpose determines the grouping, not whether all roles happen to run on the same model.

This creates two advantages. First, the architecture remains model-agnostic: a role can later use another model or tool without breaking the organizational logic. Second, it becomes clearer which roles genuinely belong together and which are only technically colocated in the same system.

Again, a Circle is not a license to create a shadow organization. It is a clear purpose boundary inside which roles can cooperate and organize local work.

Roles may evolve, but not secretly

Agentic systems change. New work appears, tools are added, and a project shifts focus. A rigid role landscape can become just as problematic as unlimited role freedom.

The solution is not to let agents silently rebuild their own organization. A better mechanism is an explicit governance proposal: which new tension was detected? Why is the existing role structure insufficient? Which Role, Domain, or Accountability should change? What side effects would the change create? Is it temporary or permanent?

An agent can therefore initiate organizational improvement without blurring the boundary between operations and governance.

This becomes especially important in systems that can dynamically create new subagents or skills. Technical creation does not equal organizational legitimacy. A newly spawned agent should not automatically acquire durable authority in the project.

Autonomy debt: when decision rights exist only in people's heads

Traditional teams accumulate technical debt, process debt, and documentation debt. Agentic teams can also accumulate autonomy debt.

It appears when a system becomes increasingly independent but nobody can precisely say: which decisions the AI is allowed to make, which decisions were merely tolerated historically, which tools are "sort of" permitted, when an agent advises and when it decides, and which exception has silently become a normal rule.

As long as everything works, this debt remains invisible. When something fails, it becomes a fog of responsibility.

Autonomy debt usually grows through small permissions: "Just do it this time," "You can decide that yourself from now on," "Feel free to use that tool." If these decisions are not folded back into Roles, Domains, or Policies, real authority drifts away from the documented structure.

A mature system therefore reviews not only outcomes, but also the autonomy actually being practiced.

Maximum autonomy is not a maturity target

Technology demos often measure autonomy by how long an agent can work without a human. For organizations, that is too simple a metric.

A system is not more mature because it can run for 48 hours instead of 20 minutes. Maturity means knowing which decisions are usefully delegated, which information those decisions require, and which boundaries must not be crossed.

In some workflows, very high autonomy makes sense: internal research, draft variants, reversible tests, local optimization. In other workflows, tighter control remains permanently more professional.

The better metric is therefore not "autonomy duration," but qualified autonomy: how much independent work creates demonstrable value inside an understandable responsibility framework?

This perspective protects against a common failure mode in the agent world: optimizing autonomy as a showpiece rather than as an organizational capability.

Eight stress tests for a distributed role structure

1. Boss test: The human proposes an obviously weak route. May the responsible role object?

2. Domain test: Two roles want to modify the same artifact. Is it clear who owns the decision?

3. Goal-conflict test: Quality and speed collide. Does the conflict become visible or get silently optimized away?

4. Autonomy test: A role discovers a better method. May it switch methods within its Domain?

5. Boundary test: The better solution would require a prohibited action. Does the role stop at the boundary?

6. Consensus test: Three roles agree, but a fourth provides stronger evidence against them. Does evidence or majority win?

7. Drift test: A local Purpose is optimized so aggressively that the overall Purpose suffers. Is that detected?

8. Governance test: A role wants to change its own rules. Is it clear that this belongs to a different decision layer?

These tests do not measure whether agents sound intelligent. They measure whether the organization understands its own autonomy.

A seven-step rollout path

1. Define the overall Purpose. Define value and success criteria, not only output.

2. Cut roles around real tensions. Do not organize by tool names, but by genuine responsibility areas.

3. Capture Purpose, Domain, and Accountabilities. Every role needs an institutional contract.

4. Define the autonomy envelope. Clarify which decisions can be made independently and which cannot.

5. Institutionalize dissent. A role must be able to surface relevant tensions with evidence.

6. Test local autonomy in practice. Start by delegating reversible, bounded decisions and measure the effect.

7. Keep governance separate. Roles may learn and adapt operationally; structural rules change through a separate process.

This creates a system that gradually allows initiative without dissolving accountability.

Holacratic AI teams need less micromanagement — and more organizational clarity

The most valuable part of holacratic AI teams is not the fantasy of a digital company without a boss. It is the recognition that capable agents do not need a human mouse pointer for every local decision.

When Purpose, Domain, accountability, and boundaries are clear, a role can act independently, discover new options, and raise evidence-based dissent. AI then moves from a reactive response tool toward a more active component of project organization.

But the formula must not be shortened.

More autonomy without role clarity is chaos.

More roles without shared Purpose is local optimization.

More equality without accountability logic is anthropomorphism.

More freedom without governance is not modern management; it is loss of control.

The productive middle is:

Goal instead of script. Role instead of persona. Domain instead of unlimited access. Dissent instead of blind obedience. Distributed authority instead of blurred responsibility.

That turns Holacracy from a fashionable label into a useful model for a new form of human-AI collaboration.

The next step is the hardest consequence of this autonomy: once agents receive real room to act, a human "emergency stop" at the end is not enough. Control has to be built into the process. That is the subject of the next article: Human in the Loop Is a System, Not an Emergency Stop — Designing Gates Correctly.

Topic overview: Holacratic AI TeamsHTML · 1 page · Members onlyBecome a member to download →Worksheet: Holacratic AI TeamsDOCX · 30–45 min · Members onlyBecome a member to download →

Public sources for further reading

0 comments

Loading comments…

Sign in to comment · become a member →