All field notes

Organizational risk

Every Agent Can Be Within Policy—and the Organization Can Still Be Overexposed

Per-action authorization cannot see the aggregate exposure created by concurrent autonomous work. Governed runtimes need a risk-budget ledger above the policy engine.

MP
Max PerfiljevFounder & CEO, AES · Architect of Autonomous Organizations
Read in Russian

An agent can be authorized to issue a refund, change a supplier record, contact a customer, or pause a production process. A thousand agents can each receive the same answer: permitted. The organization can still be taking more risk than it is prepared to bear.

This is the boundary between action policy and organizational risk appetite. A policy engine asks whether a proposed operation is allowed for this agent, this purpose, this data, and this moment. That decision remains necessary. But it is not sufficient when many individually valid operations are in flight at once.

A governed autonomous organization needs a second control: a risk-budget ledger that measures and constrains the portfolio of work already authorized to proceed. Its question is different: if this action joins the work currently underway, does the organization remain inside its approved operating envelope?

Permission is local; exposure is systemic

NIST’s AI Risk Management Framework defines risk tolerance as the level of risk an organization or AI actor is ready to bear in pursuit of its objectives. It also makes an important qualification: tolerance is contextual, specific to the use case, and likely to change over time.

That is not a property an individual action can determine alone. An action may be harmless in isolation and unacceptable in concentration. Fifty customer-account changes aimed at unrelated cases are not the same as fifty changes involving one customer segment, one jurisdiction, one supplier, or one defective upstream data source. The authorization may be identical; the organizational exposure is not.

NIST SP 800-30 calls this risk aggregation: combining discrete risks to assess overall organizational risk. It warns that separately assessed risks can exceed organizational capacity when they materialize concurrently, or when the same risk recurs. Correlation, shared causes, and repeated effects matter. Risk is not reliably additive.

For autonomous systems, concurrency makes this problem operational rather than theoretical. Agents do not merely make more decisions. They can make permitted decisions in parallel, at machine speed, against the same vulnerable process. A policy engine that sees only one request at a time can produce a portfolio no leader intended to authorize.

A ledger above the policy engine

The proposed architecture is not a requirement prescribed by NIST, the Basel Committee, Google SRE, or Kubernetes. It is a synthesis of established control patterns. The central design move is to place an organizational risk-budget ledger above per-action authorization and connect it to the execution path.

Before consequential work proceeds, the runtime estimates the action’s projected exposure and attempts to reserve capacity from relevant budgets. This is not a universal risk score. The ledger should maintain multiple budgets because financial loss, customer impact, legal exposure, operational disruption, and concentration are not naturally interchangeable.

A refund action, for example, might reserve expected financial exposure, a customer-contact capacity, and a concentration allowance for a campaign or customer cohort. A supplier-record update might consume operational and counterparty-concentration capacity. Some constraints will be quantitative. Others may be hard qualitative boundaries: no autonomous action in a particular jurisdiction, no more actions against a named counterparty until review, or no simultaneous changes to a critical process.

The runtime should record both the proposed measurement and the basis for it: the action type, scope, counterparty or process, expected magnitude, relevant policy version, and the budget state used for the decision. The objective is not to claim precise prediction. It is to make exposure visible, reservable, and governable before it becomes irreversible.

The operating cycle

  1. Leadership sets appetite, limits, ownership, review periods, and escalation rules. This is a governance decision, not a model output.
  2. The runtime independently meters projected exposure when consequential work is about to proceed. It checks several interacting budgets, including concentration and shared-dependency limits.
  3. If capacity is available, the runtime reserves it and binds the reservation to the work or business effect. A successful policy decision alone does not create unlimited portfolio capacity.
  4. As outcomes settle, the ledger reconciles the reservation. It releases unused capacity, realizes measured exposure where appropriate, and records exceptions or uncertainty.
  5. When a limit approaches exhaustion or a forward-looking breach is likely, the operating policy changes: throttle work, reduce an agent’s autonomy, narrow scope, queue the action, or escalate it to a human.

This resembles resource quota enforcement more than a reporting dashboard. Kubernetes ResourceQuota records hard limits and observed usage for a scope, then rejects new objects whose requested resources would exceed the quota. The analogy is useful because enforcement happens at admission, not after capacity has already been consumed. A risk-budget ledger should use the same basic posture: reserve before work, not merely count after harm.

Budgets must account for work that has not settled

A ledger cannot wait for final outcomes if the purpose is to govern a live organization. Many effects settle later. A payment may be disputed days after release. A customer contact may contribute to a complaint pattern only after a campaign has run. An operational change may reveal a shared dependency failure after several agents have acted.

The ledger therefore needs explicit states. At minimum, it must distinguish projected exposure, reserved exposure, realized exposure, released capacity, and unresolved exposure. It also needs expiry and reconciliation rules. Otherwise, abandoned work holds capacity forever; conversely, premature release lets the organization spend the same risk allowance twice.

This is where the design becomes a runtime concern rather than an enterprise-risk report. Reservations must be tied to durable work identities and business effects. They must survive retries, compensation, reassignment, and delayed settlement. If the system cannot tell whether exposure remains outstanding, it cannot truthfully tell whether capacity remains available.

Error budgets offer a useful discipline, not a complete model

Google SRE error budgets turn an agreed reliability objective into a measurable allowance for failure over a defined period. When the budget is exhausted, an organization can halt releases and redirect effort toward reliability. That is not a full framework for business or AI risk. It is, however, a strong operational idea: a declared tolerance should change what the system is allowed to do.

An autonomous organization should make the same connection. If the customer-impact budget for an outreach process is nearly exhausted, agents should not continue at their ordinary rate simply because each message passes content and consent checks. If exposure is concentrated in one process or counterparty, the correct response may be to stop adding correlated work even while capacity remains elsewhere.

This produces a more truthful form of autonomy. Agents retain authority within a live organizational envelope, rather than receiving a static permission that ignores what their peers are doing.

Do not centralize judgment into a fictional number

The main failure mode is false precision. Organizations should not compress every potential harm into one score and pretend that a unit of legal exposure can be exchanged cleanly for a unit of customer harm. NIST’s discussion of aggregation points toward the opposite lesson: relationships among risks matter, including correlation and cause-and-effect.

The ledger should therefore support separate dimensions, explicit concentration rules, qualitative prohibitions, and conservative handling of uncertainty. A budget can be temporarily reduced when measurement confidence is weak. A shared upstream dependency can trigger a cross-budget constraint. A hard boundary can override apparent room in every numerical allowance.

Nor should the ledger replace action policy. Policy determines whether an action is legitimate in principle. The risk-budget layer determines whether the organization has capacity to absorb that legitimate action now. These are distinct decisions, made from distinct evidence, and both belong in the execution path.

The architectural consequence

Organizations adopting autonomous agents should stop treating authorization as the last governance decision before execution. It is one decision in a larger control sequence. After an action is found permissible, the runtime still needs to ask what it adds to the active portfolio, which limits it draws down, how long that exposure remains reserved, and what behavior changes when a limit is near breach.

The result is not an attempt to eliminate risk. NIST’s definition of tolerance makes clear that risk is borne in pursuit of objectives. The purpose is to ensure that the organization—not the accidental timing of thousands of local approvals—decides how much risk it is carrying.

A policy engine can authorize an action. Only a risk-budget ledger can say whether the organization can afford all of its authorized actions together.

BUILD WITH AES

Turn architecture into an operating company.

AES connects strategy, tasks, organizational memory, knowledge, agents, people and approvals in one execution environment.