All field notes

Performance isolation

A Replica Is Not an Isolation Boundary

Separate compute does not make a volatile reader harmless. Isolation exists only until the first resource path it shares with production.

MP
Max PerfiljevFounder & CEO, AES · Architect of Autonomous Organizations
Read in Russian

A new endpoint can be useful. A read replica can be useful. A separate VM, Kubernetes namespace or ephemeral database can be useful. None of them, by themselves, prove that a workload is isolated from production.

That distinction matters as bursty readers multiply: agent-driven exploration, analytics, reporting, simulations, support tools and ad hoc internal queries. These workloads often look safe because they are read-only and run away from the primary application process. But read-only is not contention-free. A workload can avoid write locks and still consume the cache, storage bandwidth, network capacity, metadata service or account-level quota that production needs to meet its SLO.

The right architectural question is not, “Do we have replicas?” It is: “At what layer do the production and exploratory workloads first meet?”

Isolation is a property of a path

Workload isolation is often discussed as a deployment property. One team points to a distinct endpoint; another to a separate compute instance; a third to a namespace or tenant boundary. These are real forms of separation, but they answer only part of the question.

A request takes a path. It starts at a client and passes through compute, caches, network links, storage services, metadata systems, admission controls and quotas. Production and exploratory traffic may travel independently for the first several steps, then converge on a shared resource. That first shared resource is the meaningful boundary of isolation.

If both workload classes depend on the same saturated page server, a reader burst can still slow production. If they compete for a cache tier, exploratory scans can evict useful production data. If they share a network bottleneck, quota or metadata service, distinct database nodes may not prevent an incident from crossing the boundary. Separate compute absorbs compute demand; it does not automatically absorb downstream demand.

This is not a claim that shared components are intrinsically unsafe. Shared resources can be partitioned, rate-limited, scheduled and protected by admission control. The point is narrower: the word “isolated” should describe demonstrated non-propagation of contention, not merely a diagram with separate boxes.

AlloyDB makes the distinction concrete

Google’s September 24 preview of PostgreSQL for agents in AlloyDB is useful because its architectural claim reaches beyond a large number of temporary database nodes. Google says it provisions independent, read-only AlloyDB nodes in microVMs, gives them sub-second-fresh access to production state and scales them down after the work finishes. More importantly, it says agent reads use separate Colossus storage segments and do not traverse production database components. Google frames the separation as extending across compute, network and storage.

The offering is a preview under Google Cloud Pre-GA terms, may have limited support and currently requires an access request. It is not a general availability announcement. Nor should its performance figures be treated as a customer-wide guarantee: Google reports that an internal benchmark scaled from one to 1,000 agent nodes to 3 million aggregate queries per second with no measurable primary-cluster degradation. That is vendor-run evidence, not an independently reproduced result.

Still, the design claim is the important part. Google is not arguing merely that agent traffic has separate compute. It is arguing that the storage path no longer contends with the production database path. Whether a particular deployment achieves the stated result remains a question for testing. But the claim identifies the layer at which a conventional replica design can stop being sufficient.

Read offload and isolation are different products

Conventional replicas remain valuable. They can provide availability, distribute predictable reads and create useful operational separation. But their architecture may deliberately retain shared persistence components. AWS documents that Aurora reader instances connect to the same cluster storage volume as the primary. Microsoft documents that Azure SQL Hyperscale HA and named replicas use the same page servers as the primary, while geo-replicas use a separate set.

Neither fact proves inadequate isolation in any particular Aurora or Azure SQL deployment. Performance depends on the product’s resource partitioning, workload shape, quotas, admission controls and operating configuration. The examples show something more basic: replica types in the same product family can have different shared-fate topologies. “Replica” is therefore not a sufficient answer to a performance-isolation requirement.

This distinction should also prevent a common procurement mistake. A buyer may ask whether a platform supports thousands of readers, receive a credible yes, and still fail to ask what happens to the primary during an adversarial burst. Scale and isolation are related but separate claims. A system can scale readers while retaining one or more shared bottlenecks. It can also isolate a specific path while still sharing broader physical infrastructure. The relevant promise is always bounded by the resource path and the load condition being considered.

Reserve “isolated” for a workload whose contention cannot propagate across the boundary you care about.

Require an isolation-path map

Founders and operators do not need to reverse-engineer every managed service. They do need a concrete artifact for each volatile workload class: an isolation-path map. It should trace the path from a reader to the production state it accesses, and make shared dependencies visible rather than implicit.

  • List compute placement and runtime limits: where do the reader and primary execute, and which CPU, memory or host-level resources can they share?
  • Trace cache and network dependencies: can the workloads contend for cache capacity, connection pools, links, gateways or load-balancing limits?
  • Trace persistence and control dependencies: which storage servers, page services, metadata systems, replication services and recovery paths are common?
  • Identify quotas and admission boundaries: which tenant, project, account or regional limits are consumed by both classes?
  • Mark the first shared saturation point, its capacity limit, its protective mechanism and the team that owns it.

Then test the map. Hold the primary workload at a representative production load. Apply an adversarial reader burst with the query shapes, concurrency and scan patterns that the volatile workload can actually generate. Observe primary latency, error rate, throughput and any resource-specific saturation signal. Repeat at the expected quota and cache boundaries, not only at the reader endpoint.

The test does not have to prove that production can never be affected. It must establish what will happen under the loads the organization is prepared to permit. If contention propagates, the architecture may still be acceptable. Call it read offload, access isolation or compute separation, then apply rate limits and capacity plans accordingly. Honest language gives operators the right operating model.

The bottleneck defines the promise

The arrival of more autonomous and exploratory readers makes this discipline urgent, but the rule is older than AI. Google’s research on shared compute clusters identifies processor caches and memory buses as sources of performance interference. Separation at one layer has never guaranteed separation at every downstream layer.

A production system is isolated from a volatile workload only as far as its resource paths remain independently protected. A separate endpoint is an interface decision. A replica is a topology decision. A separate compute fleet is a capacity decision. They become an isolation decision only when the path to the first meaningful bottleneck is separate—or when the shared bottleneck has a demonstrated mechanism that prevents one workload’s contention from becoming the other’s outage.

That is the standard worth putting into an architecture review. Do not ask whether the reader is separate. Ask where its pressure goes.

BUILD WITH AES

Turn architecture into an operating company.

AES connects strategy, tasks, organizational memory, knowledge, agents, people and approvals in one execution environment.