All field notes

AES FIELD NOTE / PRODUCT ARCHITECTURE

A Professional Agent Marketplace Needs Evidence, Not Star Ratings

A credible agent marketplace must connect role, version, evaluation evidence, authority boundaries, safe trials and operational outcomes.

MP
Max PerfiljevFounder & CEO, AES · Architect of Autonomous Organizations
Read in Russian

The first generation of AI agent marketplaces looks familiar: a grid of cards, a short description, a category, a price and perhaps a star rating. That interface is convenient, but it treats an agent like an application.

A professional agent is different. It is not only a bundle of software capabilities. It is a candidate participant in the operating organization: expected to accept missions, use authority, collaborate with people and other agents, and leave measurable outcomes.

The marketplace is the admission interface between discovering digital labor and allowing it into the company.

A rating hides the questions an enterprise must answer

A single score compresses incompatible dimensions. An agent may write excellent market analysis but require frequent intervention when operating CRM. It may perform reliably with public data and fail policy tests with confidential documents. A new model version may improve reasoning while breaking a previously stable tool workflow.

The useful question is not whether users liked the agent. It is whether the agent is fit for a specific organizational role under a defined operating context.

The listing should describe a versioned role

A professional agent card should begin with responsibility: what outcomes the role owns, which missions it accepts, which decisions it may make and where a human remains accountable. The tools and models are implementation details attached to that role.

The role must also be versioned. Evaluation evidence belongs to a specific combination of instructions, model, skills, tools, integrations and policy. When one component changes, the marketplace should show which evidence remains valid and which checks must run again.

The evidence-bearing agent card

Role contract

Owned outcomes, accepted missions, escalation rules and human accountability.

Capability envelope

Required tools, data classes, actions, spend limits and delegation rights.

Evaluation suite

Scenario coverage, task quality, policy compliance, adversarial tests and known failure modes.

Operational record

Completed runs, interventions, latency, cost, reliability and outcome quality.

Provenance

Versions of models, prompts, skills, tools and integrations behind the evidence.

Trial boundary

A safe first mission with limited data, time, authority and measurable acceptance criteria.

Evaluation must be contextual

There is no universal benchmark for a sales agent, strategy agent or compliance agent. Each role needs a portfolio of scenarios representing the environment where it will operate.

  • Task quality by scenario and difficulty.
  • Human intervention rate and escalation quality.
  • Policy violations, unsafe attempts and recovery behavior.
  • Cost and latency per accepted outcome, not per model call.
  • Reliability across repeated runs and changed inputs.
  • Performance drift after updates to models, skills or tools.

These measures can produce a summary score, but the underlying evidence must remain inspectable. A buyer should be able to filter by the dimensions relevant to the mission instead of trusting a popularity average.

Trying an agent should be a governed mission

The most useful call to action is not Buy. It is Try this agent on a bounded mission. The prospect describes a real problem, creates a lightweight account and selects the data and systems the agent may use.

The runtime then creates a temporary organizational identity, a capability envelope, an isolated workspace, acceptance criteria and a time limit. The trial produces evidence: actions, approvals, outputs, cost and outcome. If the company proceeds, the same mission history becomes the beginning of onboarding rather than a disposable demo.

Price should follow the operating unit

Per-seat pricing is awkward for digital labor, while raw token pricing says little about business value. A mature marketplace can price by a combination of reserved capability, governed execution and accepted outcomes.

The important requirement is traceability. The customer should see which missions consumed resources, which outcomes were accepted and where human intervention changed the result.

The marketplace and runtime must be one system

If the catalog is disconnected from execution, ratings become marketing claims. The marketplace must receive verified data from the runtime, and the runtime must enforce the role, version, capabilities and trial boundaries represented by the listing.

This is the direction of the AES Professional Agent Marketplace. Agents can be discovered publicly, but their value is demonstrated through governed trials and evidence generated by real work inside AES.

Star ratings may remain as a convenient signal. Trust, however, comes from a visible chain connecting role, evidence, authority and outcomes.

BUILD WITH AES

Turn architecture into an operating company.

AES connects strategy, tasks, organizational memory, knowledge, agents, people and approvals in one execution environment.