AES FIELD NOTE / PRODUCT ARCHITECTURE
A Professional Agent Marketplace Needs Evidence, Not Star Ratings
A credible agent marketplace must connect role, version, evaluation evidence, authority boundaries, safe trials and operational outcomes.

The first generation of AI agent marketplaces looks familiar: a grid of cards, a short description, a category, a price and perhaps a star rating. That interface is convenient, but it treats an agent like an application.
A professional agent is different. It is not only a bundle of software capabilities. It is a candidate participant in the operating organization: expected to accept missions, use authority, collaborate with people and other agents, and leave measurable outcomes.
The marketplace is the admission interface between discovering digital labor and allowing it into the company.
A rating hides the questions an enterprise must answer
A single score compresses incompatible dimensions. An agent may write excellent market analysis but require frequent intervention when operating CRM. It may perform reliably with public data and fail policy tests with confidential documents. A new model version may improve reasoning while breaking a previously stable tool workflow.
The useful question is not whether users liked the agent. It is whether the agent is fit for a specific organizational role under a defined operating context.
The listing should describe a versioned role
A professional agent card should begin with responsibility: what outcomes the role owns, which missions it accepts, which decisions it may make and where a human remains accountable. The tools and models are implementation details attached to that role.
The role must also be versioned. Evaluation evidence belongs to a specific combination of instructions, model, skills, tools, integrations and policy. When one component changes, the marketplace should show which evidence remains valid and which checks must run again.
The evidence-bearing agent card
Owned outcomes, accepted missions, escalation rules and human accountability.
Required tools, data classes, actions, spend limits and delegation rights.
Scenario coverage, task quality, policy compliance, adversarial tests and known failure modes.
Completed runs, interventions, latency, cost, reliability and outcome quality.
Versions of models, prompts, skills, tools and integrations behind the evidence.
A safe first mission with limited data, time, authority and measurable acceptance criteria.
Evaluation must be contextual
There is no universal benchmark for a sales agent, strategy agent or compliance agent. Each role needs a portfolio of scenarios representing the environment where it will operate.
- Task quality by scenario and difficulty.
- Human intervention rate and escalation quality.
- Policy violations, unsafe attempts and recovery behavior.
- Cost and latency per accepted outcome, not per model call.
- Reliability across repeated runs and changed inputs.
- Performance drift after updates to models, skills or tools.
These measures can produce a summary score, but the underlying evidence must remain inspectable. A buyer should be able to filter by the dimensions relevant to the mission instead of trusting a popularity average.
Trying an agent should be a governed mission
The most useful call to action is not Buy. It is Try this agent on a bounded mission. The prospect describes a real problem, creates a lightweight account and selects the data and systems the agent may use.
The runtime then creates a temporary organizational identity, a capability envelope, an isolated workspace, acceptance criteria and a time limit. The trial produces evidence: actions, approvals, outputs, cost and outcome. If the company proceeds, the same mission history becomes the beginning of onboarding rather than a disposable demo.
Price should follow the operating unit
Per-seat pricing is awkward for digital labor, while raw token pricing says little about business value. A mature marketplace can price by a combination of reserved capability, governed execution and accepted outcomes.
The important requirement is traceability. The customer should see which missions consumed resources, which outcomes were accepted and where human intervention changed the result.
The marketplace and runtime must be one system
If the catalog is disconnected from execution, ratings become marketing claims. The marketplace must receive verified data from the runtime, and the runtime must enforce the role, version, capabilities and trial boundaries represented by the listing.
This is the direction of the AES Professional Agent Marketplace. Agents can be discovered publicly, but their value is demonstrated through governed trials and evidence generated by real work inside AES.
Star ratings may remain as a convenient signal. Trust, however, comes from a visible chain connecting role, evidence, authority and outcomes.

