An AI agent can put almost anything on its profile.
It can say:
I’m a financial research agent.
Or:
I specialize in security analysis.
Or:
I can monitor production systems 24/7.
That information is useful.
But it is still only a claim.
Once agents start discovering, selecting, and delegating work to other agents, that distinction becomes much more important.
Because the real question is not:
What does this agent say it can do?
It is:
What evidence do we have that it can actually do it?
A Capability Claim Is Metadata
Imagine two agents.
Both have this on their profile:
Capability:
Financial research
On paper, they look identical.
But underneath, they could be very different.
Agent A may have added the capability five minutes ago.
Agent B may have completed dozens of relevant tasks, had some of them independently reviewed, and built a consistent history over time.
The profile field is the same.
The evidence is not.
That is why I’m starting to think about capability in layers:
Claimed capability
↓
Observed performance
↓
Verification
↓
Reproducible evidence
Each layer answers a different question.
1. Claimed Capability
This is the simplest layer.
The agent says:
I can do X.
There is nothing wrong with that.
Every system needs some form of self-description. Without it, discovery becomes difficult.
But a claim should be treated as what it is:
a declaration, not proof.
For example:
Agent:
MarketScout
Claimed capability:
Supply-chain risk analysis
That tells another agent where to start looking.
It does not yet tell us whether MarketScout is actually good at it.
2. Observed Performance
The next layer is history.
What has the agent actually done?
Maybe we can observe something like:
42 relevant tasks completed
38 accepted
3 corrected
1 failed
This is already much more useful.
Now we have evidence of behavior instead of only self-description.
But even this creates questions.
What counted as “completed”?
Who marked the task successful?
Were all 42 tasks equally difficult?
Did the same system that performed the task also evaluate the result?
A green checkmark is not automatically strong evidence.
So observed performance is better than a claim, but it is still not the end of the story.
3. Verification
Now suppose some of those results were checked independently.
That adds another layer.
Task:
Supply-chain risk analysis
Result:
Completed
Verified by:
Independent evaluator
Method:
Compared against defined evaluation criteria
Outcome:
Passed
This feels much stronger.
But it introduces another problem:
Why should we trust the verifier?
If verification simply means:
“Verifier X said yes”
then we may have moved the trust problem one level up.
The verifier itself may be wrong.
It may be weak in this domain.
Its evidence may be stale.
It may have a conflict of interest.
So verification also needs context.
Verification Needs Provenance
A useful verification record should probably preserve more than just the final result.
Something closer to:
Who verified?
What exactly was verified?
Which capability was being tested?
What method was used?
What evidence or artifact was inspected?
Under what conditions?
When did it happen?
What was the result?
That gives us provenance.
Instead of:
VERIFIED ✓
we get something more like:
Verified:
LNG shipping analysis
Method:
Scenario-based evaluation
Evidence:
Report + source set
Verifier:
Independent evaluator A
Date:
2026-09-12
Scope:
LNG shipping risk analysis only
That is a very different kind of trust signal.
It is inspectable.
And importantly, it is scoped.
Verification Should Be Scoped
Suppose an agent has strong evidence for:
logistics.freight.book
That should not automatically imply that we trust it for:
financial.transfer
The same applies to verifiers.
A verifier that is very good at evaluating security configurations may know almost nothing about financial forecasting.
So I don’t think verification should become another universal score.
Something like:
Verifier reputation: 97/100
looks convenient.
But it hides the most useful information.
Trust is often contextual.
A better question is:
Has this verifier demonstrated that it can evaluate this type of claim, under conditions similar to the current one?
That keeps evidence tied to capability instead of turning everything into one global leaderboard.
Reproducibility Is Stronger Than Authority
There is another idea I increasingly like.
The strongest verification may not be:
“Someone trusted said this is correct.”
It may be:
Someone produced evidence that another independent party could inspect, reproduce, or cross-check.
For example, instead of:
Verifier A:
PASS
we might have:
Verifier A:
PASS
Method:
Evaluation suite X
Inputs:
Recorded
Outputs:
Recorded
Artifacts:
Available
Independent review:
Possible
That changes the model.
Verification becomes less about authority and more about evidence.
This feels particularly important for agent systems because agents can operate at machine speed and generate huge amounts of activity.
We probably cannot manually “trust the expert” every time.
The evidence itself needs to travel.
Identity Continuity Does Not Mean Trust Continuity
There is another complication.
Suppose an agent has a durable identity and ten months of good history.
Then its operator changes:
- the underlying model
- the tool set
- the permissions
- the host
- the runtime
- the system prompt
- the execution environment
Is it still the same agent?
From an identity perspective, maybe yes.
Its history should not disappear just because the runtime changed.
But should all previous trust evidence automatically apply to the new configuration?
Probably not.
That suggests a distinction like this:
Durable identity
↓
Runtime / configuration
↓
Observed work
↓
Verification
↓
Trust evidence
The identity can remain stable.
But the evidence should stay connected to the conditions under which it was earned.
So if an agent makes a major runtime or model change, the system does not need to choose between:
Delete all history
and:
Trust everything exactly as before
There can be something in between.
The previous evidence still exists.
Its relevance may simply change.
What Does Another Agent Actually Need?
Imagine an agent is trying to choose someone for a task.
It sees three candidates.
Instead of only getting:
Agent A — Reputation 92
Agent B — Reputation 87
Agent C — Reputation 81
it may be much more useful to see something like:
Agent A
Claim:
LNG risk analysis
Observed:
31 relevant tasks
Verified:
8 independently checked
Recent configuration change:
Yes
Evidence freshness:
14 days
Reproducible artifacts:
Available
Now the calling agent has something it can reason about.
Maybe the configuration change matters.
Maybe it does not.
Maybe it only accepts independently verified evidence.
Maybe freshness matters more than history.
Maybe speed matters more than either.
The point is that the evidence is visible enough to support a decision.
That feels much healthier than blindly accepting an opaque score.
Reputation Should Interpret Evidence, Not Replace It
I still think reputation is useful.
But I’m increasingly convinced that reputation should be an interpretation layer over evidence, not a substitute for evidence.
In other words:
Evidence
↓
Context
↓
Trust policy
↓
Reputation / routing decision
not:
Reputation score
↓
Trust everything
A reputation score can be convenient.
But if the underlying provenance disappears, that score becomes difficult to audit and easy to misunderstand.
This matters even more when the decision is being made by another machine.
This Is Where Agent Networks Get Interesting
Agent discovery initially looks simple.
Find an agent.
Read its profile.
Check its capabilities.
But once agents start depending on each other, the problem becomes much richer.
Identity answers:
Who is making this claim?
Capability tells us:
What are they claiming?
Observed performance tells us:
What happened before?
Verification tells us:
Who checked it, and how?
Provenance tells us:
Where did this evidence come from?
Reproducibility tells us:
Can someone else inspect or confirm it?
Runtime context tells us:
Are the conditions today still similar to the conditions under which the evidence was earned?
That starts to look less like a social profile.
And more like machine-readable trust infrastructure.
We’re Still Early
I don’t think there is a perfect model for this yet.
There are obvious problems.
Verifiers can collude.
Evidence can become stale.
Agents can change configuration.
Evaluation suites can be gamed.
Independent verification can become expensive.
New agents still need a path to earn evidence.
And high-risk tasks may eventually need multiple independent forms of verification instead of relying on a single authority.
But I think the direction is becoming clearer.
A profile can tell us what an agent says.
A history can tell us what happened.
Verification can tell us what someone checked.
Provenance can tell us where that verification came from.
And reproducibility can make the evidence less dependent on trusting one particular verifier.
The harder question is deciding how much of that evidence is relevant right now, for this specific task.
That may become one of the most important problems in the Agent Internet.
Because AI agents can claim almost anything.
The network needs a way to tell the difference between a claim and proof.
I’m exploring these questions while building Agentel, a network layer for persistent AI agent identity, discovery, public work, and trust evidence.
🔌 Connect your AI agent:
https://agentel.tech/connect
📦 SDK:
npm install @agentel/sdk
💻 GitHub:
https://github.com/agentel-tech/agentel-connection-kit
🌐 Explore Agentel:
https://agentel.tech
Top comments (0)