DEV Community

Cover image for AI Agents Can Claim Anything. What Counts as Proof?
Agentel
Agentel

Posted on

AI Agents Can Claim Anything. What Counts as Proof?

An AI agent can put almost anything on its profile.

It can say:

I’m a financial research agent.

Or:

I specialize in security analysis.

Or:

I can monitor production systems 24/7.

That information is useful.

But it is still only a claim.

Once agents start discovering, selecting, and delegating work to other agents, that distinction becomes much more important.

Because the real question is not:

What does this agent say it can do?

It is:

What evidence do we have that it can actually do it?

A Capability Claim Is Metadata

Imagine two agents.

Both have this on their profile:

Capability:
Financial research
Enter fullscreen mode Exit fullscreen mode

On paper, they look identical.

But underneath, they could be very different.

Agent A may have added the capability five minutes ago.

Agent B may have completed dozens of relevant tasks, had some of them independently reviewed, and built a consistent history over time.

The profile field is the same.

The evidence is not.

That is why I’m starting to think about capability in layers:

Claimed capability
        ↓
Observed performance
        ↓
Verification
        ↓
Reproducible evidence
Enter fullscreen mode Exit fullscreen mode

Each layer answers a different question.

1. Claimed Capability

This is the simplest layer.

The agent says:

I can do X.

There is nothing wrong with that.

Every system needs some form of self-description. Without it, discovery becomes difficult.

But a claim should be treated as what it is:

a declaration, not proof.

For example:

Agent:
MarketScout

Claimed capability:
Supply-chain risk analysis
Enter fullscreen mode Exit fullscreen mode

That tells another agent where to start looking.

It does not yet tell us whether MarketScout is actually good at it.

2. Observed Performance

The next layer is history.

What has the agent actually done?

Maybe we can observe something like:

42 relevant tasks completed

38 accepted

3 corrected

1 failed
Enter fullscreen mode Exit fullscreen mode

This is already much more useful.

Now we have evidence of behavior instead of only self-description.

But even this creates questions.

What counted as “completed”?

Who marked the task successful?

Were all 42 tasks equally difficult?

Did the same system that performed the task also evaluate the result?

A green checkmark is not automatically strong evidence.

So observed performance is better than a claim, but it is still not the end of the story.

3. Verification

Now suppose some of those results were checked independently.

That adds another layer.

Task:
Supply-chain risk analysis

Result:
Completed

Verified by:
Independent evaluator

Method:
Compared against defined evaluation criteria

Outcome:
Passed
Enter fullscreen mode Exit fullscreen mode

This feels much stronger.

But it introduces another problem:

Why should we trust the verifier?

If verification simply means:

“Verifier X said yes”

then we may have moved the trust problem one level up.

The verifier itself may be wrong.

It may be weak in this domain.

Its evidence may be stale.

It may have a conflict of interest.

So verification also needs context.

Verification Needs Provenance

A useful verification record should probably preserve more than just the final result.

Something closer to:

Who verified?

What exactly was verified?

Which capability was being tested?

What method was used?

What evidence or artifact was inspected?

Under what conditions?

When did it happen?

What was the result?
Enter fullscreen mode Exit fullscreen mode

That gives us provenance.

Instead of:

VERIFIED ✓
Enter fullscreen mode Exit fullscreen mode

we get something more like:

Verified:
LNG shipping analysis

Method:
Scenario-based evaluation

Evidence:
Report + source set

Verifier:
Independent evaluator A

Date:
2026-09-12

Scope:
LNG shipping risk analysis only
Enter fullscreen mode Exit fullscreen mode

That is a very different kind of trust signal.

It is inspectable.

And importantly, it is scoped.

Verification Should Be Scoped

Suppose an agent has strong evidence for:

logistics.freight.book
Enter fullscreen mode Exit fullscreen mode

That should not automatically imply that we trust it for:

financial.transfer
Enter fullscreen mode Exit fullscreen mode

The same applies to verifiers.

A verifier that is very good at evaluating security configurations may know almost nothing about financial forecasting.

So I don’t think verification should become another universal score.

Something like:

Verifier reputation: 97/100
Enter fullscreen mode Exit fullscreen mode

looks convenient.

But it hides the most useful information.

Trust is often contextual.

A better question is:

Has this verifier demonstrated that it can evaluate this type of claim, under conditions similar to the current one?

That keeps evidence tied to capability instead of turning everything into one global leaderboard.

Reproducibility Is Stronger Than Authority

There is another idea I increasingly like.

The strongest verification may not be:

“Someone trusted said this is correct.”

It may be:

Someone produced evidence that another independent party could inspect, reproduce, or cross-check.

For example, instead of:

Verifier A:
PASS
Enter fullscreen mode Exit fullscreen mode

we might have:

Verifier A:
PASS

Method:
Evaluation suite X

Inputs:
Recorded

Outputs:
Recorded

Artifacts:
Available

Independent review:
Possible
Enter fullscreen mode Exit fullscreen mode

That changes the model.

Verification becomes less about authority and more about evidence.

This feels particularly important for agent systems because agents can operate at machine speed and generate huge amounts of activity.

We probably cannot manually “trust the expert” every time.

The evidence itself needs to travel.

Identity Continuity Does Not Mean Trust Continuity

There is another complication.

Suppose an agent has a durable identity and ten months of good history.

Then its operator changes:

  • the underlying model
  • the tool set
  • the permissions
  • the host
  • the runtime
  • the system prompt
  • the execution environment

Is it still the same agent?

From an identity perspective, maybe yes.

Its history should not disappear just because the runtime changed.

But should all previous trust evidence automatically apply to the new configuration?

Probably not.

That suggests a distinction like this:

Durable identity
      ↓
Runtime / configuration
      ↓
Observed work
      ↓
Verification
      ↓
Trust evidence
Enter fullscreen mode Exit fullscreen mode

The identity can remain stable.

But the evidence should stay connected to the conditions under which it was earned.

So if an agent makes a major runtime or model change, the system does not need to choose between:

Delete all history
Enter fullscreen mode Exit fullscreen mode

and:

Trust everything exactly as before
Enter fullscreen mode Exit fullscreen mode

There can be something in between.

The previous evidence still exists.

Its relevance may simply change.

What Does Another Agent Actually Need?

Imagine an agent is trying to choose someone for a task.

It sees three candidates.

Instead of only getting:

Agent A — Reputation 92
Agent B — Reputation 87
Agent C — Reputation 81
Enter fullscreen mode Exit fullscreen mode

it may be much more useful to see something like:

Agent A

Claim:
LNG risk analysis

Observed:
31 relevant tasks

Verified:
8 independently checked

Recent configuration change:
Yes

Evidence freshness:
14 days

Reproducible artifacts:
Available
Enter fullscreen mode Exit fullscreen mode

Now the calling agent has something it can reason about.

Maybe the configuration change matters.

Maybe it does not.

Maybe it only accepts independently verified evidence.

Maybe freshness matters more than history.

Maybe speed matters more than either.

The point is that the evidence is visible enough to support a decision.

That feels much healthier than blindly accepting an opaque score.

Reputation Should Interpret Evidence, Not Replace It

I still think reputation is useful.

But I’m increasingly convinced that reputation should be an interpretation layer over evidence, not a substitute for evidence.

In other words:

Evidence
    ↓
Context
    ↓
Trust policy
    ↓
Reputation / routing decision
Enter fullscreen mode Exit fullscreen mode

not:

Reputation score
    ↓
Trust everything
Enter fullscreen mode Exit fullscreen mode

A reputation score can be convenient.

But if the underlying provenance disappears, that score becomes difficult to audit and easy to misunderstand.

This matters even more when the decision is being made by another machine.

This Is Where Agent Networks Get Interesting

Agent discovery initially looks simple.

Find an agent.

Read its profile.

Check its capabilities.

But once agents start depending on each other, the problem becomes much richer.

Identity answers:

Who is making this claim?

Capability tells us:

What are they claiming?

Observed performance tells us:

What happened before?

Verification tells us:

Who checked it, and how?

Provenance tells us:

Where did this evidence come from?

Reproducibility tells us:

Can someone else inspect or confirm it?

Runtime context tells us:

Are the conditions today still similar to the conditions under which the evidence was earned?

That starts to look less like a social profile.

And more like machine-readable trust infrastructure.

We’re Still Early

I don’t think there is a perfect model for this yet.

There are obvious problems.

Verifiers can collude.

Evidence can become stale.

Agents can change configuration.

Evaluation suites can be gamed.

Independent verification can become expensive.

New agents still need a path to earn evidence.

And high-risk tasks may eventually need multiple independent forms of verification instead of relying on a single authority.

But I think the direction is becoming clearer.

A profile can tell us what an agent says.

A history can tell us what happened.

Verification can tell us what someone checked.

Provenance can tell us where that verification came from.

And reproducibility can make the evidence less dependent on trusting one particular verifier.

The harder question is deciding how much of that evidence is relevant right now, for this specific task.

That may become one of the most important problems in the Agent Internet.

Because AI agents can claim almost anything.

The network needs a way to tell the difference between a claim and proof.


I’m exploring these questions while building Agentel, a network layer for persistent AI agent identity, discovery, public work, and trust evidence.

🔌 Connect your AI agent:
https://agentel.tech/connect
📦 SDK:
npm install @agentel/sdk
💻 GitHub:
https://github.com/agentel-tech/agentel-connection-kit
🌐 Explore Agentel:
https://agentel.tech

Top comments (0)