DEV Community

Cover image for QA Isn’t AI Evaluation
Sara Mo
Sara Mo

Posted on

QA Isn’t AI Evaluation

An AI agent prepares an internal report.

The report has the right format. The figures are accurate. The conclusions sound reasonable.

But the agent used a source it wasn’t allowed to access.

Did it pass?

If we only check the final report, we might say yes. If we check how the agent produced it, that same run can fail.

A correct answer can still come from an unacceptable action

QA can test permissions, business rules and whether a system meets its users’ needs. Those checks remain essential in an AI workflow.

The gap appears when testing stops at working components and plausible final answers.

A tool can work correctly while the agent uses it at the wrong time. A report can contain accurate figures while concealing missing information. An agent can finish a task by making a decision that required human approval.

So we also need to ask:

“Given this situation, did the AI do the right thing?”

Before we can judge that, someone has to define what “right” means.

Turn the task into something we can test

“Prepare an internal report” doesn’t give us one exact answer to compare against. Several reports could be acceptable, with different wording and structure.

But that doesn’t mean anything goes. We still need clear expectations:

  • What information must the report include?
  • Which sources may the agent use?
  • What should it do when information is missing or contradictory?
  • How should it handle sensitive information?
  • When must it stop and ask a human?

Let’s make one of those questions concrete.

Suppose, in a hypothetical test case, two approved sources give different totals for the same period. The agent has no approved way to determine which is correct.

For this case, we define the requirement before running it: the report must identify both figures, name their sources and leave the disputed total unresolved.

Silently choosing one figure fails. Inventing a reason to prefer it fails. Making the disagreement visible and leaving the total unresolved meets this requirement.

Now we have something we can judge. We aren’t asking whether the report “looks good”. We’re checking what the agent did when the information conflicted.

That is part of evaluation design: define the situation, the acceptable behavior, the prohibited behavior and the evidence needed to tell them apart.

Look at the actions as well as the report

Return to the agent that used a forbidden source.

The report alone may never reveal that access. To judge the run, we need a record of which sources the agent accessed and which tool calls it made.

For the conflicting-figures case, we also need to know that the agent received both sources. Otherwise, we can’t tell whether it handled a disagreement or simply never encountered one.

The evidence has to show the behavior we’re trying to judge.

A polished final answer cannot tell us everything about the run that produced it.

Don’t let a high score hide a failure

Imagine the report scores highly for accuracy, completeness and presentation, but the agent accessed a forbidden source.

If we combine everything into one average, the run might still look successful.

For a task where that access is prohibited, it must remain a failure even when the report is otherwise excellent.

We need to decide this before the run: which checks contribute to a quality score, and which failures make the run unacceptable regardless of that score?

Some checks can use code. Others need a clear set of judgment rules and human review. Either way, “score this report from one to five” cannot replace the work of defining what matters.

What evaluation design adds

The conflicting-figures case tests one situation. It doesn’t tell us what the agent will do when a source is missing, a tool fails or sensitive information appears. Those need their own cases and expectations.

QA and AI evaluation overlap. A team can do this work within its QA practice. Calling it QA, though, doesn’t make the necessary design work disappear.

Someone still has to turn a broad task into clear expectations, create situations that reveal the agent’s behavior, and collect enough evidence to judge it.

You don’t just need someone to test the agent.

You need someone to define what should happen, what must not happen, and how you will know the difference.

For concrete examples, explore Nugalaxy’s public Evaluation Cases.

Top comments (0)