A team compares two versions of an AI service. Average accuracy rises from 82% to 86%. The new version looks better, so releasing it seems like the obvious decision.
Then an engineer separates the failures by type. The system now makes fewer mistakes on simple questions, but it is twice as likely to answer confidently without a source when the question concerns billing or contracts. The average improved while the product risk increased.
This is where an applied AI engineer needs mathematics. Not as an entrance exam, and not as an attempt to compress a university program into a few weeks. Math is useful here because it connects system behavior to an engineering decision: release the new version, return it for more work, or restrict one scenario.
The most practical way to begin is to build a small evaluation loop. Every formula in it should answer one engineering question.
Define the unit of evaluation first
A weak test set often contains little more than a few pairs of query -> good answer. That is not enough for diagnosis.
For each evaluation case, store at least:
- the user's query;
- the scenario or segment;
- the identifiers of the sources the system should retrieve;
- the acceptable behavior: answer, clarify, or refuse;
- whether the result requires human review;
- the cost of this particular failure.
A plain Python dictionary is enough to represent the first version:
case = {
"query": "Can I change the billing entity after payment?",
"segment": "billing",
"expected_source_ids": ["billing-policy-17"],
"expected_behavior": "answer_with_source",
"human_review_required": False,
"failure_cost": 5,
}
The value of failure_cost is not a universal mathematical truth. It is an explicit team decision. A mistake in a low-risk reference answer might cost 1 point. An unsupported claim about a payment might cost 5. An action with irreversible consequences should be blocked by a separate policy instead of being reduced to another score.
Without this labeling, a metric quickly becomes a number with no owner.
Evaluate retrieval separately from the final answer
In a RAG system, a bad answer can have several causes. The required document never entered the index. Retrieval did not return it near the top. The model received the correct source but ignored it. Or the source itself was obsolete.
If you measure only the final answer, those failures collapse into one result.
A useful first retrieval metric is recall@k: the fraction of required sources that appears in the first k results.
def recall_at_k(cases, retrieved_by_case, k):
recalls = []
for case in cases:
expected = set(case["expected_source_ids"])
retrieved = set(retrieved_by_case[case["query"]][:k])
if expected:
recalls.append(len(expected & retrieved) / len(expected))
if not recalls:
raise ValueError("No cases contain expected_source_ids")
return sum(recalls) / len(recalls)
If recall@5 is low, changing the answer-generation prompt is premature. Check document splitting, metadata, filters, embeddings, and index freshness first.
This is where linear algebra becomes practical. An engineer needs to understand that an embedding is a vector and that cosine similarity, or another distance measure, is only a comparison rule. High similarity does not prove that a document answers the question. It only says that the selected representation and measure consider the objects similar.
Add failure types to average accuracy
Suppose the system classifies incoming requests and routes them to a department. Accuracy alone is insufficient when classes are imbalanced or different mistakes have different consequences.
At minimum, distinguish between:
- a false positive, where the system selects a scenario without enough evidence;
- a missed case, where it fails to recognize the required scenario;
- an unsupported answer;
- an unnecessary refusal, where it could have helped safely;
- a wrong escalation to a person.
You can then calculate not only the fraction of correct results, but also weighted risk:
FAILURE_COST = {
"false_positive": 2,
"missed_case": 3,
"unsupported_answer": 5,
"unnecessary_refusal": 1,
"wrong_escalation": 2,
}
def weighted_risk(results):
if not results:
raise ValueError("Results must not be empty")
total_cost = sum(
FAILURE_COST[result["failure_type"]]
for result in results
if result["failure_type"] is not None
)
return total_cost / len(results)
The process owner should agree on the cost of each failure. The formula does not make the business decision. It makes that decision visible and lets the team compare versions using the same rule.
Inspect segments before trusting the average
An aggregate metric can improve simply because the test set contains many easy cases. Break the result down by a few meaningful segments, such as:
- task type;
- data source;
- language;
- new or returning user;
- low-risk or sensitive scenario;
- refusal type.
Even a simple segment calculation tells you more than one overall number:
from collections import defaultdict
def accuracy_by_segment(results):
grouped = defaultdict(list)
for result in results:
grouped[result["segment"]].append(result["is_correct"])
return {
segment: sum(values) / len(values)
for segment, values in grouped.items()
}
If the overall result improves while a sensitive segment gets worse, that regression should drive the discussion. An average should not override a safety boundary that the team agreed on in advance.
Check whether confidence matches observed quality
Another common mistake is to treat high confidence as evidence that an answer is correct.
A better question is this: when the system reports confidence near 0.8, is it actually correct in roughly 80% of those cases?
To inspect this, group answers into confidence intervals and compare the average reported confidence with the observed fraction of correct results. A large gap indicates poor calibration.
Calibration affects product rules directly:
- at low confidence, the system asks a clarifying question;
- in a sensitive scenario, the result goes to a specialist;
- without a supporting source, the system refuses to answer;
- an automatic action requires an additional check.
You cannot choose the threshold separately from the cost of errors. In one product, an unnecessary refusal is almost harmless. In another, it destroys the product's main function.
A small sample does not justify a large conclusion
Thirty evaluation cases are better than three demonstrations, but they may still represent real traffic poorly.
The minimum discipline is straightforward:
- Do not change the test set between two versions you are comparing.
- Keep a separate set of cases that was not used to tune prompts or rules.
- Do not combine segments without inspecting them separately.
- Repeat unstable checks when the model is nondeterministic.
- Add every important production failure to the regression set.
At this point, an engineer benefits from understanding samples, means, variance, and confidence intervals. The goal is not a more impressive report. It is to avoid calling random movement in a small sample an improvement.
Build the first evaluation loop
You can build a useful first version without a complex platform:
- Collect 30 to 50 cases from real task types.
- Label the expected source, acceptable behavior, segment, and failure cost.
- Run the baseline and the new version on the same set.
- Measure retrieval, final output, refusal, and escalation separately.
- Compare the aggregate metric, segments, weighted risk, and calibration.
- Review the most expensive failures manually.
- Record the release decision together with its reasons and limitations.
This loop gives further mathematical study a concrete context. If retrieval fails, go deeper into vectors, distance measures, and ranking. If the model is confidently wrong, probability and calibration become the next step. If results vary substantially between runs, study sampling and variability more carefully.
What to study after the first loop
A useful sequence might be:
- vectors, matrices, dot products, and normalization;
- conditional probability and basic probabilistic reasoning;
- precision, recall, F1, and ranking metrics;
- samples, means, variance, and confidence intervals;
- loss functions and optimization fundamentals;
- calibration and decision thresholds.
This is no longer an abstract curriculum. Each topic connects to a failure that the engineer can already observe in the evaluation loop.
The point
An applied AI engineer does not need to wait for a complete mathematical education before measuring a system properly.
Start by choosing the unit of evaluation, separating the stages of the system, naming the failure types, and connecting them to consequences. Mathematics then stops acting as a gate. It becomes the tool that explains why a new version is genuinely better, or why a better-looking average still does not justify a release.
Top comments (0)