DEV Community

Cover image for Using Custom Metrics to Improve Agentic Experiences
Bala Madhusoodhanan
Bala Madhusoodhanan

Posted on AI-assisted

Using Custom Metrics to Improve Agentic Experiences

Intro:
Custom metrics in Microsoft Copilot Studio help teams understand the quality and outcomes of agent conversations—not simply how many conversations take place. They enable process owners and business teams to track the operational signals that matter for their particular agentic use case, such as whether users receive a resolved answer, where essential information is missing, when source content is insufficient, or when a request needs to move to a different support channel.

As outlined in Microsoft’s Custom metrics documentation, metrics can be configured using natural language. The metric owner describes the outcome they want to measure across conversations, then defines the result categories that the agent should identify. This makes the feature accessible to business and process teams, without requiring them to create a complex reporting data model or introduce manual tagging of every conversation.

Where custom metrics add value
For agentic use cases, standard analytics—such as number of sessions, escalation rate, or engagement—show activity, but may not explain why an agent did or did not help the user. Custom metrics add that business context.
They can help a process team or business owner understand:

  • Resolution quality: whether users receive a supported, complete response.
  • Information gaps: whether the agent needs essential facts before it can proceed.
  • Knowledge and content gaps: whether approved sources contain sufficient information to answer the question.
  • Data-quality issues: whether a lookup, mapping, or reference data prevents a tailored answer.
  • Demand for unsupported services: whether users expect the agent to carry out actions outside its current scope.
  • Improvement priorities: whether the right intervention is a conversation-design change, content update, data correction, integration enhancement, or clearer user guidance.

The important point is that each metric should answer one practical operational question. Categories should be clear, distinct, and based on outcomes that can be recognised in the conversation.

Example: Travel & Expense policy assistant

A Travel & Expense (T&E) policy assistant provides a useful example. Its role is to answer employee questions about T&E policy using approved policy documentation. For personal policy guidance, the assistant may need to establish the employee’s employment or base country, validate the applicable policy route, retrieve the relevant evidence, and determine whether the request is a policy question or an operational request such as making a booking.

This means that an apparently simple question—such as “Can I travel business class?”—may have several possible outcomes:

The assistant can provide a grounded answer from the relevant policy.
It needs a missing fact, such as the employee’s base country or band.
The country-to-policy mapping cannot be validated.

The policy documentation does not adequately answer the question.
The user is requesting an action that needs a service desk or booking channel.

Example Metric 1: Context and Routing Readiness

Element Definition
Purpose Identifies whether the assistant can establish the context needed to give a tailored answer.
What do you want to measure? Whether personal policy questions can be successfully routed to the correct policy scope. Measure whether the user provides the necessary context and whether the required policy mapping or reference-data lookup is successfully validated, unavailable, disabled, duplicated, malformed, or missing.
Result category Definition
Successfully routed The conversation concerns a personal policy question and the assistant confirms that the required context and valid policy route are available, allowing it to provide an appropriately tailored answer.
Essential context needed The conversation concerns a personal policy question, but the user has not provided a necessary fact required to determine their applicable policy or entitlement. The assistant asks for that fact.
Routing or reference-data issue The user provides the required context, but the assistant cannot validate the applicable route because the required lookup or mapping is missing, disabled, duplicated, incomplete, malformed, unavailable, or returns an error.
General information request The user asks a general policy or document-information question that does not require personal context or tailored policy routing.

Example Metric 2: Answer Resolution Outcome

Element Definition
Purpose Shows whether the assistant resolved the substantive user question and, where it did not, identifies the reason.
What do you want to measure? The outcome of substantive user questions. Classify whether the assistant provides a grounded answer using applicable approved sources, needs an essential clarification, cannot answer because relevant source content is missing or incomplete, or directs the user to an appropriate operational or human-support channel for requests outside its capability.
Result category Definition
Resolved with approved evidence The assistant answers the substantive question using applicable approved source content, including material conditions, approvals, limitations, or citations where available.
Clarification required The assistant cannot make a case-specific determination because an essential fact is missing. It asks a concise question for the information required to continue.
Knowledge or evidence gap The assistant has sufficient user context, or the question is general, but cannot establish an answer because the relevant approved source content is missing, incomplete, unavailable, conflicting, or does not address the requested point.
Operational or human support needed The user requests an action outside the assistant’s capability, or requires human assistance. The assistant directs the user to the appropriate approved support route.

What this tells the business owner

In the T&E example, the desired outcome is an increasing proportion of Resolved with approved evidence. The remaining categories indicate where improvement should be focused:

Metric outcome Likely focus area Example in the T&E agent
Clarification required Conversation design and user guidance Explain earlier that country or HR band may be needed
Knowledge or evidence gap Policy content and knowledge governance Update a missing local policy rule or clarify conflicting guidance
Operational or human support needed Service design and integration roadmap Improve booking or service-desk hand-off; assess future integration
Routing or reference-data issue Data ownership and technical operations Correct country-policy mappings or resolve lookup access issues

Closing remarks:

Custom metrics help process teams shift from asking, “Is the agent being used?” to asking, “Is the agent producing the intended business outcome, and what prevents it when it does not?”

The T&E policy assistant is only one example. The same pattern can be applied to procurement, order-to-cash, supply-chain, HR, IT support, customer service, or any agent that depends on user context, business rules, enterprise knowledge, and hand-offs to other services.

As the feature is currently in preview, custom metrics initially display a rolling seven-day view because the analysis is computationally intensive. Once sufficient data has been collected, and hopefully have a view expanding to show up to 90 days of metric history.

Teams should begin with a weekly review cadence, monitor meaningful changes in the category mix, and assign improvement actions to the appropriate owner—such as the agent owner, process team, policy/content owner, data steward, or service owner.

Important: Custom metrics are calculated from the conversation transcript. The outcome described in “What do you want to measure?” must therefore be visible or clearly evidenced in the conversation—for example, through the user’s request, the agent’s response, or an explicit outcome such as a resolved answer, clarification request, routing failure, or escalation.

Top comments (0)