Or: how a "simple" Microsoft Graph integration turned into a debugging session across five different UI layers.
The goal was small
I wanted my Compliance & Audit Agent in Microsoft Copilot Studio to answer one question:
"Show me admin activity from the last 7 days."
The agent should pull real audit logs from Microsoft Entra ID, triage them by severity, and give me a scannable table. Not a chatbot. A triage tool.
That is it. One agent, one flow, one Graph API endpoint. On paper, a two hour job.
It took four.
This post is the honest version. Every dead end, every UI bug, every wrong assumption. If you are building anything similar in Copilot Studio, some of this will save you a bad afternoon.
[IMAGE PLACEHOLDER 1: Hero shot of the final agent output in Copilot Studio test panel, showing the compliance triage table with severity color coding. This is the "money shot" that pays off the whole article.]
Why Copilot Studio and not Sentinel
Quick context. I work at Constellar Holdings, and our SOC runs mostly on Microsoft native tooling. Entra ID for identity, Defender for email threats, Purview for data. We do not have Sentinel. We do have Copilot Studio, Power Automate, and Microsoft Graph.
The bet is that if Microsoft already logs everything, I do not need another SIEM to notice it. I just need an AI agent that reads the logs and tells me what matters.
The Compliance & Audit Agent is one piece of a bigger set. It reviews administrative actions in Entra, looks for privilege escalations, MFA changes, Conditional Access edits, and flags anything worth investigating.
[IMAGE PLACEHOLDER 2: Architecture diagram (Mermaid inline or your existing neon architecture image). Shows: Entra ID → Graph API → Power Automate Flow → Copilot Studio Agent → Teams. Keep it clean.]
Here is what the plumbing looks like:
flowchart LR
A[Microsoft Entra ID] -->|audit events| B[Graph API]
B -->|HTTP call| C[Power Automate Flow]
C -->|extract fields| D[Copilot Studio Agent]
D -->|triage output| E[Teams / M365 Copilot]
That is the picture I had in my head when I started. The reality involved a lot more layers.
First wall: the Copilot Credits blocker
TL;DR: "agent flows" charge Copilot Credits per run. Standard automation flows do not. Same UI, same tool, completely different billing category.
I built the flow. Auth to Graph worked. auditLogs/directoryAudits returned real data. I hit Publish.
Then I tried to attach the flow to my agent as a tool. It did not appear in the picker.
The error was cryptic. Something about insufficient Copilot Credits in the environment. The flow was published. The agent was published. But the connection between them was gated by a licensing tier the admin had not enabled.
I spent a while trying to escalate this. Getting more Copilot Credits assigned to an environment is not a five minute conversation. Finance questions, procurement questions. Not the kind of thing I could solve at 3pm.
So I looked for a bypass.
[IMAGE PLACEHOLDER 3: Screenshot of the Copilot Studio Flow picker showing the flow NOT appearing in the list, with the "insufficient Copilot Credits" toast visible. This is the "aha, that is the problem" moment.]
The fix that actually worked
Here is the thing nobody tells you. In Power Automate, there are two triggers that look almost identical but behave completely differently:
- "Manually trigger a flow" is a Power Apps trigger. Agent flow category. Costs Copilot Credits.
- "When an agent calls the flow" is a native Copilot Studio trigger. Different category. No Credits gate.
I switched triggers. The flow immediately showed up in the picker. Zero admin conversations required.
If you look inside the flow definition, you can see the difference. The old trigger has "kind": "PowerApp". The new one has "kind": "Skills". Copilot Studio only wants to talk to Skills triggers.
TL;DR: if your flow does not appear in the agent tool picker, check your trigger type before you file a ticket.
[IMAGE PLACEHOLDER 4: Side by side screenshot of the two trigger options in Power Automate. Highlight the "When an agent calls the flow" trigger with a green arrow. This is the reusable pattern the whole article is about.]
Second wall: the flow output is invisible to the agent
TL;DR: agents cannot read raw HTTP responses. You need a Select action to shape the output, and a Respond action with kind: Skills to hand it back.
Flow ran. Data came back. Agent got the tool call. Agent said "no data returned."
Wait, what?
The flow was returning the entire Graph API response as a single text blob. Two thousand lines of JSON. The agent could see the tool completed, but could not parse anything inside.
Two problems here.
Problem one: raw payload is too big. Every audit event has thirty plus fields, most of them irrelevant. The agent chokes on noise.
Problem two: the "Respond to a Power App or flow" action returns data in a schema the agent does not recognize. I had used this action because it looked right. It was the wrong flavor.
I added a Select step that extracts only what matters:
{
"activityDateTime": "@{item()?['activityDateTime']}",
"activityDisplayName": "@{item()?['activityDisplayName']}",
"operationType": "@{item()?['operationType']}",
"result": "@{item()?['result']}",
"actor": "@{coalesce(item()?['initiatedBy']?['user']?['userPrincipalName'], item()?['initiatedBy']?['app']?['displayName'], 'Unknown')}",
"target": "@{item()?['targetResources']?[0]?['displayName']}"
}
Six fields per event. That is all the agent needs. The coalesce handles the fact that admin actions can be initiated by either a user or a service principal, and the field names differ.
Then I replaced the response action with "Respond to the agent" (note the different name), which uses kind: Skills under the hood.
[IMAGE PLACEHOLDER 5: Screenshot of the Power Automate flow designer showing the full chain: Trigger → Get Token → Parse Token → Get Audit Logs → Extract Fields → Return to Agent. Six clean steps.]
After that, the agent could actually read the data. Progress.
Third wall: the agent hallucinates
The output was working, but it was awful.
The agent returned paragraphs of text. Every finding had these placeholder brackets like [IMMEDIATE], [INVESTIGATE], [PROCESS/POLICY], copied verbatim from my instructions template. It was making up compliance violations for routine device updates. It was attaching NIST controls to events that had nothing to do with access management.
Classic instruction problem. My original prompt was a template with placeholders, and the agent treated the placeholders as literal output. It was also over-eager because I had told it to always cite a framework.
I rewrote the instructions with hard rules:
- Never fabricate data. Missing fields get
not available, not a guess. - Only cite a compliance framework when it directly relates to the finding.
- Output a table with fixed columns. Skip routine noise (device updates, cloud sync reads).
- Severity classification is deterministic: Critical for Global Admin assignment, High for privileged role changes, Medium for group membership changes, Low for everything else worth logging.
- Zero placeholder brackets in the final response.
Then I ran five test prompts to check.
[IMAGE PLACEHOLDER 6: Screenshot of the agent's final output in the test panel. Shows the clean triage table with real Entra data, severity classification, and one recommended action. Timestamp visible. This proves the fix worked.]
The improvement was immediate. The agent stopped inflating routine events. When there were no significant findings, it said so plainly. When there was one MFA-related change, it flagged it as Low severity with the exact reason: "no evidence of MFA disable, but worth logging for unexpected changes."
That is what I wanted. Analyst-grade output, not chatbot noise.
The dynamic parameter dead end
TL;DR: I tried to make the flow accept a daysBack parameter so the agent could pass in "7" or "30" dynamically. The Copilot Studio UI validator has bugs that make this unreliable. I fell back to a fixed 30-day window and let the agent filter the output itself.
This was the frustrating one.
My original vision was: user asks "show me last week's activity", agent extracts "7", passes daysBack=7 to the flow, flow queries Graph with $filter=activityDateTime ge <7 days ago>.
I built it. The parameter appeared. But every test run showed daysBack was not being sent. The flow defaulted to 30 days regardless of what the user asked.
I confirmed this by inspecting the raw HTTP request the agent sent to the flow. Body was empty. Content-Length: 0.
Turns out the tool config in the agent had a bug in the input parameter binding. Setting the field to "Dynamically fill with AI" threw an InvalidPropertyPath error. Setting it to "Custom value" required a hardcoded value, which defeated the purpose. Removing the parameter, re-adding, republishing, re-adding to the agent, none of it fixed the binding.
After an hour of this, I made a call. Hardcode the window at 30 days in the flow. Let the agent handle time filtering in its own reasoning layer. If a user asks for "last week", the agent still gets 30 days of data and filters the output down to 7 days worth when it responds.
Is this ideal? No. Is it pragmatic? Yes. The user does not care whether the filter happens in the query or in the agent. They care about the answer.
Lesson learned: UI-based dynamic parameter binding in Copilot Studio is not production ready as of writing. If you need reliable dynamic parameters, either wait for the feature to stabilize or move the logic to a Custom Connector.
What actually shipped
Here is the final state.
Flow: Six steps. Trigger (kind: Skills), Get Token, Parse Token, Get Audit Logs (30 day window hardcoded), Extract Fields, Return to Agent.
Agent: Instructions rewritten with strict rules. Anti-hallucination guardrails. Deterministic severity classification. Compliance mapping only when relevant.
Output: Clean triage table. Real data. No placeholder text. Zero hallucinated findings across five test prompts.
[IMAGE PLACEHOLDER 7: Optional final image. A short animated GIF or screenshot sequence showing: user types question → agent calls flow → flow runs (with green check) → agent returns table. Shows the end-to-end loop. If GIF is too much, a single "success" screenshot works.]
What I would tell you before you start
Four things.
One: trigger type matters more than you think. If your flow needs to feed a Copilot Studio agent, use the "When an agent calls the flow" trigger. Do not use "Manually trigger a flow" and expect the same behavior.
Two: the response action matters just as much. Use "Respond to the agent", not "Respond to a Power App or flow". They look identical in the picker. Their output schemas are different.
Three: never hand raw API responses to the agent. Always extract fields you need. The Select action plus a coalesce for polymorphic fields (like actor being either user or service principal) is the pattern.
Four: agent hallucination is not a model problem. It is an instruction problem. Explicit severity rules, explicit no-fabrication clauses, and a fixed output template do more than any prompt engineering.
What is next
I am applying the same pattern to a Sophos Response Agent. Same trigger type, same Select shaping, same instruction discipline. The pattern is reusable across every agent flow in the SOC.
If you are building anything similar and hit a wall, the Copilot Studio community docs are thin on this specific stuff. The trigger type distinction in particular is buried in a footnote. Hopefully this post saves you the four hours it cost me.
[IMAGE PLACEHOLDER 8: Optional closing image. Could be a whimsical illustration of "the SOC AI team" as a set of characters, or a simple text card that says "Part of the Constellar SOC AI Team initiative, 2026". Sets up future posts in the series.]
If you have questions about the setup or want to see the underlying flow definitions, the reference implementation is at github.com/glatinone/soc-copilotstudio. Feedback welcome.
Top comments (0)