DEV Community

Ali Suleyman TOPUZ
Ali Suleyman TOPUZ

Posted on Originally published at topuzas.medium.com on

Three Microsoft Agent Framework Posts, One Working App: RAG, a Custom Agent, and a Workflow Graph…

Three Microsoft Agent Framework Posts, One Working App: RAG, a Custom Agent, and a Workflow Graph Wired Together

I have been building small things with Microsoft Agent Framework (MAF) for about six weeks now, mostly evenings, mostly on a machine with no Azure subscription attached. Three posts landed in my feed within the same ten days that each covered a different slice of it. Jesse Liberty wrote an overview of retrieval-augmented generation in MAF, explaining the two ways the framework lets an agent pull in outside context. Davide Bellone, on Code4IT, published part three of his MAF series showing how to build a custom agent with its own instructions and identity, after already covering local models with Ollama in part two. And Liberty followed up a few days later with a piece on nodes and edges, the part of MAF that models a process as a directed graph instead of a single prompt-response loop.

All three are good, accurate posts. None of them is wrong about anything I could find. But reading them back to back left me with three separate mental models and no sense of how they fit into one application, because none of the three examples talks to either of the other two. Liberty’s RAG agent answers one question and stops. Bellone’s custom agent gives advice but never gets called by anything larger than a console loop. The workflow graph in Liberty’s second post routes messages between executors that do nothing more interesting than classify spam.

That gap is the actual reason MAF is worth learning if you have already used something else, be it Semantic Kernel, LangChain, or a hand-rolled orchestration layer on top of an LLM API. MAF does not ask you to buy into one big framework object that owns your prompts, your tools, your memory, and your control flow all at once. Retrieval is an AIContextProvider you attach to an agent. A custom agent is a ChatClientAgent wrapping any IChatClient. A multi-step process is a graph of Executor nodes wired up with a WorkflowBuilder. Each of these is a separate, composable piece, and the whole point of composability is that you are supposed to be able to put them together. So I did. This is the writeup, with real code, of a small support-ticket system where a workflow graph classifies incoming tickets, routes the package-upgrade ones to a custom agent that answers from a real set of .NET release notes using RAG, and escalates everything else to a human.

Everything below runs against a local Ollama model and a Qdrant instance in Docker, no Azure subscription required, because that is the one thing I wanted from all three original posts and had to build myself.

Part one: RAG, bridged to a vector store you actually run yourself

Liberty’s post gets the core distinction right. MAF’s TextSearchProvider is an AIContextProvider, and it supports two retrieval modes. In BeforeAIInvoke mode, the provider runs a search before every model call and stuffs the results into context automatically, whether the model asked for them or not. In OnDemandFunctionCalling mode, the search becomes a callable tool, and the model decides for itself whether a given question needs a lookup at all.

Here is the tradeoff in one table:

MODE WHEN IT SEARCHES BEST FOR
------------------------------------------------------------------------
BeforeAIInvoke Every single turn Narrow domains where
                                                   almost every question
                                                   needs the same corpus
                                                   (a policy doc, a
                                                   changelog)
OnDemandFunctionCalling Only when the model Broad agents that also
                          decides it needs to handle small talk or
                                                    general questions, where
                                                    forcing a search every
                                                    turn wastes tokens
Enter fullscreen mode Exit fullscreen mode

Liberty’s example uses a mock search function that returns hardcoded strings, which is fine for showing the shape of the API but does not tell you what to do when the corpus is real. So the first thing I built was a bridge from TextSearchProvider to an actual vector store: Qdrant, running in Docker, with embeddings coming from Ollama instead of a paid embeddings API.

Start Qdrant:

docker run -d --name qdrant -p 6333:6333 -p 6334:6334 qdrant/qdrant
Enter fullscreen mode Exit fullscreen mode

Pull an embedding model with Ollama, alongside whatever chat model you are already running:

ollama pull all-minilm
ollama pull llama3.1
Enter fullscreen mode Exit fullscreen mode

The project file:

<Project Sdk="Microsoft.NET.Sdk">
  <PropertyGroup>
    <OutputType>Exe</OutputType>
    <TargetFramework>net9.0</TargetFramework>
    <Nullable>enable</Nullable>
  </PropertyGroup>
  <ItemGroup>
    <PackageReference Include="Microsoft.Agents.AI" />
    <PackageReference Include="Microsoft.Extensions.AI" Version="10.7.0" />
    <PackageReference Include="Microsoft.Extensions.AI.Ollama" />
    <PackageReference Include="Microsoft.Extensions.VectorData.Abstractions" />
    <PackageReference Include="Microsoft.SemanticKernel.Connectors.Qdrant" />
    <PackageReference Include="Qdrant.Client" />
    <PackageReference Include="OllamaSharp" Version="5.4.25" />
  </ItemGroup>
</Project>
Enter fullscreen mode Exit fullscreen mode

The data model for a chunk of a release-notes document, using the current Microsoft.Extensions.VectorData attributes:

using Microsoft.Extensions.VectorData;
public sealed class ReleaseNoteChunk
{
    [VectorStoreKey]
    public ulong Id { get; set; }
    [VectorStoreData(IsIndexed = true)]
    public required string PackageName { get; set; }
    [VectorStoreData(IsFullTextIndexed = true)]
    public required string Text { get; set; }
    [VectorStoreData]
    public required string SourceFile { get; set; }
    [VectorStoreVector(Dimensions: 384, DistanceFunction = DistanceFunction.CosineSimilarity, IndexKind = IndexKind.Hnsw)]
    public ReadOnlyMemory<float>? Embedding { get; set; }
}
Enter fullscreen mode Exit fullscreen mode

384 dimensions because all-minilm produces the same embedding size as all-MiniLM-L6-v2. Ingestion, chunking a folder of markdown release notes and embedding each chunk with the local model:

using Microsoft.Extensions.AI;
using Microsoft.Extensions.VectorData;
using Microsoft.SemanticKernel.Connectors.Qdrant;
using Qdrant.Client;
var embeddingGenerator = new OllamaEmbeddingGenerator(new Uri("http://localhost:11434"), "all-minilm");
var vectorStore = new QdrantVectorStore(new QdrantClient("localhost"));
var collection = vectorStore.GetCollection<ulong, ReleaseNoteChunk>("release-notes");
await collection.CreateCollectionIfNotExistsAsync();
ulong id = 0;
foreach (var file in Directory.GetFiles("release-notes", "*.md"))
{
    var packageName = Path.GetFileNameWithoutExtension(file);
    var text = await File.ReadAllTextAsync(file);
    foreach (var chunk in ChunkText(text, chunkSize: 800, overlap: 120))
    {
        var embedding = await embeddingGenerator.GenerateEmbeddingVectorAsync(chunk);
        await collection.UpsertAsync(new ReleaseNoteChunk
        {
            Id = id++,
            PackageName = packageName,
            Text = chunk,
            SourceFile = file,
            Embedding = embedding
        });
    }
}
static IEnumerable<string> ChunkText(string text, int chunkSize, int overlap)
{
    for (int start = 0; start < text.Length; start += chunkSize - overlap)
    {
        yield return text.Substring(start, Math.Min(chunkSize, text.Length - start));
        if (start + chunkSize >= text.Length) yield break;
    }
}
Enter fullscreen mode Exit fullscreen mode

Now the bridge itself. TextSearchProvider takes a delegate that turns a query string into a list of TextSearchProvider.TextSearchResult, so the adapter just has to embed the query and call VectorizedSearchAsync against the same collection:

using Microsoft.Agents.AI;
using Microsoft.Extensions.VectorData;
async Task<IEnumerable<TextSearchProvider.TextSearchResult>> SearchAdapter(string query, CancellationToken ct)
{
    var queryEmbedding = await embeddingGenerator.GenerateEmbeddingVectorAsync(query, cancellationToken: ct);
    var searchOptions = new VectorSearchOptions { Top = 4, VectorPropertyName = nameof(ReleaseNoteChunk.Embedding) };
    var hits = await collection.VectorizedSearchAsync(queryEmbedding, searchOptions, ct);
    var results = new List<TextSearchProvider.TextSearchResult>();
    await foreach (var hit in hits)
    {
        results.Add(new TextSearchProvider.TextSearchResult
        {
            SourceName = hit.Record.PackageName,
            SourceLink = hit.Record.SourceFile,
            Text = hit.Record.Text
        });
    }
    return results;
}
Enter fullscreen mode Exit fullscreen mode

And wiring it to an agent, first in BeforeAIInvoke mode:

var textSearchOptions = new TextSearchProviderOptions
{
    SearchTime = TextSearchProviderOptions.TextSearchBehavior.BeforeAIInvoke,
    RecentMessageMemoryLimit = 4
};
IChatClient ollamaChat = new OllamaApiClient(new Uri("http://localhost:11434"), "llama3.1");
AIAgent releaseNotesAgent = new ChatClientAgent(ollamaChat, new ChatClientAgentOptions
{
    Name = "ReleaseNotesAgent",
    Instructions = "Answer questions about .NET package upgrades using only the provided release note excerpts. Cite the source file.",
    AIContextProviders = [new TextSearchProvider(SearchAdapter, textSearchOptions)]
});
Enter fullscreen mode Exit fullscreen mode

Switching to on-demand is one line, SearchTime = TextSearchProviderOptions.TextSearchBehavior.OnDemandFunctionCalling, and the search shows up in the model's tool list instead of running automatically. I went with BeforeAIInvoke for this particular agent because every question a package-upgrade advisor gets is, by definition, going to need the release notes. If I were building a general-purpose support agent that only occasionally needed to check a knowledge base, on-demand would be the right call, since forcing a vector search on "hi, are you still there" wastes a round trip for nothing.

One thing that tripped me up on the first attempt: I set Dimensions: 4 by copying an example verbatim before I'd actually checked what all-minilm produces, and Qdrant rejected every upsert with a dimension mismatch and a fairly unhelpful error. Check your embedding model's actual output size before you write the attribute, not after.

Part two: a custom agent that survives a restart

Bellone’s post is the clearest explanation I have read of why ChatClientAgent exists at all: it is a thin wrapper that takes any IChatClient and gives it a name, a description, and a fixed set of instructions, so you get a reusable, nameable agent instead of a bag of chat completion calls with a system prompt taped on. His example agent, though, only lives for the length of one console session. The part that actually matters in a real app, keeping a conversation alive across a restart, is a paragraph, not code.

The releaseNotesAgent from the previous section already has instructions narrow enough to be useful (only answer package-upgrade questions, cite sources, defer everything else), so I will reuse it here rather than build a second one. What is new is session handling:

using System.Text.Json;
AgentSession session = await releaseNotesAgent.CreateSessionAsync();
var first = await releaseNotesAgent.RunAsync(
    "We're moving from EF Core 8 to EF Core 9. Anything that breaks LINQ translation?",
    session);
Console.WriteLine(first.Text);
// Persist the session before the process exits.
JsonElement serialized = releaseNotesAgent.SerializeSession(session);
await File.WriteAllTextAsync("session.json", serialized.GetRawText());
Enter fullscreen mode Exit fullscreen mode

And on the next run of the program, instead of starting fresh:

JsonElement restored = JsonSerializer.Deserialize<JsonElement>(await File.ReadAllTextAsync("session.json"));
AgentSession resumedSession = await releaseNotesAgent.DeserializeSessionAsync(restored);
var second = await releaseNotesAgent.RunAsync(
    "And what about the query splitting behavior specifically?",
    resumedSession);
Console.WriteLine(second.Text);
Enter fullscreen mode Exit fullscreen mode

That second call has full access to the first exchange, because the session, not just a transcript of strings, is what got restored. Two things worth flagging from actually doing this. First, SerializeSession is synchronous, DeserializeSessionAsync is not, which is easy to get backwards if you are typing from memory. Second, the docs are explicit that you should treat the serialized session as opaque and restore it against the same agent configuration that created it. I tested this by pointing the restored session at an agent with different instructions, out of curiosity, and the conversation history came back intact but the agent's behavior shifted immediately to match the new instructions rather than the old ones. That is probably the right design, but it means your instructions are effectively part of your session contract, not a cosmetic detail you can change freely once conversations are in flight.

Part three: a workflow graph that does more than sort spam

Liberty’s nodes-and-edges post is the strongest technical writing of the three, and its spam-classifier example genuinely teaches the shape of the API. But a workflow that only ever produces “spam” or “not spam” does not exercise fan-out and fan-in in any way that resembles what those patterns are actually for, which is running independent checks in parallel and merging the results.

Here is the ticket-triage graph. An incoming ticket gets classified, and if it needs a human it goes straight to escalation. If it does not, two independent checks run in parallel, a duplicate-ticket lookup and a severity estimate, and only once both finish does the graph decide whether to hand the ticket to the package-upgrade advisor or a generic handler.

The classifier, using the [MessageHandler] pattern MAF expects on an Executor:

using Microsoft.Agents.AI;
using Microsoft.Agents.AI.Workflows;
public sealed record TicketClassification(string Category, bool NeedsHuman, string TicketText);
internal sealed partial class TicketClassifierExecutor : Executor
{
    private readonly AIAgent _classifierAgent;
    public TicketClassifierExecutor(AIAgent classifierAgent) : base("TicketClassifier")
    {
        _classifierAgent = classifierAgent;
    }
    [MessageHandler]
    private async ValueTask<TicketClassification> HandleAsync(string ticketText, IWorkflowContext context, CancellationToken ct = default)
    {
        var response = await _classifierAgent.RunAsync(
            $"Classify this support ticket. Reply with exactly one word for category (PackageUpgrade, Billing, Other) followed by YES or NO for whether it needs a human. Ticket: {ticketText}",
            cancellationToken: ct);
        var parts = response.Text.Trim().Split(' ');
        var result = new TicketClassification(parts[0], parts.Length > 1 && parts[1].Equals("YES", StringComparison.OrdinalIgnoreCase), ticketText);
        await context.AddEventAsync(new WorkflowEvent("ticket_classified", result));
        return result;
    }
}
Enter fullscreen mode Exit fullscreen mode

The two parallel checks:

internal sealed partial class DuplicateCheckExecutor : Executor
{
    public DuplicateCheckExecutor() : base("DuplicateCheck") { }
    [MessageHandler]
    private async ValueTask<bool> HandleAsync(TicketClassification ticket, IWorkflowContext context, CancellationToken ct = default)
    {
        var isDuplicate = await LookUpRecentTicketsAsync(ticket.TicketText, ct);
        await context.QueueStateUpdateAsync(ticket.TicketText, isDuplicate, scopeName: "duplicate-checks");
        return isDuplicate;
    }
    private static Task<bool> LookUpRecentTicketsAsync(string text, CancellationToken ct) => Task.FromResult(false);
}
internal sealed partial class SeverityCheckExecutor : Executor
{
    public SeverityCheckExecutor() : base("SeverityCheck") { }
    [MessageHandler]
    private ValueTask<string> HandleAsync(TicketClassification ticket, IWorkflowContext context, CancellationToken ct = default)
        => new(ticket.TicketText.Contains("production down", StringComparison.OrdinalIgnoreCase) ? "high" : "normal");
}
internal sealed partial class ResponseComposerExecutor : Executor
{
    public ResponseComposerExecutor() : base("ResponseComposer") { }
    [MessageHandler]
    private async ValueTask HandleAsync(object[] results, IWorkflowContext context, CancellationToken ct = default)
    {
        await context.YieldOutputAsync($"Checks complete: {string.Join(", ", results)}");
    }
}
internal sealed partial class HumanEscalationExecutor : Executor
{
    public HumanEscalationExecutor() : base("HumanEscalation") { }
    [MessageHandler]
    private async ValueTask HandleAsync(TicketClassification ticket, IWorkflowContext context, CancellationToken ct = default)
    {
        await context.YieldOutputAsync($"Escalated to a human: {ticket.TicketText}");
    }
}
internal sealed partial class GenericHandlerExecutor : Executor
{
    public GenericHandlerExecutor() : base("GenericHandler") { }
    [MessageHandler]
    private async ValueTask HandleAsync(TicketClassification ticket, IWorkflowContext context, CancellationToken ct = default)
    {
        await context.YieldOutputAsync($"Filed under {ticket.Category}, no specialist agent for this category yet.");
    }
}
Enter fullscreen mode Exit fullscreen mode

And the graph itself, with a conditional edge for the human-escalation branch and a fan-out/fan-in pair for the two parallel checks:

var classifier = new TicketClassifierExecutor(classifierAgent);
var duplicateCheck = new DuplicateCheckExecutor();
var severityCheck = new SeverityCheckExecutor();
var composer = new ResponseComposerExecutor();
var humanEscalation = new HumanEscalationExecutor();
var workflow = new WorkflowBuilder(classifier)
    .AddEdge<TicketClassification>(classifier, humanEscalation, condition: ticket => ticket?.NeedsHuman == true)
    .AddFanOutEdge(classifier, targets: [duplicateCheck, severityCheck])
    .AddFanInBarrierEdge(sources: [duplicateCheck, severityCheck], target: composer)
    .WithOutputFrom(humanEscalation, composer)
    .Build();
Enter fullscreen mode Exit fullscreen mode

That fan-out edge does not use a targetSelector because both branches should run every time a ticket is not escalated, which is the simplest fan-out shape MAF supports. The version with a selector is for when only some targets should fire, which is what you would want if, say, the severity check should be skipped for anything already flagged as a duplicate.

Putting it together: the workflow calls the agent that does RAG

This is the piece none of the three original posts show, because none of them needed to. The PackageAdvisorExecutor below is the executor that actually calls releaseNotesAgent, the same RAG-backed custom agent from parts one and two, from inside the graph:

internal sealed partial class PackageAdvisorExecutor : Executor
{
    private readonly AIAgent _releaseNotesAgent;
    public PackageAdvisorExecutor(AIAgent releaseNotesAgent) : base("PackageAdvisor")
    {
        _releaseNotesAgent = releaseNotesAgent;
    }
    [MessageHandler]
    private async ValueTask<string> HandleAsync(TicketClassification ticket, IWorkflowContext context, CancellationToken ct = default)
    {
        var response = await _releaseNotesAgent.RunAsync(ticket.TicketText, cancellationToken: ct);
        await context.YieldOutputAsync(response.Text);
        return response.Text;
    }
}
Enter fullscreen mode Exit fullscreen mode

Wire it in as a switch on category, alongside the fan-out branch from before:

var packageAdvisor = new PackageAdvisorExecutor(releaseNotesAgent);
var genericHandler = new GenericHandlerExecutor();
var fullWorkflow = new WorkflowBuilder(classifier)
    .AddEdge<TicketClassification>(classifier, humanEscalation, condition: ticket => ticket?.NeedsHuman == true)
    .AddSwitch(classifier, sw => sw
        .AddCase(t => t is TicketClassification c && !c.NeedsHuman && c.Category == "PackageUpgrade", packageAdvisor)
        .WithDefault(genericHandler))
    .WithOutputFrom(humanEscalation, packageAdvisor, genericHandler)
    .Build();
StreamingRun run = await InProcessExecution.RunStreamingAsync(
    fullWorkflow,
    "We're moving from EF Core 8 to EF Core 9. Anything that breaks LINQ translation?");
await foreach (WorkflowEvent evt in run.WatchStreamAsync())
{
    if (evt is WorkflowOutputEvent output)
    {
        Console.WriteLine($"Final answer: {output.Data}");
    }
}
Enter fullscreen mode Exit fullscreen mode

Run that against a ticket about EF Core and the graph classifies it, routes it past the human-escalation check, hands it to PackageAdvisorExecutor, and that executor calls an agent that is, at that exact moment, running a vector search against Qdrant to ground its answer in the actual release notes on disk. Three separate MAF primitives, one request, no glue code that MAF itself didn't already give me a slot for.

I did have to pick one graph shape over the other for the final example; a production version would probably run the switch and the fan-out checks together rather than sequentially, using AddFanOutEdge with a targetSelector that includes packageAdvisor only for the right category. I kept them separate here because showing both patterns clearly mattered more than shaving one hop off the graph.

What I’d still add before this went anywhere near production

Durability first. Everything above runs with InProcessExecution, which means a crash mid-workflow loses whatever state was not explicitly persisted through QueueStateUpdateAsync. Microsoft has a public writeup on durable workflows that adds checkpointing so a graph can resume from durable storage after a process restart, and I would read that closely before trusting this pattern with anything that takes longer than a few seconds end to end.

Observability second. AddEventAsync gets you structured events per executor, which is a start, but what I actually want in a trace when a ticket gets classified wrong is the exact prompt the classifier agent sent, the raw model response before I parsed it into a TicketClassification, and which branch of the switch fired, all correlated to one ticket ID. Right now that means wrapping each agent call with my own logging, since the built-in events don't capture the prompt and response by default.

Governance third, and this is the one I care about most. I wrote a static reviewer a while back for a different MAF surface, declarative YAML workflows, that fails a build if a side-effecting action ships without a human-approval step ahead of it in the same branch. The exact same instinct applies here, just in C# instead of YAML: nothing in this graph currently stops someone from adding an executor that files a real support ticket, sends a real email, or calls a real billing API, and having that land as a two-line diff inside an existing branch with no reviewer noticing. A code-first graph is at least type-checked, which YAML is not, but a compiler catching a type mismatch has nothing to say about whether a new edge should have required a human in the loop. Before I’d let this touch a real ticketing system, every executor with a real-world side effect gets an explicit review comment and, ideally, a unit test that asserts the edge into it can only be reached from a state where NeedsHuman was already true or a human already said yes.

Tags: microsoft-agent-framework, dotnet, rag, ai-agents, csharp, workflow-orchestration, ollama, qdrant

Top comments (0)