Five write-ups, one week, and the patterns that kept showing up when the hype was stripped out
I spent a week reading everything I could find about shipping AI features in .NET, and I noticed something a bit humbling. The articles I trusted most were almost never about the model. They were about the boring layer around it: where the abstraction boundary sits, which NuGet packages you are still allowed to use for free, how you stop a code-generating agent from quietly wrecking your architecture, and how you keep a text-to-SQL bot from running DROP TABLE because someone phrased a question creatively.
So this is a survey, not a tutorial. I am pulling from five write-ups (listed at the bottom), adding my own opinions and doubts, and sketching the code where it helps. One honest note before we start: the code below is my own sketch of the patterns, not copied from those articles, and I have not run every snippet on every platform. Where I am extrapolating, I will say so.
Here is the map:
PATTERN WHAT IT PROTECTS YOU FROM
-------------------------------------------------------------------------
M.E.AI abstraction boundary Model and vendor lock-in
Non-silent fallback policy Surprise cloud calls and privacy leaks
Package guardrails Licensing shocks, hallucinated packages
Architecture fitness tests AI PRs that pass tests but rot the design
Validated text-to-SQL pipeline Writes, injections, and confident nonsense
1. Draw the boundary at Microsoft.Extensions.AI
The single most repeated idea across the sources was also the least exciting one: put Microsoft.Extensions.AI (M.E.AI) at the edge of your application and keep every provider SDK below it. Your app code talks to IChatClient and IEmbeddingGenerator>. Adapters for OpenAI, Azure, Ollama, or something on the phone sit underneath.
I was skeptical of this for a while. Abstractions over fast-moving APIs have a bad track record, and I have been burned by "we will just swap the provider later" more than once. What changed my mind is how thin the surface is. IChatClient is basically messages in, response out, plus streaming, plus a GetService escape hatch. There is not much to get wrong, and the middleware pipeline (logging, caching, function invocation, telemetry) composes on top of it instead of being smeared across your providers.
// dotnet add package Microsoft.Extensions.AI
// dotnet add package Azure.AI.OpenAI
// dotnet add package Microsoft.Extensions.AI.OpenAI
using Azure;
using Azure.AI.OpenAI;
using Microsoft.Extensions.AI;
builder.Services.AddChatClient(sp =>
new AzureOpenAIClient(
new Uri(builder.Configuration["Foundry:Endpoint"]!),
new AzureKeyCredential(builder.Configuration["Foundry:Key"]!))
.GetChatClient(builder.Configuration["Foundry:Deployment"]!)
.AsIChatClient())
.UseLogging()
.UseOpenTelemetry();
The local alternative is the reason I now keep this boundary even in prototypes. Ollama gives you an IChatClient with no API key and no bill:
ollama pull llama3.1
ollama serve
// dotnet add package OllamaSharp
using OllamaSharp;
builder.Services.AddChatClient(
new OllamaApiClient(new Uri("http://localhost:11434"), "llama3.1"));
Nothing above the registration changes. Same for embeddings: register an IEmbeddingGenerator> and your retrieval code never learns which vendor produced the vectors. My one warning is about embeddings specifically. Swapping the embedding model means re-embedding your corpus, so the abstraction saves you code but not migration work. Do not let the interface convince you the swap is free.
2. On-device AI in MAUI, without marrying one model
The most interesting application of that boundary is on-device AI. One of the sources walks through building it in .NET MAUI, and the question it opens with is the right one: how do you add AI to a cross-platform app without making the rest of the app depend on one model, one operating system, or one experimental API?
The platform reality is messy. Apple Intelligence on iOS and macOS, Gemini Nano through ML Kit on Android, and Phi Silica on Copilot+ Windows PCs are three different runtimes with three different availability stories. The source makes a point I think is underrated: device and runtime differences make a single static compatibility assumption unsafe. Whether a model is present is a runtime question, not a compile-time one.
So the shape is one IChatClient per platform, registered behind conditional compilation, plus a policy client on top that decides what happens when the local model is unavailable.
public enum CloudFallback { Never, AskUser, AllowWithNotice }
public sealed class OnDeviceFirstChatClient(
IChatClient onDevice,
IChatClient cloud,
CloudFallback policy,
IFallbackNotifier notifier) : IChatClient
{
public async Task<ChatResponse> GetResponseAsync(
IEnumerable<ChatMessage> messages,
ChatOptions? options = null,
CancellationToken ct = default)
{
try
{
return await onDevice.GetResponseAsync(messages, options, ct);
}
catch (Exception ex) when (ex is NotSupportedException or PlatformNotSupportedException)
{
switch (policy)
{
case CloudFallback.Never:
throw new AiUnavailableException("On-device model unavailable.", ex);
case CloudFallback.AskUser:
if (!await notifier.RequestConsentAsync(ct))
throw new AiUnavailableException("User declined cloud fallback.", ex);
break;
case CloudFallback.AllowWithNotice:
notifier.ShowBanner("Using cloud AI because on-device AI is unavailable.");
break;
}
return await cloud.GetResponseAsync(messages, options, ct);
}
}
// GetStreamingResponseAsync follows the same shape.
public IAsyncEnumerable<ChatResponseUpdate> GetStreamingResponseAsync(
IEnumerable<ChatMessage> messages, ChatOptions? options = null, CancellationToken ct = default)
=> throw new NotImplementedException();
public object? GetService(Type serviceType, object? key = null) => onDevice.GetService(serviceType, key);
public void Dispose() { onDevice.Dispose(); cloud.Dispose(); }
}
The part I care about is the word the source uses: non-silent. Fallback should be a product decision, not a catch block someone wrote on a Friday. If a user picked your app because "your notes never leave the device," a quiet cloud call is a breach of the promise even if the code is technically graceful. Make the policy an explicit enum, surface it in the UI, and log every fallback so you can see how often your local path actually fails.
My doubt: these platform APIs are young, several are still experimental, and behavior differs across OS versions and hardware. I would treat the adapters as the volatile part and keep them small enough to delete. The boundary is what makes that cheap.
3. The package supply chain got sharper teeth
This is the section that would have saved me the most pain if I had read it a year ago, because two separate problems are now colliding.
Problem one: free packages stopped being free
A run of very popular .NET libraries changed their terms. The codewithmukesh article lists the picture as of late August 2026, and I would treat it as the starting point for your own audit rather than legal advice:
PACKAGE LAST FREE LINE WHAT CHANGED
--------------------------------------------------------------------------
AutoMapper 14.0.0 (MIT) RPL-1.5 or commercial license
MediatR 12.5.0 (Apache-2.0) RPL-1.5 or commercial license
MassTransit v8 line (Apache-2.0) v9 is commercial
FluentAssertions 7.x (Apache-2.0) v8 under Xceed license, free
only for non-commercial use
Polly Still BSD-3 Open Source Maintenance Fee
(a fee model, not a relicense)
AutoMapper and MediatR have a free community tier for smaller organizations, and the article puts the revenue threshold at 5 million dollars in annual gross revenue. Check the current terms yourself, especially for the Polly fee start date, which that article dates to 16 November 2026. Fee structures move, and I would rather you read the maintainers' own pages than trust my summary of someone else's summary.
The reason this matters more with AI in the loop: coding agents were trained on years of code where AddAutoMapper and IMediator were the default answer. They will keep reaching for them, and they will happily bump a version to the latest, which is exactly the version that may carry the new license. A human might notice the license notice in the release notes. An agent will not.
If you decide to move off, there are alternatives worth evaluating, such as Mapperly (source-generated mapping), a source-generator mediator library or plain handler classes, and a community fork of FluentAssertions or Shouldly. I am not endorsing any of them blindly. I am saying that "pin the last free version forever" is a debt strategy, not a plan.
Problem two: models invent packages
The USENIX Security 2025 paper "We Have a Package for You!" tested 16 code-generating LLMs across 576,000 code samples and found 205,474 unique hallucinated package names. On average, 19.7 percent of package references pointed at packages that do not exist. The gap between model types is the part to remember: at least 5.2 percent for commercial models versus 21.7 percent for open-weight ones.
Two caveats so I do not oversell it. The study looked at two ecosystems, and C# and NuGet were not among them, so read it as directional evidence for .NET, not a measured NuGet rate. And a hallucinated name is only dangerous when someone registers it and waits, which is the "slopsquatting" attack the paper describes. That is not hypothetical enough to ignore, and it is cheap to defend against.
Guardrails that cost almost nothing
The codewithmukesh piece has a practical set, and I agree with most of it. First, make vulnerable packages break the build instead of scrolling past in a warning list:
<!-- Directory.Build.props -->
<Project>
<PropertyGroup>
<NuGetAudit>true</NuGetAudit>
<NuGetAuditMode>all</NuGetAuditMode>
<NuGetAuditLevel>low</NuGetAuditLevel>
<WarningsAsErrors>$(WarningsAsErrors);NU1903;NU1904</WarningsAsErrors>
</PropertyGroup>
</Project>
NU1903 is a high-severity advisory and NU1904 is critical. Second, turn on Central Package Management so versions live in exactly one file and an agent cannot sneak a version into a project file:
<!-- Directory.Packages.props -->
<Project>
<PropertyGroup>
<ManagePackageVersionsCentrally>true</ManagePackageVersionsCentrally>
<CentralPackageTransitivePinningEnabled>true</CentralPackageTransitivePinningEnabled>
</PropertyGroup>
<ItemGroup>
<PackageVersion Include="Microsoft.Extensions.AI" Version="9.9.0" />
</ItemGroup>
</Project>
With CPM on, NU1008 fires when a PackageReference tries to carry its own Version, and NU1109 shows up when a transitive pin conflicts with a centrally defined version. Those two warnings are the tripwires that catch an agent editing the wrong file. Third, Package Source Mapping, because NU1507 warns when you have several feeds under CPM and no mapping:
<!-- nuget.config -->
<configuration>
<packageSources>
<clear />
<add key="nuget.org" value="https://api.nuget.org/v3/index.json" />
<add key="internal" value="https://pkgs.example.com/nuget/v3/index.json" />
</packageSources>
<packageSourceMapping>
<packageSource key="nuget.org">
<package pattern="*" />
</packageSource>
<packageSource key="internal">
<package pattern="Acme.*" />
</packageSource>
</packageSourceMapping>
</configuration>
Mapping matters for the hallucination problem specifically: it stops a public package named like your internal one from being resolved by accident.
Fourth, give the agent a way to check reality instead of guessing. Microsoft ships a NuGet MCP server that lets an agent look up whether a package exists and what its latest version is. With the .NET 10 SDK you can run it through dnx:
{
"servers": {
"nuget": {
"command": "dnx",
"args": ["NuGet.Mcp.Server", "--source", "https://api.nuget.org/v3/index.json", "--yes"]
}
}
}
I have written before about when MCP is overkill, and this is one of the cases where I think it earns its keep: it is read-only, it fronts a public registry, and it replaces a guess with a lookup. Pair it with a written rule in your agent instructions file listing banned packages, and a CI step that greps Directory.Packages.props for them.
4. Architecture fitness tests for AI-generated pull requests
Here is the failure mode that worries me most, because it is invisible in the usual dashboards. An agent opens a pull request. Every test is green. Coverage went up. And the change injects a DbContext straight into a controller because that was the shortest path to making the feature work. Nothing is broken. Everything is slightly worse.
Saurav Kumar's write-up on architecture gates for Copilot-generated PRs argues for treating architecture rules as executable tests, and I think that is right. Reviewers get tired, and the volume of AI-generated diffs is not going down. Rules that live in a wiki do not survive contact with a fast agent. Rules that fail the build do.
NetArchTest lets you express those rules in plain xUnit:
// dotnet add package NetArchTest.Rules
using NetArchTest.Rules;
public class ArchitectureTests
{
private static readonly Assembly Domain = typeof(Acme.Domain.Order).Assembly;
private static readonly Assembly Web = typeof(Acme.Web.Program).Assembly;
[Fact]
public void Domain_does_not_depend_on_infrastructure()
{
var result = Types.InAssembly(Domain)
.ShouldNot()
.HaveDependencyOnAny("Acme.Infrastructure", "Microsoft.EntityFrameworkCore")
.GetResult();
Assert.True(result.IsSuccessful, Describe(result));
}
[Fact]
public void Controllers_do_not_touch_data_access_directly()
{
var result = Types.InAssembly(Web)
.That().HaveNameEndingWith("Controller")
.ShouldNot().HaveDependencyOn("Microsoft.EntityFrameworkCore")
.GetResult();
Assert.True(result.IsSuccessful, Describe(result));
}
private static string Describe(TestResult r) =>
"Violations: " + string.Join(", ", r.FailingTypeNames ?? []);
}
The legacy problem, and the ratchet
This is where most teams give up. You add the rule to a five-year-old codebase and it reports 214 violations on day one, so someone disables the test and the whole idea dies. The fix I find most convincing is a ratchet: record today's violations as a baseline, fail the build on any new one, and also fail when a baseline entry disappears, so improvements get locked in.
private static void AssertRatchet(string rule, TestResult result)
{
var path = Path.Combine(AppContext.BaseDirectory, "arch-baseline", $"{rule}.txt");
var baseline = File.Exists(path)
? File.ReadAllLines(path).Where(l => l.Length > 0).ToHashSet()
: new HashSet<string>();
var current = (result.FailingTypeNames ?? []).ToHashSet();
var added = current.Except(baseline).ToList();
var healed = baseline.Except(current).ToList();
Assert.True(added.Count == 0,
$"New violations of '{rule}': {string.Join(", ", added)}");
Assert.True(healed.Count == 0,
$"Fixed. Remove from baseline so it cannot regress: {string.Join(", ", healed)}");
}
Make sure the baseline files are copied to the test output directory in the csproj. The count only ever goes down, and the AI agent gets an unambiguous error message it can act on, which is a nice side effect. An agent that sees "New violations of 'controllers-no-ef': OrdersController" will usually fix it on the next attempt.
One caution: the sources include benchmark-style numbers for how much these gates catch, and some of them are explicitly illustrative in their own text. I am not repeating any of them. Measure your own false-positive rate before you make the gate mandatory.
5. A production text-to-SQL pipeline that cannot write
The last pattern is the one with the sharpest edges. Satwiki De’s write-up on text-to-SQL agentic AI for enterprise analytical tasks describes a pipeline built on Azure AI Foundry with DeepEval-based evaluation, and the shape is one I would defend: not one big prompt, but five stages with a hard gate in the middle.
question
|
v
[Intent detector] analytical question? or chit-chat / write request?
|
v
[Query generator] LLM proposes SQL using an allowlisted schema
|
v
[Query validator] deterministic code, NOT an LLM. Rejects writes
| and multi-statement SQL.
v
[Read-only executor] read-only login, timeout, row cap
|
v
[Report builder] LLM summarizes rows into a report
The design decision I like best is that the validator is ordinary code. Asking a second model whether the first model’s SQL is safe is asking a probabilistic system to guard a deterministic risk. Parse the SQL and check the tree.
// dotnet add package Microsoft.SqlServer.TransactSql.ScriptDom
using Microsoft.SqlServer.TransactSql.ScriptDom;
public sealed record Verdict(bool IsValid, string Reason)
{
public static Verdict Ok() => new(true, "");
public static Verdict Reject(string why) => new(false, why);
}
public sealed class ReadOnlySqlValidator(IReadOnlySet<string> allowedTables)
{
public Verdict Validate(string sql)
{
var parser = new TSql160Parser(initialQuotedIdentifiers: true);
using var reader = new StringReader(sql);
var fragment = parser.Parse(reader, out var errors);
if (errors.Count > 0)
return Verdict.Reject($"Does not parse: {errors[0].Message}");
var statements = ((TSqlScript)fragment).Batches
.SelectMany(b => b.Statements).ToList();
if (statements.Count != 1)
return Verdict.Reject("Exactly one statement is allowed.");
if (statements[0] is not SelectStatement select)
return Verdict.Reject("Only SELECT statements are allowed.");
if (select.Into is not null)
return Verdict.Reject("SELECT INTO writes data and is not allowed.");
var cteNames = select.WithCtesAndXmlNamespaces?.CommonTableExpressions
.Select(c => c.ExpressionName.Value)
.ToHashSet(StringComparer.OrdinalIgnoreCase)
?? new HashSet<string>(StringComparer.OrdinalIgnoreCase);
var visitor = new TableVisitor();
fragment.Accept(visitor);
foreach (var table in visitor.Tables)
if (!allowedTables.Contains(table) && !cteNames.Contains(table))
return Verdict.Reject($"Table '{table}' is not on the allowlist.");
return Verdict.Ok();
}
private sealed class TableVisitor : TSqlFragmentVisitor
{
public List<string> Tables { get; } = new();
public override void Visit(NamedTableReference node) =>
Tables.Add(node.SchemaObject.BaseIdentifier.Value);
}
}
Two things to note. A SELECT INTO is still a SelectStatement in the parse tree, so the explicit Into check is not decoration. And parsing is a first line of defense, not the only one. The executor is the second, and it should not depend on the validator being perfect:
public async Task<IReadOnlyList<Dictionary<string, object?>>> ExecuteAsync(string sql, CancellationToken ct)
{
// Connection string uses a login with db_datareader only, plus:
// Application Intent=ReadOnly
await using var conn = new SqlConnection(_readOnlyConnectionString);
await conn.OpenAsync(ct);
await using var cmd = new SqlCommand(sql, conn) { CommandTimeout = 15 };
await using var reader = await cmd.ExecuteReaderAsync(ct);
var rows = new List<Dictionary<string, object?>>();
while (await reader.ReadAsync(ct) && rows.Count < 1000)
{
var row = new Dictionary<string, object?>();
for (var i = 0; i < reader.FieldCount; i++)
row[reader.GetName(i)] = reader.IsDBNull(i) ? null : reader.GetValue(i);
rows.Add(row);
}
return rows;
}
If the validator has a bug, the database login still cannot write. That layering is the point. For local development, a SQL Server container gives you a real target without touching anything shared:
docker run -e "ACCEPT_EULA=Y" -e "MSSQL_SA_PASSWORD=Dev_Passw0rd!" \
-p 1433:1433 -d mcr.microsoft.com/mssql/server:2022-latest
The orchestrator is short, and the retry loop feeds the validator's rejection reason back to the generator, which in my experience of this general pattern fixes most first-attempt mistakes:
public async Task<Report> RunAsync(string question, CancellationToken ct)
{
var intent = await _intent.DetectAsync(question, ct);
if (intent.Kind != IntentKind.Analytical)
return Report.Refusal(intent.Reason);
string? feedback = null;
for (var attempt = 0; attempt < 3; attempt++)
{
var sql = await _generator.GenerateAsync(question, intent, feedback, ct);
var verdict = _validator.Validate(sql);
if (!verdict.IsValid) { feedback = verdict.Reason; continue; }
var rows = await _executor.ExecuteAsync(sql, ct);
return await _reports.BuildAsync(question, sql, rows, ct);
}
return Report.Refusal("Could not produce a safe query for that question.");
}
Evaluating it with DeepEval
DeepEval is a Python library, so the pragmatic setup is a small pytest project that calls your .NET service over HTTP in CI. I would split the tests by what they can be checked with. Anything deterministic gets a deterministic assertion: "delete all orders" must be refused, a SELECT followed by a second statement that drops a table must be rejected, and the generated SQL should return the same rows as a hand-written gold query. Use an LLM judge only for the fuzzy part, the report narrative.
# pip install deepeval httpx pytest
import httpx
from deepeval import assert_test
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, LLMTestCaseParams
faithful = GEval(
name="Report matches the data",
criteria="The report must only state facts supported by the returned rows. "
"Penalize invented numbers or trends.",
evaluation_params=[LLMTestCaseParams.INPUT,
LLMTestCaseParams.ACTUAL_OUTPUT,
LLMTestCaseParams.RETRIEVAL_CONTEXT],
threshold=0.7,
)
def test_revenue_report_is_faithful():
r = httpx.post("http://localhost:5000/ask",
json={"question": "Revenue by region last quarter"}).json()
case = LLMTestCase(
input="Revenue by region last quarter",
actual_output=r["report"],
retrieval_context=[str(r["rows"])],
)
assert_test(case, [faithful])
Point the judge at your Foundry deployment with deepeval set-azure-openai, or at a local model with deepeval set-ollama llama3.1 while you iterate. Be aware that a small local judge is a noisy judge, so use it to debug the harness and a stronger model for the numbers you actually trust.
Start here: a checklist for your first production agent
If I were handing a .NET team one page, it would be this:
- Put IChatClient and IEmbeddingGenerator at your app boundary, and keep provider SDKs out of business code.
- Register an Ollama-backed client for local development so nobody needs a key to run the tests.
- Decide your fallback policy in writing (never, ask, or allow with notice) before writing any fallback code.
- Turn on NuGet audit and make NU1903 and NU1904 fail the build.
- Enable Central Package Management and Package Source Mapping, and treat NU1008, NU1109, and NU1507 as real errors.
- Audit your dependency list for AutoMapper, MediatR, MassTransit, FluentAssertions, and Polly against their current license terms.
- Add a banned-packages rule to your agent instructions and a CI grep that backs it up.
- Write five NetArchTest rules for your worst architecture violations, and ratchet them if the codebase is old.
- Give the agent a read-only tool for your data, and enforce read-only at the database login, not just in the prompt.
- Split evals into deterministic checks first and LLM-judge checks second, and run them in CI.
None of this is glamorous, and that is exactly why I trust it. The teams that worry me are not the ones with a weaker model. They are the ones with a great model and no guardrails around what it is allowed to touch.
Sources I read for this survey:
- “The Hardest Part of Building AI in .NET Is Not the AI,” AI in Plain English
- “On-Device AI in .NET MAUI Without Model Lock-In,” blog.devops.dev
- codewithmukesh.com, “Stop Your AI Agent Picking the Wrong .NET Libraries”
- Saurav Kumar, “Building Architecture Gates for Copilot-Generated Pull Requests in .NET,” c-sharpcorner.com
- Satwiki De, “Building Text-to-SQL Agentic AI for enterprise analytical tasks,” Medium
- Spracklen et al., “We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs,” USENIX Security 2025
Tags: dotnet, csharp, ai, softwarearchitecture, nuget, azure, maui, devops
Top comments (0)