Skip to content

IKRC Insights

Adding AI to the .NET Application You Already Run

The model client arrives as one more injected dependency inside the caching, tracing and failure handling the application already has. Its place in that pipeline, and the data it is given, are the decisions to make on purpose.

Take an illustrative internal quoting screen. It reads a handful of tables, applies a discount rule nobody wants to touch, and prints a PDF. Somebody now wants it to draft the covering note that goes out with the quote, from inside the same application.

A registration connects the model client: AddChatClient goes into the same startup file that already registers the repository and the mail sender, and the client resolves into the same service the screen calls. The rest of the change follows from what that client is. It is a network call that is billed by the token, can be slow or unavailable, and can produce wrong text, so the screen needs to behave sensibly when the note is missing, and someone has to decide which customer data the prompt may contain.

The change lands in startup

Register the client the way the documentation does, with a builder, and read that builder as the list of things that happen around every model call. Microsoft's IChatClient documentation states the shape directly: IChatClient instances can be layered to create a pipeline of components that each add functionality, and those components can come from Microsoft.Extensions.AI, other NuGet packages or your own code. AddChatClient followed by UseDistributedCache and UseOpenTelemetry is the integration surface for caching and tracing, and the cache in question is whatever IDistributedCache the application already registered. Registration also puts the client under checks the host already runs: in the Development environment, the default service provider verifies that scoped services are not resolved from the root provider and not injected into singletons, so a per-request dependency captured by a long-lived client fails on a developer machine.

These are delegating clients. Each wraps the next, so registration order is nesting order: whatever is registered first sees the request first and the response last. The pipeline sample in that documentation carries a comment inviting you to explore changing the order of the intermediate Use calls. In the sample as printed, the cache comes first, function invocation second and telemetry third.

What does a cache hit skip?

UseDistributedCache layers a DistributedCachingChatClient around the client below it. When a novel chat history is submitted, it forwards the request, caches the response and returns it. When the same history is submitted again and a cached response is found, it returns that response without forwarding the request along the pipeline. On a hit, nothing registered after the cache runs.

Put the cache outside function invocation, as in the printed sample, and a cached answer comes back without the tool underneath it running. For a question about what a policy document says, that is the saving you wanted. For a question about an order balance, the answer was generated earlier and the call that would have read the current balance was skipped. Nothing errors and nothing retries. With telemetry registered after the cache, as in the same sample, the hit never reaches that inner telemetry client either, so it records no model-call span for the request; instrumentation outside the cache, such as the incoming web request's own span or custom logging, can still record that the request happened.

The key is broader than the last thing a user typed. The GetCacheKey reference says the messages, the ChatOptions and any additional values are serialized to JSON to compute it, so a change to options such as the tool list produces a different key. Where a conversation is stateless and the application resends the growing history on every turn, each turn is a new key; repeated identical single-shot prompts, such as classification or extraction, are where hits can occur. By default, caching is skipped when the options carry a ConversationId, and the reference warns that the generated key is not guaranteed to be stable across library releases, so a package upgrade can leave existing cache entries unused.

Nothing in that key comes from the ambient request. The tenant or user the application resolved from the HTTP context or a scoped service is not part of it unless the application adds it, so two callers who send the same messages with the same options can receive the same cached response. For a policy question everyone may read, that is fine. Where the answer depends on what the caller is allowed to see, a cached answer must not cross that access boundary: scope the key to the tenant or user, or do not cache. A cached response also skips everything inside the cache, so the only permission check for a request cannot live in a tool or service the cache sits in front of.

Choose the order deliberately and write the reason next to the registration. Cache outside function invocation only for prompts that do not read live data. For a prompt that reads a current order balance, the clearer default is to leave response caching off for that call. Putting the cache inside function invocation keeps that invocation layer active, but a tool runs only when the inner response requests it. That placement alone does not guarantee a fresh balance lookup. Test the returned tool calls and cache keys against the freshness requirement. If the traces have to account for every model call the application made, register telemetry outside the cache.

The boundary between the client and the business rule

The injected client belongs behind the service that already owns the discount rule. Hand the model the computed quote and ask for prose about it. Handing it the raw rows and an instruction to work the discount out asks the model to repeat a calculation the application already does correctly.

The shortcut is easy to picture. Someone injects the chat client into the controller and passes the rows straight through. A note can then describe a discount the rule did not apply and go out under the company name, with no validation failure in the logs because nothing was validated.

Keep the model call downstream of the rule; a task-specific agent is this same boundary with a tool list attached. The client reads the output of the calculation, produces text and hands it back. When the AI service is unavailable, the quote still prints without the covering note, because the note is an addition to a workflow that already worked.

What reaches the trace

UseOpenTelemetry gives the model call the same tracing the rest of the application has, written against the OpenTelemetry semantic conventions for generative AI. The next question is what ends up inside the span, because a prompt assembled from a customer record contains the customer record.

The default is narrow. EnableSensitiveData defaults to false, and the reference says what that covers: telemetry includes metadata such as token counts and excludes raw inputs and outputs, meaning message content, function call arguments and function call results. That default can be changed from outside the code. It becomes true when the OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT environment variable is set to "true", case-insensitive, and environment variables are set by whoever configures the host. Setting the property explicitly in code overrides the variable, which keeps the decision in the codebase when somebody else configures the deployment.

Related Reading

For the data layer underneath a model call, read Before You Add AI, Fix the Data Retrieval Layer.

Contact IKRC

Your next software project.

Connection Lost

Attempting to reconnect to the server...