Skip to content

IKRC Insights

Why a Task-Specific Agent Beats a Chatbot Bolted Onto Your App

A general chat box leaves the person using it to work out what to ask and which system holds the answer. An agent built for one named workflow starts with that workflow's context and a fixed set of tools.

A chat box added to an operations screen can be asked anything. The person using it still has to work out what to ask, which of the systems behind that screen holds the answer, and whether the reply is safe to act on. The application can supply the workflow context and enforce action checks.

A task-specific agent starts from a workflow the application already runs: a fixed sequence of reads, one or two writes, and a record that shows the work was done. The agent is given the operations that workflow needs and nothing else, and it is finished when the application can see that record.

The Tool List Is the Scope

Scope written into a system prompt reads like a rule, and the model treats it as guidance. What the agent can actually reach is the Tools collection on the chat options it was handed. The FunctionInvokingChatClient reference describes the loop: when a response from the inner client contains a function call, the client invokes the matching AIFunction from Tools, sends the result back, and repeats until there are no more function calls or another stop condition is met. The prompt plays no part in which functions exist.

That makes scoping a list in the code that builds the client. Write down the workflow, then the smallest set of operations that completes it, and register exactly that set. The name and description attached to each tool are what the model uses to choose between them; the function calling quickstart passes both when the function is created so the model can tell what it is for.

Overlapping descriptions make that choice ambiguous. Two tools described as "get order information" and "fetch order details" give the model nothing to separate them, and if one of them also writes, the description should say so plainly.

What Does Support See When a Tool Call Fails?

Suppose a tool that reads an invoice throws because the account number it was given does not exist. It is easy to assume the exception ends the interaction. By default it does not: when a function invocation fails, FunctionInvokingChatClient keeps making requests to the inner client, optionally including the exception information, so the model can try other arguments. The person watching the screen sees a pause while that happens.

Three published defaults in Microsoft.Extensions.AI 10.9.0 govern the next step. MaximumConsecutiveErrorsPerRequest is 3; after that many consecutive failing iterations the exception is rethrown to the caller, and setting it to zero rethrows immediately. MaximumIterationsPerRequest is 40 and includes the initial request, so it is the whole round-trip budget for one question. An iteration can contain more than one tool call, so neither number tells you exactly how many lookups ran. IncludeDetailedErrors is false, which means the model receives a generic error message and the raw exception stays available to application code on the function result.

Turning detailed errors on sends exception text to the model, and the reference warns that raw exception information may be disclosed to external users. Staff authentication in front of an internal agent does not make that text safe either: exception messages can carry connection details, internal names and data the person asking was never shown. Leave the default in place unless a specific tool's errors have been reviewed.

The useful record goes to access-controlled application logs. Capture the tool name, a correlation identifier and the arguments the task actually needs, redacted where they hold customer data, before the loop moves on. Set both limits explicitly in code, then run the failure path in a test, so the first time anyone sees the retry behavior is not in production.

The Identity the Tool Runs Under

A tool is ordinary application code, and what it runs as can end up settled by whoever writes the first one. AIFunctionArguments carries the IServiceProvider that FunctionInvokingChatClient was given, so a client built through standard dependency injection passes that provider into the function, where it can resolve anything in the container. Resolving a service from that provider says nothing about who is asking. The caller's identity has to be passed in deliberately, for example through ChatOptions.AdditionalProperties, which a function can read from FunctionInvokingChatClient.CurrentContext, and then enforced by the service the tool calls.

Consider two versions of the same invoice tool. One calls the invoice service the screens already use, passes the current user, and returns what that user is entitled to see. The other opens its own connection with an integration account and returns whatever that account can read. They look identical in a demo. The second one can describe accounts the user was never allowed to open, and because no permission check ran, no failed check exists to find later.

Having a tool also does not authorize the business action behind it. If the workflow ends in a write, the service that performs it runs the same permission check it runs for a person at a screen, with the actual user and the arguments it received, and the record names that user as well as the agent. The user's permission to make that kind of write is not approval of this one. Where the write needs a person's approval, such as paying an invoice, the agent can prepare it, and the application commits it only after an authorized person approves those exact details on the application's own screen. When the data underneath is stale or split across copies, the tool inherits that problem too; fixing the retrieval layer comes first.

Definition of Done

Give the agent a finish line the application can check: a row that exists, a status that changed, a document that was filed. A fluent answer in the transcript is not evidence that any of those happened.

Where workflow state changes are recorded, count completions and subsequent reopenings from that history. If those events are missing, add the instrumentation before comparing agent versions. Capture the baseline before changing prompts or tool descriptions.

Related Reading

Before You Add AI, Fix the Data Retrieval Layer covers the layer underneath the tools. Customer-facing MCP servers: permissions, approvals and retries covers exposing a system to outside agents on purpose.

Contact IKRC

Your next software project.

Connection Lost

Attempting to reconnect to the server...