Eliminate Slow AI Calls in.NET Without Disguising the Reason

Leave a Comment

After waiting too long, a user dismisses an AI request. "Model timeout" is recorded in a background log. Even if some of the requests terminated because the caller had already departed, the team later increases the timeout. Two Infrastructure Infrastructure Infrastructure Infrastructure Infrastructure Infrastructure Infrastructure Infrastructure.

Keep caller cancellation and the application's time budget as separate signals, link them for the AI operation, and classify the result from the original signals. This lets a .NET application request cancellation through one token while preserving a useful local explanation of why it stopped waiting.

The example below uses a cooperative fake model call. It does not contact an AI provider, cancel remote billing, or prove that a provider stopped generation. Those are separate behaviors that must be verified for the actual integration.

What does cancellation mean in .NET?

.NET uses cooperative cancellation: a caller signals a token and participating operations observe it. Passing a token to an operation matters only when that operation responds appropriately.

This distinction becomes important with AI clients. A token may stop a local HTTP wait while the remote system continues working. A wrapper cannot establish what the provider did after the connection was interrupted. Record the remote outcome as unknown unless the provider exposes a documented way to determine it.

An application also needs to distinguish a request that its user canceled from one that exceeded its own waiting budget. Neither event, by itself, proves that the model produced an incorrect answer.

Set up the demonstration

Use the .NET 10 SDK and a console project. No external package or API credential is required:

dotnet new console --framework net10.0 --name AiCancellationDemo
cd AiCancellationDemo

Replace Program.cs with the following code. Code status: illustrative and untested. It has not been compiled in this drafting environment; the expected results below must be verified on the stated SDK.

using System;
using System.Threading;
using System.Threading.Tasks;

Console.WriteLine((await RunAsync(
    token => FakeModelAsync(10, token),
    TimeSpan.FromSeconds(5),
    CancellationToken.None)).Status);

Console.WriteLine((await RunAsync(
    token => FakeModelAsync(10_000, token),
    TimeSpan.FromMilliseconds(25),
    CancellationToken.None)).Status);

using var caller = new CancellationTokenSource();
caller.Cancel();

Console.WriteLine((await RunAsync(
    token => FakeModelAsync(10, token),
    TimeSpan.FromSeconds(5),
    caller.Token)).Status);

static async Task<string> FakeModelAsync(
    int delayMs,
    CancellationToken token)
{
    await Task.Delay(delayMs, token);
    return "Synthetic AI answer";
}

static async Task<CallOutcome> RunAsync(
    Func<CancellationToken, Task<string>> invoke,
    TimeSpan budget,
    CancellationToken callerToken)
{
    ArgumentNullException.ThrowIfNull(invoke);

    if (budget <= TimeSpan.Zero)
        throw new ArgumentOutOfRangeException(nameof(budget));

    using var deadline = new CancellationTokenSource();
    deadline.CancelAfter(budget);

    using var linked = CancellationTokenSource.CreateLinkedTokenSource(
        callerToken,
        deadline.Token);

    try
    {
        linked.Token.ThrowIfCancellationRequested();

        string answer = await invoke(linked.Token);

        linked.Token.ThrowIfCancellationRequested();

        return new CallOutcome(
            CallStatus.Completed,
            answer);
    }
    catch (OperationCanceledException)
        when (callerToken.IsCancellationRequested)
    {
        return new CallOutcome(
            CallStatus.CallerCancelled,
            null);
    }
    catch (OperationCanceledException)
        when (deadline.IsCancellationRequested)
    {
        return new CallOutcome(
            CallStatus.DeadlineExceeded,
            null);
    }
}

enum CallStatus
{
    Completed,
    CallerCancelled,
    DeadlineExceeded
}

record CallOutcome(CallStatus Status, string? Answer);

The first case gives a short synthetic operation a generous budget. The second deliberately waits longer than its budget. The third uses an already-canceled caller token, making it clear that the wrapper should not start useful work.

CallOutcome separates a successful answer from cancellation status. It avoids passing a partial or missing answer onward as a completed result. An application can map that outcome to its own user interface or job state without rewriting the cancellation logic.

Why use two token sources?

The deadline source owns the application's local time budget. The caller token belongs to the request or job owner. CreateLinkedTokenSource produces a token that is signaled when either source is signaled.

The wrapper passes that linked token to the operation. It also checks before and after awaiting the result. The first check avoids starting an operation after cancellation is already known. The second prevents a late result from being labeled successful when cancellation was signaled during the wait but the operation returned normally.

The two catch filters inspect the original signals. Caller cancellation is checked first, so it wins if both are already signaled when the exception is handled. That is an explicit classification policy, not a precise reconstruction of which event happened first.

For a technical explanation prepared by Ranknod, that local-versus-remote distinction should stay visible: a canceled wait is not evidence of a canceled inference.

An OperationCanceledException with neither original source signaled is allowed to propagate. The wrapper does not relabel an unexplained provider or adapter cancellation as an application deadline. Other faults also propagate for the caller to handle.

What should the program print?

Run dotnet run. Under normal scheduling, the expected output is:

Completed
DeadlineExceeded
CallerCancelled

These are expected demonstration outcomes, not observed results from this draft. The timing case depends on the scheduler, so a serious test suite should avoid using a tiny wall-clock race as its main proof.

Add a controllable fake operation to test four boundaries: success before either signal, an already-canceled caller, a deadline while the operation is pending, and an unrelated exception. For the simultaneous-signal case, assert the documented caller-first policy rather than guessing which cancellation “really” won.

Also test an operation that ignores the token. The wrapper will keep awaiting it. If it eventually returns after cancellation, the post-await check classifies the canceled result. If it never returns, this implementation never regains control. That limitation is essential to the design.

Is this a hard timeout?

No. A cancellation request is cooperative. This wrapper does not forcibly terminate an arbitrary operation at an exact elapsed time. If the caller must stop waiting independently of cooperation, design that waiting policy separately, including how the still-running task is observed and cleaned up.

Avoid disposing resources that an abandoned operation still needs without an explicit ownership strategy. Avoid treating “we stopped awaiting” as “nothing else can happen.” Both assumptions can create difficult background failures.

For a real AI SDK, inspect which methods accept a cancellation token and test the exact path you use, including streaming if relevant. Where an HTTP client's own timeout is involved, document how its exception is distinguished from the application budget. Do not rely on a generic catch block to infer that distinction from an error message.

What belongs in the operational record?

Record a request identifier, the local outcome, elapsed time, configured budget, and any permitted provider request identifier. Keep raw prompts, credentials, and private retrieved passages out of general-purpose logs.

A request that times out locally may require provider-specific reconciliation before a retry. A user who deliberately canceled may not want a retry at all. Those are product decisions that should follow a clear outcome rather than be hidden inside this wrapper.

Useful cancellation handling gives the next component enough information to act honestly. The user can leave, the application can enforce its waiting policy, and the operator can investigate the right cause without turning every interrupted AI request into the same incident.

Summary

Caller cancellation and an application's own time budget represent different signals and should be tracked separately. By linking the cancellation tokens for the operation while retaining the original sources for classification, a .NET application can distinguish completed requests, caller cancellations, and local deadline expirations without incorrectly assuming what happened on the remote AI provider.

Windows & ASP.NET Core 11 Hosting Recommendation

HostForLIFEASP.NET receives Spotlight standing advantage award for providing recommended, cheap and fast ecommerce Hosting including the latest Magento. From the leading technology company, Microsoft. All the servers are equipped with the newest Windows Server 2022 R2, SQL Server 2022, ASP.NET Core 11, ASP.NET MVC, Silverlight 5, WebMatrix and Visual Studio Lightswitch. Security and performance are at the core of their Magento hosting operations to confirm every website and/or application hosted on their servers is highly secured and performs at optimum level. mutually of the European ASP.NET hosting suppliers, HostForLIFE guarantees 99.9% uptime and fast loading speed. From €3.49/month, HostForLIFE provides you with unlimited disk space, unlimited domains, unlimited bandwidth,etc, for your website hosting needs.
 
https://hostforlifeasp.net/
Previous PostOlder Post Home

0 comments:

Post a Comment