Your Long MCP Tool Call Dies at the Client Timeout

A tool that takes four minutes to run cannot answer inside a client request that gives up after sixty seconds. The old fix was to keep the connection open and pray.

Viral RuparelNo. 65 of 6511 min read

A team I was helping had a perfectly good MCP tool that generated a quarterly report: pull the numbers, run the model over them, render a PDF, upload it, return a link. It worked every time in testing. In production it failed about a third of the time, and always the same way: the agent reported that the tool had errored, but the report showed up in the bucket a minute later anyway. The tool had not errored. It had finished. The client had simply stopped waiting.

The render took between forty seconds and three minutes depending on how much data the quarter held. The client gave the call sixty seconds and then closed the request as a timeout. Everything the tool did after that second sixty was invisible to the agent, which meant the agent either gave up or, worse, called the tool again and kicked off a second render on top of the first.

This is the oldest problem in remote calls wearing new clothes. A tool call is a request and a response, and a request has a deadline. For most tools that is fine. For any tool whose honest answer is "this takes a few minutes," a single request is the wrong shape, and no amount of keep-alive heartbeat fixes the fundamental mismatch. The 2026-07-28 MCP spec finally gives this case a real protocol, the Tasks extension, and it is worth building on correctly rather than bolting a job queue onto the side of your server.

On the left, a client sends tools/call to a server, a sixty-second timeout fires, and the server keeps rendering a report the client never sees; on the right, the server answers tools/call immediately with a task handle, writes task state to a shared store, and the client polls tasks/get at the given interval until the task reaches completed and carries the result.
Figure. A blocking tool call that times out beside a task handle the client polls

The failure: a tool call outlives the request that made it

Walk through what actually happened to that report tool. The client opens a request, sends tools/call for generate_report, and starts a timer. The server begins the render. At sixty seconds the client's timer fires, the client tears down the request, and from its side the call is over: either an error surfaces to the model or the turn simply stalls. The server, holding no knowledge that anyone stopped caring, keeps rendering, finishes at ninety seconds, and writes a perfectly good PDF that no one is listening for.

Every workaround for this is bad in a specific way. You can stream a trickle of progress events to keep the connection warm, which some clients allow, but now your tool's correctness depends on a heartbeat staying alive across proxies, load balancers, and a client that may cap total stream duration anyway. You can make the tool return fast with a "started" message and expose a second tool, check_report_status, that the model is supposed to call later, but now you have taught the model a two-step dance it will sometimes forget, and you are hand-rolling state tracking that every such tool reinvents. You can shorten the work, which is not an option when the work is genuinely a three-minute render.

The deeper issue is that the client's timeout is not wrong. It is correct for the common case and protective: a client that waits forever on a hung tool is a client that hangs the whole agent. The mistake is forcing a long job through a surface built for a short one. That same tension shows up whenever an agent touches slow systems, which is why bounding deadlines and propagating cancellation matters even when nothing is broken. Tasks does not remove the deadline. It changes what the deadline applies to.

What the Tasks extension actually changes

Tasks began life as an experimental feature baked into the MCP core, and production use made it clear the design needed room to move, so the 2026-07-28 release lifted it out into a versioned extension identified as io.modelcontextprotocol/tasks and standardized as SEP-2663. The identifier matters because the whole release is built around extensions you opt into rather than one monolithic protocol, which is the same philosophy behind the move that made MCP servers stateless and horizontally scalable.

The core idea is small. A tool that would otherwise block can instead answer the tools/call immediately with a durable task handle, then do the real work in the background. The client, seeing a handle rather than a finished result, follows the task over time: it reads the current state, waits, reads again, and collects the result once the task is done. The spec post describes the poll side around a tasks/get method, with a companion tasks/update for the control surface.

A task carries a small lifecycle. It is working while the server grinds, it can enter input_required if it needs something from the user mid-flight, and it ends in exactly one of three final states: completed, failed, or cancelled. The last three are terminal, which is the property the client's loop keys off: poll until the state is final, then stop. The Upstash write-up lays out this state set clearly and is a good companion read.

Two rules keep this from becoming chaos. First, the client opts in: it declares support for the Tasks extension in its capabilities, and a server must never hand a task to a client that did not ask for one, because that client has no idea what to do with a handle. Second, the server sets the pace. Each task carries a pollIntervalMs that tells the client how often it is welcome to ask, so a server rendering a slow report can say "check me every five seconds" and not get hammered once a second by an eager client.

Returning a task instead of blocking

Here is the shape of a tool handler that decides, per call, whether the work fits in a request or needs a task. The cheap path stays synchronous. The expensive path spawns a task and returns the handle.

// The tool decides synchronously: small job answers inline,
// large job becomes a task the client will follow.
async function handleGenerateReport(args: ReportArgs, caps: ClientCapabilities) {
  const estimate = estimateRenderSeconds(args); // cheap heuristic on the inputs

  if (estimate < 10) {
    // Fits comfortably inside one request. No task needed.
    return { content: [{ type: "text", text: await renderNow(args) }] };
  }

  // Long job. Refuse to start a task the client cannot follow.
  if (!caps.extensions?.["io.modelcontextprotocol/tasks"]) {
    return { isError: true, content: [{ type: "text",
      text: "This report is too large to return in one call. Enable Tasks." }] };
  }

  const task = await tasks.create({ kind: "generate_report", args });
  startRenderInBackground(task.id, args); // fire and forget; see the next section
  return { task: { taskId: task.id, status: "working", pollIntervalMs: 5000 } };
}

Two things in that handler are easy to skip and expensive to skip. The capability check is not optional politeness: if you hand a task to a client that never opted in, you have returned something it will treat as a malformed result, and you are back to a failed tool call with extra steps. And the synchronous fast path is worth keeping, because wrapping a two-second lookup in a task adds a poll round trip and a store write for no reason. Tasks earns its cost on work that is actually long.

Polling without hammering the server

The client side is a loop with a bound. It reads the task, and if the state is not final it waits the interval the server asked for and reads again. The bound is yours to set, because "poll forever" is how a stuck task becomes a stuck agent.

async function awaitTask(taskId: string, deadlineMs: number): Promise<TaskResult> {
  const start = Date.now();
  let interval = 1000; // server will tell us the real cadence on the first read

  while (Date.now() - start < deadlineMs) {
    const task = await client.request("tasks/get", { taskId });
    interval = task.pollIntervalMs ?? interval;

    if (task.status === "completed") return task.result;
    if (task.status === "failed")    throw new TaskFailed(task.error);
    if (task.status === "cancelled") throw new TaskCancelled();
    if (task.status === "input_required") {
      // Mid-flight the task needs a value it could not infer. Collect and send.
      await client.request("tasks/update", { taskId, inputResponses: await ask(task) });
    }
    await sleep(interval);
  }
  // We hit our own ceiling. Cancel so the server stops spending on work nobody waits for.
  await client.request("tasks/cancel", { taskId }).catch(() => {});
  throw new TaskTimeout();
}

Notice that the client's timeout did not vanish, it moved up a level. A single tasks/get is a fast request with a normal short deadline, and the long wait is now a series of those fast reads under a ceiling the client chose deliberately. When that ceiling is hit, the client cancels the task rather than abandoning it, so the server can stop a render nobody will read. This is the same discipline as streaming progress to keep perceived latency low: the human and the model both want to know the thing is alive, and a task that reports working with a sane interval gives them that without a held-open pipe.

The task store is the whole design

The part that separates a real Tasks implementation from a demo is where the task lives. In the handler above, tasks.create and startRenderInBackground imply a store and a worker, and if that store is a Map in one process, you have built the same trap the stateless spec was trying to kill. The client polls tasks/get, a load balancer routes that poll to a different replica, and the replica that gets it has never heard of the task. It returns "unknown task," the client treats the job as lost, and meanwhile the first replica finishes a report into the void.

So the task state has to live somewhere every replica can read and write: a row in Postgres, a record in Redis with a sane TTL, a document in whatever durable store you already run. The task id is the key, the status and pollIntervalMs and eventually the result are the value, and the background worker updates that row as it progresses. Any replica serving a poll just reads the row. This is not a new idea, it is exactly the durable execution contract where you record each finished step so a crash resumes instead of restarting, applied to the task's own lifecycle. Tasks gives you the protocol surface; durable execution keeps the work behind it honest across a deploy.

-- The task store is the contract. Any replica reads and writes this row.
-- status in ('working','input_required','completed','failed','cancelled')
CREATE TABLE mcp_task (
  task_id        text PRIMARY KEY,
  kind           text NOT NULL,
  status         text NOT NULL DEFAULT 'working',
  poll_interval_ms integer NOT NULL DEFAULT 5000,
  result         jsonb,            -- populated only on 'completed'
  error          text,             -- populated only on 'failed'
  lease_owner    text,             -- which worker owns the run right now
  lease_expires  timestamptz,      -- so a dead worker's task can be reclaimed
  updated_at     timestamptz NOT NULL DEFAULT now()
);

The lease columns are the detail people leave out and regret. If two workers can pick up the same task, they will double-render; a lease with an expiry lets exactly one worker own a run while letting another reclaim it if the owner dies. And because the client may retry the original tools/call after a network blip, the underlying work needs idempotency keys so a retry does not kick off a second job. A task handle makes the long case possible; idempotency makes it safe.

Events: waking up instead of being polled

Polling is fine, but it is not free: a task that takes ten minutes at a five-second interval is a hundred and twenty wasted reads if nothing interesting happened. The natural next step is for the server to tell the client when something changes rather than answer the same question over and over. The 2026-07-28 work includes a sketch for exactly this, a Triggers and Events design where the client registers a callback with a signing secret and the server POSTs a signed event when the task changes state, which wakes the agent instead of making it ask.

I am flagging it rather than building on it, because as of this writing that events design is still a sketch in an incubation repo without a settled SEP, and real clients are only starting to support even the stable Tasks surface. Treat events as the direction, not the dependency. If you build on Tasks today, build the polling path, keep the interval generous, and leave a seam where an event callback can short-circuit the wait later. When you do accept server-initiated callbacks, the signature check is not optional, for the same reason a resume token must be verified: an unauthenticated "your task is done" POST is a thing anyone can send you.

The parts that bite

A task handle backed by nothing durable is a worse lie than a timeout. If your task lives in process memory and the pod restarts, the client polls into a void and the job is genuinely lost, with none of the "it finished anyway" grace the blocking version accidentally had. Do not ship Tasks until the state is in a shared store that survives a deploy. This is the single most common way the pattern goes wrong.

Client support is uneven and you cannot assume it. The extension is opt-in by design, and several popular clients did not support Tasks at launch. That is why the capability check in the handler returns a clear error instead of a handle: a tool that silently assumes Tasks will simply break against a client that never declared it. Check capabilities on every call and degrade honestly.

Polling cost and latency trade against each other. A tight interval wastes requests and load; a loose one adds dead time between "done" and "the client noticed." Set pollIntervalMs from the actual work, and let it change over the task's life if you can estimate better partway through. A render that is clearly ten minutes out should not be polled every second.

A task is not an approval workflow. The input_required state is for a value the task genuinely cannot proceed without, the same way tool-call validation and recovery handles a missing argument. If what you actually need is a human to authorize a risky action, that is a different concern with its own audit needs, and overloading a task into a full approval engine will hurt. Keep the task about doing the work.

Cancellation has to actually stop the work. Returning cancelled in the store while the background worker keeps rendering defeats the point and keeps spending tokens and compute on output nobody will read. Wire the cancel through to the worker so a client that gave up stops the bill, which is the whole reason the client cancels on its own ceiling instead of walking away.

The takeaway

The report tool was never broken. It was the right tool forced through the wrong shape: a multi-minute job squeezed into a single request that a careful client was right to abandon at sixty seconds. The Tasks extension fixes the shape rather than the symptom. The server answers immediately with a durable handle, the real work runs in the background against a store every replica can see, and the client follows the task with cheap polls under a ceiling it chose, cancelling cleanly when it has waited long enough. The client's timeout stops being a bug and becomes a bound on a fast read.

Build the plain version first and resist the urge to wait for events. A task row with a lease, a worker that updates status and writes the result on completion, a handler that keeps short work synchronous and only spawns a task when the work is genuinely long, and a poll loop with a real ceiling. That is a day of work, and it turns a tool that fails a third of the time into one that finishes every time and tells you so.

If you have an MCP tool that quietly depends on a client waiting longer than it should, and you are not sure whether to reach for Tasks, a job queue, or durable execution behind it, book a consultation call and we can map the long-running tools in your server before the next timeout turns into a double-render.

ShareXLinkedInviralruparel.com/blog/mcp-tasks-long-running-tool-calls

Questions this essay answers

01What is the MCP Tasks extension?

Tasks is an official extension to the Model Context Protocol, identified as io.modelcontextprotocol/tasks and based on SEP-2663, that lets a server answer a slow tool call without blocking. Instead of holding the request open until the work finishes, the server returns a durable task handle right away and keeps running in the background. The client then polls for progress and collects the result when the task reaches a final state. It shipped as part of the 2026-07-28 specification, which moved Tasks out of the experimental core into its own versioned extension.

02Why did MCP need Tasks when tool calls already worked?

A normal tool call is a single request and response, and clients cap how long they will wait, often somewhere around thirty to sixty seconds. Any tool whose real work takes longer than that, a video render, a large export, a build, a human approval, could not fit inside one request. Teams worked around it by streaming keep-alive noise or splitting the tool in two by hand. Tasks makes the long-running case first class, so the call returns immediately and the slow work happens behind a handle the client can check.

03How does a client get the result of an MCP task?

The server returns a task handle in place of the result, and the client polls a tasks/get style method to read the task's current state. Each task carries a pollIntervalMs that tells the client how often to ask, so the server controls the pressure on itself. The task moves through working and possibly input_required, then settles into one of three final states: completed, failed, or cancelled. On completed, the client reads the result the task carries. A separate events sketch would let the server push a signal instead of being polled, but that part is still early.

04Do I still need durable execution if I use MCP Tasks?

Yes, they solve different halves of the same problem. Tasks is the protocol surface: how the client learns the work is long and how it collects the answer later. Durable execution is what keeps the work itself alive across a crash or a deploy, by recording each finished step so a restart resumes instead of starting over. A task handle backed by an in-memory job that vanishes on the next pod restart is a worse lie than a blocking call. Put the task state in a shared, durable store and the two patterns compose cleanly.

Working on this in production?I help teams make agents like this reliable. Thirty minutes, no pitch.
Viral Ruparel

Generative AI consultant and full-stack architect. Ex-IBM, former Director of Engineering at a YC W20 startup. Writes about what breaks when agents meet reality.