An MCP tool call can succeed while the job is unfinished
A support agent calls a tool to refund an order. The tool returns a refund identifier. The agent tells the customer the refund is complete.
In this hypothetical workflow, the payment provider has accepted the request for processing. The customer’s money has not arrived yet.
The tool did its job. The agent’s summary still went too far.
This is an evidence problem as much as a wording problem. If the only durable record says “refund successful,” a later reviewer cannot tell whether that means the tool ran, the provider accepted a refund, or the customer received the funds.
MCP makes tools available to agents. The meaning of a business outcome still has to come from the systems responsible for that outcome.
Tool results have a specific scope
The MCP tools specification distinguishes protocol errors from tool execution errors. A tool can return a result containing an execution error, indicated by isError: true. Receiving a result message is therefore not the same as successful execution.
Even a result with no reported tool error has a limit. It tells the client what the server returned. A successful submission to a downstream service may still require later processing. A valid output schema tells you the result has the expected structure; it does not establish that every claim in the result is correct.
The MCP Tasks extension, in its revision dated July 28, 2026, makes this separation explicit: a task can reach completed while its tool result contains isError: true. That status reflects protocol completion, not a successful business outcome. The extension is versioned separately and continues to change, so check the revision your implementation uses. The lesson for audit design holds regardless: inspect the result as well as the terminal status.
Tool annotations also need care. A tool described as read-only or idempotent is making a behavioral claim. The MCP specification tells clients to treat annotations as untrusted unless they come from trusted servers. An annotation is no substitute for the access controls and enforcement around a consequential call.
For audit purposes, record the tool result at its actual level of certainty. “Provider accepted refund request” is useful. “Customer refunded” requires evidence supporting that later state.
Approval, invocation, and outcome are separate events
For the refund example, review should distinguish at least these events:
- An approval system permitted a refund for a specific order and amount.
- The tool server received an invocation with specific arguments.
- The server reported a downstream response with a refund identifier.
- The provider later reported a status for that refund.
If the customer or another system supplies a receipt observation, that is another event with another source. It should retain that source.
These distinctions matter when the amounts differ. Suppose the approval allows ₹2,000, but the tool invocation requests ₹2,500. A record of a valid approval is insufficient unless the review compares its scope with the actual invocation.
The same applies to time and target. An approval for one account cannot establish authority for another. An expired approval may remain a validly signed artifact while being unsuitable for a later action. Authorization enforcement belongs in the execution path; review must preserve enough evidence to assess what that path reported doing.
What useful MCP tool audit evidence contains
Start with a consequential tool, such as changing an entitlement or submitting a refund. Preserve enough information to answer a concrete review question.
| Review question | Material to preserve |
|---|---|
| Which tool was invoked? | Tool identity and the service exposing it |
| What action was requested? | Necessary arguments or a documented, inspectable representation |
| What approval applied? | Approval reference, scope, and relevant validity period |
| What did the server report? | Result, error state, and downstream references |
| Was this a retry? | Logical operation reference and a distinct attempt reference |
| Who can assess the outcome? | The responsible downstream system and its native evidence |
This is a review design, not a proposed MCP message schema. The exact mapping depends on the tool and the record format in use.
Signing a tool-server record makes its statement portable and protects its integrity. It does not make the tool server an independent witness to everything beyond it. If an exporter signs a statement derived from a tool log, identify the exporter as the signer and preserve the transformation it performed.
An argument digest is useful when the reviewer has the corresponding representation and can recompute it. Without that material, it is a commitment to content the reviewer cannot inspect. Report that limit rather than presenting the digest as proof that the arguments were acceptable.
Timeouts deserve their own evidence
Imagine the refund request reaches the provider, but the tool server times out before receiving the response. The agent retries.
The first timeout does not establish that no refund was created. The second invocation does not establish that two refunds were created. Resolving the uncertainty requires the provider’s state and the workflow’s retry semantics.
Preserve both attempts. Connect them to the same intended operation without collapsing their different observations. Keep idempotency enforcement in the tool and downstream service; signed records document attempts and reported results but cannot prevent duplicate execution.
The uncomfortable state is sometimes the accurate one: “Request sent; downstream result unknown.”
An agent should be able to carry that uncertainty into its response instead of converting it into success or failure for conversational convenience.
Make the recipient part of the design
A tool audit trail becomes useful when someone can act on it. For the refund workflow, that might be a finance operator checking duplicate refunds or a merchant answering a customer’s escalation.
Design the handoff for that recipient. Include the records they need, identify the issuers they accept, and preserve provider evidence under its own verification rules. Explain what is missing. Do not require a full transcript if scoped evidence answers the question.
PEAC Protocol provides signed interaction records that can accompany MCP workflows. Originary’s work is to help select, issue, verify, and hand off those records for review. Neither the record nor its signature takes over the tool’s execution or the payment provider’s state machine.
For a first implementation, pick one write operation and one failure worth investigating. Run the normal path, an approval mismatch, and a timeout followed by a retry. Give the resulting evidence to the person responsible for reconciliation.
The acceptance criterion is whether they can distinguish an authorized request, a reported execution, and an unresolved outcome. A tool call counter cannot make that distinction for them.