Engineering
Design the failure paths before the agent demo
Use explicit task states, bounded retries, effect-aware recovery, and useful human handoffs to keep agent workflows understandable when execution goes wrong.
Distinguish reasoning errors from execution failures
An agent may interpret input, arrange steps, invoke tools, and explain results. Each stage has a different success condition. A generated response can contain a misunderstanding, and a successful tool response can still leave the business task incomplete.
Examine whether the input is sufficient, whether the judgement has support, whether the action completed, and whether the outcome meets the task requirements. These separate questions help identify the appropriate recovery path.
Start failure planning with the task’s external effects. Reading information and modifying records have different recovery requirements. For consequential actions, the system needs to distinguish what is known to have happened from what remains uncertain.
Let task states describe confirmed facts
Useful states can distinguish pending work, active execution, missing information, required approval, completion, and an uncertain outcome. They should help the application and its operators select the next action rather than merely supply a progress label.
State changes need conditions and supporting evidence. Waiting for additional information should not be treated as a tool error, and generating a proposed action should not be treated as approval to execute it.
Longer tasks also need references to confirmed intermediate results. Recovery can then continue from valid work while checking that earlier results still apply to the current input and source versions.
Base retries on the meaning of the action
A retry is safe only when its effects are understood. Idempotent behaviour aims to make repeated requests produce the same intended effect as one request; it does not make every external action inherently occur once.
Where supported, a stable request identifier can associate attempts with one operation. The implementation must also handle the identifier’s scope, conflicting intent, and the relationship between deduplication records and the action’s result.
A timeout means that a definite response was not obtained within the waiting period. It does not establish that the action failed. Check an uncertain write against the external record before repeating it.
Give automatic recovery a stopping condition
Transient connectivity problems and persistent input errors require different responses. The former may recover through limited retries; the latter usually require correction or clarification. Permission failures should return to an identity and access check.
Bound retries by attempts, elapsed time, and resource use, with an appropriate delay between attempts. Once a limit is reached, the task should enter a defined stopped or handoff state rather than continue indefinitely.
Model-led replanning also needs boundaries. Recovery should not grant permission to expand access, change the task’s purpose, or bypass an approval requirement simply because the original approach did not work.
Distinguish reversal from compensation
A task may stop after some actions have taken effect. Its recovery plan should identify those results and determine which can be reversed, which can be corrected, and which require a compensating action. Compensation can itself have business consequences.
Do not assume that recovery restores a world in which nothing happened. Information already communicated or records already consumed by another process may require an explicit correction and further coordination.
Complete necessary checks before significant actions where possible, and retain result references that support reconciliation. Resolving uncertainty becomes harder when several external actions have already depended on it.
Make human recovery part of the workflow
A useful handoff includes the task goal, relevant inputs, confirmed steps, external action references, unresolved issues, and permitted next options. Present the information needed for a decision rather than forwarding an undifferentiated log.
The person taking over also needs clear authority. Specify what can be edited, whether execution can resume, and when further review is needed. Human involvement does not automatically make an uncertain repeated action safe.
Record the resolution and update the task state after recovery. Later processing should recognise that the issue has been handled, and the recovery record should remain available for improving the workflow.
Test whether the system can be operated through failure
Validation should cover missing inputs, malformed responses, timeouts, repeated requests, interrupted execution, and attempts made after recovery. Check the resulting external state as well as the message shown to the user.
Operational monitoring should distinguish automatic recovery, human resolution, and unresolved work. Completion, recovery time, repeated effects, and review effort describe different aspects of reliability.
Revisit the relevant failure paths when task volume, tools, or permissions change. An agent becomes manageable through maintained states, bounded actions, and clear recovery procedures, not through a successful demonstration alone.