original research

How Do You Know an AI Agent Finished the Job, Correctly?

An agent saying it is done is a claim, not a fact. Outcome Assurance treats completion as something the evidence has to establish.

Successful tool calls, completed workflows, and clean audit trails do not prove an agent achieved the outcome. A new research paper proposes Outcome Assurance: a task-level framework for verifying that agentic work actually finished, correctly.

Key takeaways

  • Execution is a process; outcome is a state of the world. A successful tool call, a finished workflow, or a complete audit trail does not establish that the intended outcome exists.
  • Assurance attaches to the task, not the model or the agent. Decision, Action, Execution, and End-State Assurance answer different questions and catch different failures.
  • Completion should be a verdict over the task's current assurance state, including claims that may have gone stale during a wait and assurance returned by delegated child agents.

Original source

Read the full research paper

This article adapts 'Outcome Assurance for Agentic Systems: How Do You Know an AI Agent Finished the Job, Correctly?' The full paper includes the formal definitions, the Completion Contract, delegation and revalidation semantics, eleven falsifiable hypotheses, and the proposed experimental design. It is a working draft and reports no experimental results.

Download the research paper (PDF) →

An AI agent reports that it has issued the refund, updated the record, and closed the ticket. The workflow shows green. Every tool call returned success. The audit trail is complete.

Did the job actually get done, correctly?

For a single model response, you can mostly answer that by examining the output. For agentic work, you cannot. An agent's task may span decisions, external actions, approvals, waits, retries, delegated sub-agents, and state changes across several systems. Each of those steps can succeed while the outcome fails.

I have written a research paper that proposes a framework for this problem. I call it Outcome Assurance. This article is the operator's version. The full paper (PDF) has the formal treatment.


Execution Is Not Completion

The core distinction is simple. Execution is a process. Outcome is a state of the world produced by that process. For consequential work, assurance of the process cannot substitute for assurance of the outcome.

The ways this breaks are familiar to anyone who has run distributed systems:

  • A tool call succeeds, but its effect fails downstream.
  • A transaction commits in one system while another required system stays unchanged.
  • An approval expires while execution is suspended, and the agent resumes anyway.
  • A timeout leaves it unknown whether a side effect happened, so a retry may duplicate it.
  • A child agent succeeds while the parent's outcome remains incomplete.
  • Every planned step executes, and the required final state still does not exist.

None of these are uniquely AI problems. What agents add is probabilistic decision-making, dynamically generated execution paths, delegation, and real-world side effects, all at once. And the evidence is that agents asserting success when the environment says otherwise is common, and that LLM judges are poor at catching it.

Each Existing Control Answers a Different Question

Outcome Assurance does not replace evaluation, authorization, runtime verification, observability, or provenance. Its premise is that each of these establishes a different claim, and none of them alone establishes completion.

  • Evaluation estimates how a system behaves across a population of runs. A 95% success rate on a benchmark says nothing about whether this task instance finished.
  • Authorization says an action was permitted. An authorized payment can still fail.
  • Runtime verification checks execution traces against specifications. A trace that says "write succeeded" is not proof that the value persisted.
  • Provenance records who did what. A perfectly complete audit trail can faithfully describe an incorrect result.
  • Self-report tells you about the agent's execution state. It is not evidence about the world.

Recent agent-systems work already gates completion on independent evidence instead of agent self-report. The paper adopts that as established practice. Its question is what has to be added on top of it.

Assure the Task, Not the Agent

The first design decision is where to draw the assurance boundary. The model is too narrow; one task may use several. The agent is too broad; one agent performs many kinds of work with very different consequences. The execution path is the wrong object; two different paths can legitimately reach the same acceptable outcome.

The paper makes the task the unit of assurance: an objective, its constraints and acceptance conditions, and the lifecycle it moves through.

That leads to a practical principle: assurance requirements attach to consequential work in context, not to the agent. The same agent might read an account, recommend a $100 refund, and execute the $100 refund. Those three operations should not carry the same evidence, approval, or verification requirements just because one agent performs them.

Four Layers, Four Different Failures

The framework separates four assurance surfaces. They are not a mandatory four-stage pipeline. One mechanism can check several. They are separated because each establishes a different proposition and catches a different class of failure.

Decision Assurance. Was a consequential judgment made from admissible, valid inputs, under the applicable decision rules, with uncertainty and escalation handled as required? Generated rationale and self-reported confidence can be recorded, but neither establishes correctness.

Action Assurance. Is this specific side effect admissible to attempt: within authority, policy, scope, approval, rate and budget limits, with idempotency and compensation in place? A satisfied verdict here means "allowed to try," not "effect achieved."

Execution Assurance. Should the task continue, pause, retry, compensate, or stop, given its dependencies, checkpoints, deadlines, and the current state of authority and policy? This is not the same as durable execution. A durable runtime can perfectly preserve a stale approval or an unsafe retry decision.

End-State Assurance. Do authoritative observations of the external world (not the agent's report) satisfy the task's outcome predicates at verification time?

A correct end state is necessary but not always sufficient. If a balance needed to move from $1,000 to $900, an unauthorized change satisfies the end state through an unacceptable process. An authorized, conformant process can also fail downstream and leave the balance at $1,000. You need both the process claims and the outcome claim.

Assurance Decays Over Time

Agentic tasks wait: for approvals, external events, human responses, rate-limit resets, or child tasks. While a task is suspended, the policy version, the approval, the authority grant, or the underlying facts may change. The task's identity does not.

So a claim that was well supported at 9:00 may not be usable at 2:00. Restoring execution state is not the same as restoring assurance. The paper proposes that durable agentic systems preserve not only execution history but their assurance claims: the evidence behind each one, what it depends on, how long it is valid, and what would invalidate it. When something changes, the system revalidates only the claims that depend on it rather than starting over.

Delegation Transfers Work, Not Accountability

When a parent agent delegates to a child, two things move differently across the boundary:

Authority attenuates. Assurance obligations do not.

If the parent task is onboarding a supplier and the child performs sanctions screening, delegating the screening does not remove the parent's requirement that the supplier pass a valid screening before onboarding completes. A child can stay within its authority and still produce insufficient evidence. It can also produce a correct result through an unauthorized process.

The paper's added rule: delegated assurance must be reconciled by the parent before it can support the parent's completion. "Every child reported success" is not a completion argument.

Completion Is Its Own Verdict

Taken together, completion becomes a separate, task-level claim. It is evaluated through a Completion Contract that is specified before execution, by a verifier the executing agent cannot unilaterally control. That contract evaluates the authoritative end state together with the process and delegated claims it relies on, and admits only the claims that are still usable at verification time.

The verdict has three values: Verified, NotVerified, or Indeterminate. Indeterminate is a legitimate answer. If a settlement has a ten-minute window and has not cleared at minute ten, that is not failure. It is a reason to wait.

None of these mechanisms is new on its own. Evidence-gated completion, dynamic assurance cases, sagas, and runtime verification all have strong precedent. The paper's claim is their composition at the task level: completion evaluated over a current assurance state that survives time and delegation.

Reliability Is a Different Question

Outcome Assurance answers whether one task reached a verified outcome. Outcome Reliability answers how often a defined agent configuration produces verified outcomes across a defined population of tasks under stated conditions.

That distinction matters for deployment decisions. Conventional task-success and consistency metrics can rank configurations differently from a reliability profile built on verified completion, because the conventional metrics count false completions as successes. The paper recommends reporting a reliability profile, not a single score: first-pass, eventual, and on-time reliability, under nominal and stressed conditions.

What the Paper Does Not Claim

This is a proposal to be tested, not an established result. The paper reports no experimental data. It specifies eleven falsifiable hypotheses and one experimental program, each with a stated condition under which the simpler approach should win. If Outcome Assurance does not measurably reduce false completion relative to strong evidence-gated completion, the simpler architecture should be used.

It also does not solve alignment, bad specifications, or bad evidence. An agent can satisfy an incomplete specification and still produce the wrong real-world result. An authoritative system of record can be wrong. And assurance has a cost. Low-consequence tasks should not carry a full assurance case.

What to Change Now

You do not need the full framework to act on its central distinction.

  • Stop accepting "done" from the agent. For any consequential task class, require completion to be confirmed against authoritative external state by a verifier the agent does not control.
  • Write acceptance conditions before execution. If you cannot state what "complete" means for a task class, you cannot verify it, and the agent will define it for you.
  • Revalidate on resume. Long-running agent tasks should recheck approvals, authority, and policy versions when they resume, not just restore state.
  • Make parents reconcile children. In multi-agent systems, child results should return structured evidence that the parent checks against its own requirements.
  • Report verified outcomes, not runs. Board and operating dashboards should track the rate of verified completions, false completions caught, and unresolved tasks, not task volume or tool-call success.

Agentic systems change the unit of evaluation. The question is no longer only whether a model produced a correct response. It is whether a task involving decisions, actions, external systems, time, and other agents produced an acceptable result under the conditions that governed it.

Download the full research paper (PDF).