Agentic Work Needs a Contract, Not Just a Goal
Agentic AI cannot run on goals alone. This post explains why enterprise agents need explicit work contracts that define responsibility, expected outcomes, acceptance criteria, and execution limits before autonomy can become predictable.
TECHNOLOGY & ARCHITECTURE
10/1/20267 min read


Why a Goal Is Not Enough
One of the easiest ways to start using an AI agent is also one of the most dangerous.
We give it a goal.
“Investigate why this invoice has not been paid and resolve the issue.”
At first, this sounds reasonable. It has a business object. It has a problem. It gives the agent room to think. That is usually the promise of agentic AI: don’t prescribe every step, let the agent reason its way through the work.
But the moment we try to use this in an enterprise system, the instruction starts to look less clear.
What does “resolve” mean?
Does it mean identify the reason for the delay? Gather evidence? Contact the supplier? Update the invoice record? Recommend the next step? Decide the case can be closed? Approve payment?
A human may understand the boundary because they bring organizational context with them. They know their role, authority, escalation path, and the habits of the process. An agent does not inherit all of that automatically. It receives the goal, the context we provide, the tools it can access, and whatever boundaries the system has actually made explicit.
If those boundaries are not clear, the agent has to infer them while doing the work.
That is a problem.
We wanted the agent to use probabilistic reasoning to decide how to pursue a responsibility. We did not want the agent to decide what responsibility had been delegated in the first place.
The opposite mistake is also possible. We can remove ambiguity by scripting every step.
Retrieve the invoice. Compare it with the purchase order. Check receipt status. Inspect supplier correspondence. If the amount differs, ask for clarification. If receipt evidence is missing, route the case for review.
That may be a perfectly valid workflow. But then we should be honest about what we have built. We have automated a procedure. We have not really delegated agentic work.
The useful space sits between these two extremes.
An agent should not receive an arbitrary natural-language goal and decide for itself what work it owns. But the system should also not prescribe every internal step so tightly that the agent has no meaningful reasoning left to perform.
This is why agentic work needs a contract.
The First Boundary: What Work Can This Agent Accept?
The first part of that contract is the kind of responsibility the agent is designed to accept. I call this an Assignment Type.
An Assignment Type is an explicit, bounded operational responsibility that a particular agent has been designed to accept. It is narrower than an open-ended goal, broader than a workflow procedure, and stable enough for another system to invoke through a contract.
That “particular agent” part matters.
Almost all leading language models can talk across many enterprise domains in the same conversation. But an enterprise agent is not just a model answering questions. It is enterprise software acting in an enterprise context. If it reads enterprise data, changes enterprise records, communicates externally, consumes governed systems, or influences a business outcome, the enterprise remains responsible for what happened.
A finance investigation agent should not accept every finance-related request merely because the underlying model can discuss invoices, payments, suppliers, procurement, and accounting. The agent’s responsibility has to be narrower than the model’s language capability.
The model may have broad knowledge. The agent should have bounded responsibility.
This is why a generic `/execute` interface is so tempting and so dangerous. It treats the agent as if its responsibility can be discovered at runtime from whatever natural-language goal it receives. That may be fine for a personal assistant or exploratory tool. For enterprise work, it is not a valid enterprise contract.
For example, “resolve accounts payable issues” is too broad. It names an area where something may be wrong and asks the agent to work out what responsibility it is supposed to carry.
At the other extreme, “compare the invoice amount to the purchase-order amount” is too narrow. That may be a useful step, but it is not the full responsibility.
A useful Assignment Type has to sit between those two extremes. It should name a bounded responsibility the agent can accept, without reducing that responsibility to a single procedural step.
The Assignment Type answers only the first question:
What kind of work may this agent accept?
It does not yet describe the particular work in front of us.
Turning Responsibility Into a Specific Assignment
That is where the Work Assignment comes in.
A Work Assignment is a contextualized instance of an Assignment Type. It binds the agent’s responsibility to a specific enterprise situation.
In the invoice case, this is where the abstract responsibility becomes a particular invoice, a particular purchase order, a particular supplier, and a defined set of relevant records. The agent now has a real piece of work. But the system has still not told the agent exactly how to reason. The path remains open inside a bounded assignment.
Even that is not enough.
The caller also needs to know what result should come back.
The Result Must Be More Than a Paragraph
This is the expected outcome.
For predictable autonomy, the expected outcome should define a structured result the delegated responsibility must return so the invoking process can decide what happens next, without giving that next decision to the agent.
In the invoice case, the agent should not simply return a paragraph saying, “The discrepancy has been resolved.” That tells the surrounding workflow too little. It should return an investigation disposition: what discrepancy was found, what explanation is supported, what evidence was reviewed, what remains unresolved, and how much confidence or uncertainty attaches to the finding.
For example, the result category might be:
discrepancy explained by supported evidence;
discrepancy explained but evidence incomplete;
conflicting records found;
supplier clarification required;
insufficient evidence within assignment limits;
out of scope for this assignment.
The exact labels do not matter. They will vary by use case. What matters is that the invoking process receives a result it can interpret.
Natural language still matters. Humans need to understand the reasoning, evidence, and uncertainty. But natural language should not be the only output. If the agent returns only prose, the system has to interpret the prose to decide what happened.
Structured output and natural language serve different purposes.
The structured result lets the invoking process route the work, apply rules, detect unresolved cases, trigger review, and avoid treating uncertainty as closure. The natural-language explanation lets a human understand why the result was produced.
Predictable autonomy needs both.
Who Decides Whether the Work Is Good Enough?
Then comes acceptance criteria.
The expected outcome defines what must come back. Acceptance criteria define what must be true for that return to count as fulfilment of the assignment.
This matters because an agent can return the right kind of result and still fail the responsibility.
It may cite no evidence. It may classify a discrepancy as explained while leaving the key mismatch untouched. It may report high confidence when supporting records are incomplete. It may blur confirmed facts, assumptions, and unresolved gaps.
Acceptance criteria should be predefined and attached to the expected outcome structure. They may require judgment to evaluate, but they should not be vague after-the-fact opinions about whether the agent “did well.” In the invoice case, they should make clear what evidence must be cited, how uncertainty must be handled, and when the agent is not allowed to claim that the discrepancy has been explained.
The agent can see these criteria. It can even report how its result addresses them. But the agent should not be the final authority that decides whether its own work fulfilled the contract.
That evaluation belongs outside the agent’s self-declaration.
How Far Should the Agent Be Allowed to Continue?
Finally, the contract needs execution constraints.
Even if the responsibility is clear, the work instance is clear, the expected outcome is clear, and the acceptance criteria are clear, the agent still needs limits on how far it may continue.
An invoice investigation agent should not contact the supplier every day until someone replies. It should not keep expanding the investigation indefinitely because another record, another email thread, or another follow-up might improve confidence.
The contract must bound the pursuit of the work, not only define the work.
Execution constraints may include attempt limits, deadlines, scope limits, cost or effort limits, and return conditions. In the invoice case, this is where the contract should define how many supplier clarifications may be attempted, which records may be inspected, what cutoff matters, and where the investigation must stop.
A technical retry is not the same as another meaningful business attempt. Retrying a failed API call to retrieve invoice data is not the same as contacting the supplier again.
These distinctions need to be in the contract.
The Contract Behind Agentic Work
A sample invoice-discrepancy contract can therefore look like this.
Assignment Type
Investigate invoice and purchase-order discrepancy.
This defines the bounded responsibility the agent is designed to accept.
Work Assignment
Investigate the discrepancy between invoice INV-123 and purchase order PO-456 for supplier ABC Components, using the available invoice, purchase order, receipt records, and supplier correspondence.
This turns the responsibility into a specific piece of work in the enterprise context.
Expected Outcome
Return an investigation disposition that identifies the discrepancy, states the result category, cites evidence reviewed, records confidence or uncertainty, and lists unresolved gaps.
This tells the invoking process what kind of result must come back. The result category might be “discrepancy explained by supported evidence,” “discrepancy explained but evidence incomplete,” “supplier clarification required,” “insufficient evidence within assignment limits,” or “out of scope for this assignment.”
Acceptance Criteria
The result must identify the discrepancy, cite the records used, distinguish confirmed facts from assumptions, use a permitted result category, and avoid claiming a supported explanation when required evidence is missing.
This defines what must be true for the returned result to count as fulfilment of the assignment.
Execution Constraints
Supplier clarification may be attempted only within the permitted contact limit and time window. The assignment must return before the payment cutoff. The agent may inspect invoice, purchase-order, receipt, and supplier-correspondence records, but must not open a broader vendor-risk investigation.
This defines how far the agent may pursue the work before returning control.
None of this makes the agent’s reasoning deterministic. That is not the goal.
The goal is to make the delegation explicit. Inside the contract, the agent can still reason, investigate, adapt, compare evidence, reject one hypothesis, follow another, and produce a finding the surrounding system did not know in advance.
But the enterprise should not have to discover during execution what work the agent thought it owned, what result it thought was enough, or how long it thought it could continue.
Agentic work needs a contract.
Not because we want to remove reasoning, but because we need to know where the reasoning is allowed to operate.