Workflow / decision guides
Before You Give AI a Work Queue, Define What Good Looks Like
Cihan's view: TRY one bounded workflow with a small set of good examples, explicit failure checks, and a named reviewer; SKIP unattended actions when nobody can define or inspect a good result.
An AI agent can now do more than answer a question. Depending on the product and permissions, it may review leads, summarize support requests, generate reports, update tickets, or run on a schedule.
That is useful. It is also where many AI projects become vague. “Let the agent handle this queue” sounds like a plan, but it hides the important question: what counts as a good result, and who checks it?
Thesis: before you automate a recurring work queue, build a small review loop. Define the expected result, collect representative examples, inspect failures, and keep a named human owner at the consequential decision point.
Evidence label: evidence-based editorial using vendor documentation and public risk-management guidance. It is not a report of a personal tool test or a customer result.
The strongest case for moving faster
There is a reasonable counterargument. If a task is repetitive and reversible, adding a formal evaluation process can feel like paperwork. A team may learn more by letting an agent handle a small batch than by spending a week designing a perfect test. For low-risk work, speed matters.
That argument is partly right. You do not need a research lab to automate a weekly report. You do need a way to notice when the report is wrong before it becomes someone else’s decision.
OpenAI describes this distinction in its business guidance on evaluations. Its recommendation is to turn a business objective into explicit examples of desired and undesired outputs, test against realistic inputs, analyze errors, and keep measuring after launch. The point is not to predict every failure. It is to make failures visible enough to improve the system.
Anthropic makes a similar distinction between workflows and agents. Workflows follow predefined paths. Agents dynamically choose their process and tool use. Anthropic recommends starting with the simplest solution and adding agentic complexity only when the flexibility is worth the extra cost, latency, and unpredictability.
The practical translation is simple: do not buy autonomy when a checked template will do.
A small review loop for a real work queue
Use this before automating a recurring task such as classifying inbound requests or drafting a weekly operations report.
1. Write the output contract
Describe the result in plain language. For example:
Each request receives one category, a short reason tied to the source text, an urgency flag only when the stated rules support it, and an “unclear” label when the evidence is insufficient.
Then write what the system must not do. It must not invent an SLA, infer a customer priority from tone alone, or close a request without approval.
This is more useful than “be accurate.” Accuracy needs a task, a boundary, and a visible failure condition.
2. Assemble a small reference set
Collect representative examples from the real task, with permission and sensitive details removed. Include ordinary cases, ambiguous cases, and a few expensive-to-mishandle cases.
For each example, record the expected category and the reason a knowledgeable reviewer would accept it. This becomes a working reference set, not a claim that the examples cover every future situation.
OpenAI calls a curated set of examples a “golden set” and recommends using it with error analysis. You can start much smaller for an internal experiment. The important part is that the examples reflect the team’s judgment, not just the model’s preferred wording.
3. Review the risky boundary
Separate preparation from action. It may be acceptable for an agent to draft a reply or group requests. Sending the reply, changing a customer record, approving spend, or closing a case may need a human checkpoint.
NIST’s Generative AI Profile recommends defining roles and responsibilities, documenting relevant data and system information, and planning ongoing monitoring and periodic review. That is governance language, but it maps neatly to office work: name the reviewer, keep the source trail, and decide how often the process is checked.
What to measure without building a dashboard
For a first trial, track four things in a simple table:
- Accepted: the output can be used after normal review.
- Corrected: the structure is useful but a person had to fix it.
- Escalated: the input was ambiguous or outside the rules.
- Unsafe or unsupported: the output invented, overreached, or took an action it should not take.
Do not compress these into one impressive accuracy percentage. A system that is usually correct but occasionally sends an unsupported promise may need a different control than one that merely mislabels a low-risk request.
Also record review effort. If the agent produces a polished draft that takes longer to verify than the original task, the workflow may be saving keystrokes rather than time.
The limit: review can become theatre
A human-in-the-loop label is not a safety system by itself. If the reviewer sees hundreds of outputs at once, has no source links, or is expected to approve everything in seconds, the checkpoint is decorative.
Anthropic’s 2026 research on agent autonomy found that experienced users tend to auto-approve more often while also interrupting more frequently when they do intervene. That is a useful warning: familiarity may change how people supervise an agent. A team should monitor what reviewers actually inspect, not merely document that a reviewer exists.
TRY, SKIP, USE
TRY a bounded, reversible queue with a small reference set, explicit “unclear” handling, and a named reviewer.
SKIP unattended actions when nobody can define a good result, inspect the source, or undo the change.
USE the workflow more broadly only after the review record shows which errors occur, how costly they are, and what control catches them.
Sources and next step
- How evals drive the next chapter in AI for businesses — OpenAI
- Building effective agents — Anthropic
- Measuring AI agent autonomy in practice — Anthropic
- NIST Generative AI Profile
- Choose your first AI workflow at work
For a first experiment, choose one queue you can pause and reverse. Keep the source, the output, the reviewer decision, and the correction in the same record. That is enough evidence to decide whether more autonomy is earned.