2 min readPivot Scale Labs
Where AI actually earns its place in an operations team
Most AI projects fail because they start with a model instead of a job. How to pick the first AI task in a business, and the guardrails that make it safe.
- AI
- automation
- operations
- AI agents
Most AI implementations in a small or mid-sized business stall for the same reason: they begin with a model and go looking for a task. What actually works is the reverse: start with a job that already has an owner, an input and an output, then ask whether AI does that job better than a person and a spreadsheet.
The three questions worth asking first
- Is the work reading, drafting or classifying? Those are the three things current models do reliably at volume. If the task is deciding, negotiating or judging edge cases, it is not an AI task yet.
- Is there a source of truth to read from? An AI that reads your own documents, tickets and records is useful. One that guesses from general knowledge is a liability.
- Can a human review the output cheaply? If reviewing the draft takes longer than writing it, the AI has made the process slower.
If a task passes all three, it is worth scoping. Most candidate tasks fail at least one.
Start in draft-only mode
The first version of any agent should produce a draft that a human approves, nothing else. This is not caution for its own sake; it is how you collect the evidence that decides whether the agent should ever be allowed to send, file or update anything by itself.
Set the guardrails before the first prompt:
- Scope. What the agent may read, and what it may never touch.
- Output shape. Draft only, with the source of every claim attached.
- Logging. Prompt, sources and output recorded so any decision can be reconstructed.
- Cost ceiling. A volume cap, so a runaway loop cannot become an invoice.
- Owner. One named person accountable for its behaviour.
Judge it on a test set, not a demo
A demo shows the happy path. Before anything touches live work, build a small evaluation set, twenty or thirty real examples with the correct output for each, and measure against it. A pass rate you can quote is worth more than a video of the agent working once.
Then widen the scope slowly
Once the agent clears its evaluation set in draft mode, you can let it do more: send a reply without review for the categories it always gets right, file documents autonomously, escalate the rest. Widening scope is a decision backed by evidence, not a launch.
The businesses that get value from AI are almost never the ones with the most ambitious plan. They are the ones that picked one boring, high-volume task, automated it properly, and then moved to the next one.
If you want a second opinion on which task to attack first, bring us the list.
