Why workflow builders stall at 12 automations
The canvas is not the problem. The problem is that every automation needs an owner, and the person who owns them is the same person you need for everything else.
What we learn from running an agent on real service desks: where autonomy should be granted first, why workflow projects stall, and how to measure any of this without flattering ourselves.
Every failed agent deployment we have looked at shares one decision: somebody switched it on across the board before anyone had watched it work. Autonomy is not a feature you enable, it is a position you arrive at per ticket category, per client, on evidence you collected yourself.
The canvas is not the problem. The problem is that every automation needs an owner, and the person who owns them is the same person you need for everything else.
A $52,000 req that never gets posted is a cleaner return than any efficiency percentage. Here is how to work out whether your ticket mix actually supports deferring it.
Most published automation rates include drafted replies and human-approved actions. Strip those out and the number halves. A resolved ticket has three properties, and all three should be checkable.
An attacker does not need your tenant. They need one helpdesk technician having a bad afternoon. This is why MFA method changes stay gated in Theo permanently, in every configuration.
We grade Theo per work type and score a correct fix executed without the required approval as zero. The interesting half of an eval is where it escalated, not where it won.
Volume, reversibility and how narrow the correct outcome is. Three tests that put password resets and account unlocks first and offboarding last, on every desk we have deployed on.
A log of tool calls tells you what executed. A note tells the next technician what was wrong, what was deliberately not done, and what to watch for. Only one of those builds trust.
If your classes are wide enough to absorb anything, your category reporting cannot tell you where the week went. Notes from building a taxonomy of 17 domains and 142 work types.
Most SLA breaches on an L1 queue are ordinary tickets that arrived at 7pm. The fix is not working faster during the day; it is not having a backlog at 8am.
Book 30 minutes. We'll map your stack, look at your actual ticket mix, and tell you which categories Theo takes in week one.