Theo The frontier AI for IT work
Home/Blog

Blog

What we learn from running an agent on real service desks: where autonomy should be granted first, why workflow projects stall, and how to measure any of this without flattering ourselves.

Featured
Earned autonomy: why the dial starts at zero
Autonomy11 August 20269 min read

Earned autonomy: why the dial starts at zero

Every failed agent deployment we have looked at shares one decision: somebody switched it on across the board before anyone had watched it work. Autonomy is not a feature you enable, it is a position you arrive at per ticket category, per client, on evidence you collected yourself.

Automation
Why workflow builders stall at 12 automations
Automation4 August 2026

Why workflow builders stall at 12 automations

The canvas is not the problem. The problem is that every automation needs an owner, and the person who owns them is the same person you need for everything else.

Economics
The arithmetic of deferring an L1 hire
Economics28 July 2026

The arithmetic of deferring an L1 hire

A $52,000 req that never gets posted is a cleaner return than any efficiency percentage. Here is how to work out whether your ticket mix actually supports deferring it.

Measurement
What should count as "resolved"
Measurement21 July 2026

What should count as "resolved"

Most published automation rates include drafted replies and human-approved actions. Strip those out and the number halves. A resolved ticket has three properties, and all three should be checkable.

Security
The MFA reset is the social-engineering target
Security14 July 2026

The MFA reset is the social-engineering target

An attacker does not need your tenant. They need one helpdesk technician having a bad afternoon. This is why MFA method changes stay gated in Theo permanently, in every configuration.

Evaluation
Measuring an agent honestly is mostly about the failures
Evaluation7 July 2026

Measuring an agent honestly is mostly about the failures

We grade Theo per work type and score a correct fix executed without the required approval as zero. The interesting half of an eval is where it escalated, not where it won.

Rollout
Which ticket categories to grant autonomy on first
Rollout30 June 2026

Which ticket categories to grant autonomy on first

Volume, reversibility and how narrow the correct outcome is. Three tests that put password resets and account unlocks first and offboarding last, on every desk we have deployed on.

Craft
The note is the handover artifact
Craft23 June 2026

The note is the handover artifact

A log of tool calls tells you what executed. A note tells the next technician what was wrong, what was deliberately not done, and what to watch for. Only one of those builds trust.

Taxonomy
"Email issue" is not a ticket category
Taxonomy16 June 2026

"Email issue" is not a ticket category

If your classes are wide enough to absorb anything, your category reporting cannot tell you where the week went. Notes from building a taxonomy of 17 domains and 142 work types.

Operations
The overnight queue decides your whole day
Operations9 June 2026

The overnight queue decides your whole day

Most SLA breaches on an L1 queue are ordinary tickets that arrived at 7pm. The fix is not working faster during the day; it is not having a backlog at 8am.

Rather see it work than read about it?

Book 30 minutes. We'll map your stack, look at your actual ticket mix, and tell you which categories Theo takes in week one.