Everyone is talking about higher output. We're talking about real impact.
The last article was about governance: who implements, who reviews, and who is accountable. That leaves the most uncomfortable question of the entire series: Is it actually worth it?
The ROI question leads to made-up numbers
“What’s the ROI for us?” sounds straightforward. In practice, you estimate how long something used to take, measure how long it takes now, multiply that by an hourly rate, and you have a result. But the whole thing is based on memory.
We know this from personal experience. At the AIC Group, we started working with our agents without first conducting any systematic measurements. Today, we have the distinct impression that we’ve become faster. We can only partially substantiate this.
Three questions can help. None of them is enough on its own.
1. Objectively: How long did it take before?
The approach: Take a specific process and measure it—duration, correction loops, frequency. Repeat this after each phase of the agent’s development, including the time spent on testing and rework. If you only track the time spent on development, you’re only measuring half the work. This yields a range rather than a single percentage, and it allows you to determine when further development is no longer worthwhile.
The limit: If the agent runs, the state before that is gone. What remains is a feeling—and that feeling is deceptive. In a controlled experiment conducted by the research organization METR in 2025, experienced developers using AI tools took 19 percent longer but were convinced afterward that they had been 20 percent faster. A small study using the tools available at the time. The gap between perceived and measured results remains instructive nonetheless.
What you can do to catch up: Ticket systems and file versions often still reflect the turnaround times from the past. And a handful of similar tasks—sometimes with an agent, sometimes without—can at least give you a rough idea.
Our status: This is where we have the biggest gap.
2. Subjectively: What has become easier for you since then?
Some of the effects don’t show up in any measurements. A task you’ve been putting off suddenly gets done on Monday morning. A draft takes shape without you having to stare at a blank document for twenty minutes first.
That’s not proof of speed. It answers a different question: whether the agent will stick around. Where the relief is noticeable, usage grows on its own. Where it merely saves time, usage stagnates. Part of it can even be quantified: Is the agent still being opened daily after three months?
You’ll find out the rest through conversation—on a regular basis and in concrete terms. What has become easier for you since then? What are you doing now that you used to put off? Where do you still run into problems?
Our situation: a lot of experience, but very little of it written down.
3. Result: Does quality hold up as volume increases?e?
The third question focuses on the end result. In our marketing department, these are the reach, clicks, and interactions of the posts the agent helped create. Each post is fed back into our planning along with its key metrics.
The caveat: These figures measure the quality of the result, not the contribution of AI. In our analyses, formats featuring humans perform significantly better than product explanations. This is an insight into formats. Attributing it to AI would be convenient but incorrect.
What the numbers can do: show whether quality holds up as quantity increases. A simple metric for each agent who submits drafts fits this purpose: How much of the draft survives the revision process?
And here’s a principle from software development: a small, fixed set of test tasks that is run again after every adjustment to the Prompt system or the knowledge base. It reveals whether an improvement in one area has caused a deterioration elsewhere. Our own agent does not yet have this set of test tasks.
Our booth: This is where we’ve come the farthest.
Three Questions to Take With You
How long did it take before? Measurement before the agent runs, including testing time. Missed it? Analyze old system data or do a quick comparison with and without the agent.
What has become easier for you since then? Checking regularly and tracking usage.
Does quality hold up with higher volume? Focus on consistency rather than growth, and validate adjustments with fixed test tasks.
Only when all three are combined do they answer the question from the beginning—without having to make up a number.
And then it starts all over again
Whatever the measurement reveals becomes the basis for the next adjustment: a gap in the knowledge base, a guardrail that’s too restrictive, or a task that the agent would be better off not taking on at all.
This brings all three together: Readiness determines whether the team can handle it. Governance determines who is responsible. Performance determines whether it works.
Next time, we’ll look back at everything that’s come together in this series. And look ahead.
Write to us. We’d be happy to share our approach.