How do you measure a synthetic agent? The metrics that matter
You measure a synthetic agent on four fronts: business results (the work it gets done), quality (how well it does it), control (whether it follows its rules and when it needs a person) and health (whether it is available in the real channel). Set a baseline for each metric before you start; without a “before,” there is no honest way to say it improved.
Why is measuring an agent different from measuring a tool?
A tool is judged on whether it works. An agent holds a role, so it is measured the way a role is: by the work it gets done, the quality of that work and how far it can be trusted. And because it acts for the business, it must also be controllable: resolving cases isn't enough; you need to know when and how.
Which results metrics do you use?
The ones the team the agent supports already uses, not metrics invented for the agent. That way results compare with what came before. Examples by role:
- Sales: leads handled and qualified, meetings booked, time to first response, opportunities passed to the closing team.
- Customer support: requests resolved without handoff, time to resolution, customer-reported satisfaction.
- Collections: effective contacts, payment commitments logged, payments received after contact.
- Front desk and scheduling: appointments booked and confirmed, cancellations rescheduled, no-shows.
- Admin and accounting: documents processed, reconciliations closed, items that needed correction.
How do you measure quality?
By having a person review real conversations against written criteria: did it answer what was asked, use the right information, sound like the company, deliver what it promised? It is done on an ongoing sample, not once, and each error found is logged along with the rule that was corrected. Automated scoring can help decide what to read first, but it doesn't replace a person reading.
What shows the agent is following its rules?
- Handoffs to a person: how many cases it escalates and why. If it never escalates, something is going unseen; if it escalates everything, it isn't doing the job.
- Human interventions: how often a person takes over a conversation and what triggered it.
- Out-of-rule actions: attempts to do something the rules don't allow, and whether they were stopped.
- Complaints and corrections: cases where a customer or the team had to correct the agent.
- Log coverage: that every action has a record with a date and a result.
How do you know the agent is healthy?
By measuring in the channel it serves, not just on the server: that it receives messages, replies and does so on time. An agent can have a running server and still be silent to the customer. You watch availability, response time, errors in the integrations with your systems and the alerts sent to a person when something failed. Monitoring is part of Synthetic OS.
What do you compare the results against?
A baseline: how the process performed before the agent, measured the same way. If that “before” doesn't exist, build it in the first weeks, even from a manual sample. Then compare against the same role or against a sample handled by people. This page carries no benchmark numbers on purpose: what is reasonable in collections isn't in front-desk work, and a blanket figure would be made up. A vendor who promises a percentage without knowing your process is selling an expectation, not a measurement.
Who sees the metrics, and how often?
The process owner, not just the vendor. A periodic report with the metrics above and room to decide: which rule to change, what to expand, what to stop. A dashboard nobody reads governs nothing. How this ties to control is in how a synthetic agent is governed.
What are the warning signs?
- Results metrics improve, but complaints rise too.
- The agent almost never hands a case to a person.
- Handoffs cluster around a type of request that wasn't in its rules.
- Nobody on the business side has read real conversations in weeks.
- The numbers the vendor presents can't be traced back to the log.
What resolution rate is good?
There is no figure that fits everyone: it depends on the role, the process and what counts as “resolved.” What helps is defining in advance what counts, measuring the baseline and improving against it.
How long until the results show?
It depends on the role and the volume. The sensible approach is to agree up front on what will be measured, how often it will be reviewed and what changes follow from what you see.
Can I measure what the agent does without being technical?
Yes. The results metrics are your own team's, and quality review is reading conversations against clear criteria. The technical side, such as availability and integration errors, is watched by the system, which alerts a person.