Judge agents on outcomes (meetings, pipeline) and quality (approval rate, escalation quality, guardrail pass rate) — not raw activity. The Agents dashboard shows both; activity without outcomes is a targeting problem, not a success.
Outcome metrics
Qualified meetings booked, opportunities influenced, pipeline value attributed, and reply-to-meeting conversion. These are the numbers that justify the agent's existence.
Quality metrics
Draft approval rate (target 95%+ in review mode), guardrail pass rate, escalation precision (were escalations genuinely human-needed?), and reply-classification accuracy on your spot checks.
Efficiency metrics
Cost per meeting, accounts engaged per meeting, and sends per reply. Compare across agents and against your historical human-SDR baselines — that comparison is the ROI story.
What not to celebrate
Send volume and activity counts are inputs, not results. An agent sending 3,000 emails for 2 meetings has a targeting or message problem that volume dashboards can hide — always anchor reviews on the outcome column.