The AI receptionist goes live on a Monday. By Friday, the owner wants to know if it is working. The usual answer is a dashboard of call volume and average call length. Those numbers move every week, and neither tells you whether you booked more work or lost it.
Pick a short list before launch and write down how each is calculated. Six numbers and one weekly habit will carry the first 30 days.
Answer rate, counted by hour
Answer rate is the share of inbound calls that reached a live voice instead of voicemail or a ring-out. Count it by hour and by day, not as one monthly figure. An overall 97% can hide a dead stretch every Friday after 4 p.m., and that stretch is where a roofing owner loses a storm lead or a broker loses a spot quote.
Also count calls where the caller hung up within ten seconds. Those people heard the greeting and left, so the opening line is the problem.
Time from call to human follow-up
Answering fast is the easy part. What decides revenue is how long a person on your team takes to act on what the AI collected. Measure the gap between the end of the call and the first human touch: a callback, a text, a quote, a dispatch. Track the median and the slowest 10% of calls, because the slow tail is where leads die.
The best-known evidence is old. A 2007 study by MIT's James Oldroyd with InsideSales.com looked at web-generated leads and reported that the odds of contacting a lead called within 5 minutes versus 30 minutes dropped 100 times, and the odds of qualifying it dropped 21 times ([lead response summary](https://www.onecavo.com/wp-content/uploads/2015/11/MIT-InsideSales.com_Lead-Response-Management.pdf)). That study is vendor-linked, old, and about web forms, not phone calls. Treat it as a direction, not a forecast.
Share of calls with a complete record
Define "complete" before you count. For a roofing company it might be name, callback number, property address, and the reason for the call. For a brokerage it might be company, callback number, lane, equipment, and pickup date. Then pull every record from the week and ask one question: could someone act on this without calling the person back to re-ask basics?
Report the share that passes, and tally which field is missing most often. If it is always the address, fix the question, not the people.
Appointments booked or quotes created
This is the outcome number. Count appointments set, inspections scheduled, or quote requests created in your CRM or load system, and divide by the calls that were real prospects. Leave out wrong numbers, vendors, and spam, or the rate will look worse than it is.
If the AI collected the details and a human booked the job the next morning, that still counts. Label it, so you can see which conversions happen on the call and which after follow-up.
Escalations
Count how many calls the AI passed to a person. For each one, check whether the reason matched a trigger you wrote down, whether a human picked up, and how long the caller waited.
There is no right percentage. A rate near zero can mean the AI is mishandling calls it should have passed along. A very high rate can mean the rules are too nervous to be useful. Read ten of them and decide.
Calls the AI should not have handled
This is the number a dashboard never shows you. Keep a running list of calls where the AI should have stopped and transferred, or never taken the call at all. Common examples are an existing customer with an active job, an angry caller, a carrier reporting a breakdown, a claim question that needs a licensed person, or a caller who asked for a human twice.
Each entry is a gap in your rules. Count them weekly and tag the cause. A cause that repeats is your next script change.
Vanity numbers
Total call volume tells you how busy the month was. Average handle time misleads, because a short call can be a caller who gave up. "Calls handled by AI" looks good on a slide and says nothing about whether the right ones were handled. Sentiment scores are fine to glance at, but they do not replace reading the call.
If a number cannot lead to a decision about a rule, a field, or a staffing hour, drop it from the weekly report.
Sample transcripts every week
Every week, pull 15 to 20 calls. Do not choose only the short ones or the ones that went well. Make sure the sample includes:
- A few escalations - A few calls that ended without a booking - A few from the slowest follow-up cases - A few picked at random
Read the transcript, or listen to the audio, with the CRM record next to it. Mark each call as handled correctly, handled with a small error, or should not have been handled. Write one sentence on what went wrong. The owner or office manager should do this, not only the vendor.
Confirm your call recording and disclosure setup with your own counsel before you store audio or transcripts. State rules on recording differ, and this article is not legal advice.
When to change the script
Do not edit the script after one bad call. Change it when the same failure appears in at least three calls in the weekly sample, or when one failure is serious enough that you would rather turn the rule off than wait a week. That bar is a judgment call, not a standard. Pick yours and keep it.
Change one thing at a time and write down the date. Then check the same number the following week. If you rewrite five things at once, you will not know which one helped.
The NIST AI Risk Management Framework describes this kind of discipline for any AI system. It calls for regular tracking of risks in deployed settings (MEASURE 3.1) and for post-deployment monitoring plans that include user input, override, incident response, and change management (MANAGE 4.1) ([NIST AI RMF 1.0](https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf)). A one-page change log is a small version of that.
If you want help setting up the first-30-days scorecard for an AI receptionist, [book a call with Chosen AI Solutions](https://chosenai.co/book).
Sources
- Oldroyd, James B. and Dave Elkington / InsideSales.com. Lead Response Management study summary. [PDF](https://www.onecavo.com/wp-content/uploads/2015/11/MIT-InsideSales.com_Lead-Response-Management.pdf) - National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023. [nist.gov](https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf)