AI Customer Service on WhatsApp: The 24-Hour Window Runs It
Answering inside the window costs nothing. Answering late costs a template.
Written by
xcale Team
Equipo xcale · The xcale Team

AI customer service on WhatsApp is usually pitched as a headcount saving. For an SMB in Latin America the real argument is narrower and easier to check: since 1 November 2024 Meta made service conversations free for every business, so anything you answer inside the 24-hour window the customer opens costs you nothing. Once that window closes, the same answer costs an approved template. Answering fast stopped being a service metric and became a cost line.
This guide covers what has to be decided before you switch anything on: which questions an AI can answer, how to build it, where it has to stop, and what to measure afterwards.
Why customer service on WhatsApp is a clock problem
xcale
Try xcale free for 7 days
Your agent configured, connected to your stack, and answering on WhatsApp — in hours, not weeks.
WhatsApp has no queue and no session. It has one thread that never closes and one clock that does. Three of Meta's rules define the whole operation:
- When a customer messages you, a 24-hour customer service window opens. Inside it you reply in free text — no templates, no formatting rules. It is an ordinary conversation.
- Since 1 November 2024 those service conversations are free for all businesses, per Meta's own pricing documentation. There is no charge for replying to someone who wrote to you first.
- Since 1 July 2025 Meta bills per delivered message, not per 24-hour conversation. Outside the window you can only write using an approved template, and that one is billed.
Read it straight: the channel gives you the answering and charges you for the interrupting. A business that clears the day's questions within the day runs this channel at close to zero marginal cost. A business that replies three days later pays a template every time it has to reopen the thread — and arrives after a decision the customer already made with someone else.
There is one exception in your favour: chats that start from a click-to-WhatsApp ad open a 72-hour window in which everything you send is free. If you buy traffic, that is three full days of support with no per-message cost. Rates, template categories and platform limits are covered in the WhatsApp Business API guide.
Which questions an AI can answer, and which it can't
The honest answer has nothing to do with the AI. It depends on whether your business has already decided what the answer is.
An agent can resolve, unsupervised, anything that is already written down somewhere:
- Order status, delivery times and coverage by area.
- Hours, location, payment methods, and what someone needs in order to buy or book.
- Price and availability for anything in the catalogue.
- Warranty, exchange and return policy — where one exists and is written.
- How to use, install or activate something you have already documented.
And there is a short list that never gets automated. It is shorter than most owners fear and less negotiable than they would like:
- Complaints and angry customers. Someone upset who gets an automated reply gets more upset. That escalates on the first message, not the third.
- Goodwill and exceptions. Waiving a shipping fee, extending a warranty, refunding money — that is a business decision, not a piece of copywriting.
- Professional advice. Diagnoses, dosages, clinical guidance, legal or tax opinions.
- Bad news. A delayed order, a cancellation, a price increase. That gets communicated by a person with a name.
- Anything that isn't written down anywhere. If the business has no policy, no system can invent one. The correct answer there is "let me confirm that with the team."
The test isn't whether the AI can write the answer. It's whether your business has decided what the answer is.
The general order of what to automate first — and why the first reply outperforms everything else — is in how to automate WhatsApp.
How to set up AI customer service on WhatsApp, step by step
Five steps, in this order. The first three require buying nothing.
1. Count the questions, not the features. Open the last two hundred chats and sort them by topic. You are not looking for an average; you are looking for the short list of questions that comes back every week. That list is the scope of the project and its definition of done.
2. Write the answers before you buy anything. This is the actual work, and it is where most of these projects die. Every repeated question needs a written answer with its exception and its limit attached: what shipping to a rural address costs, what happens if the box arrived open, how long the exchange period really is. In xcale that lives as documents and FAQs the agent retrieves by semantic search rather than keyword match, and you can assign a different knowledge base to each agent.
3. Define escalation before you switch it on. Write the rules by topic (complaints, refunds, anything health-related) and by signal (the customer asks for a person, repeats the same question, changes tone). An agent without escalation rules isn't a theoretical risk — it's an angry customer talking to a machine.
4. Decide who owns the thread. On WhatsApp there is one permanent conversation, so somebody has to be able to take it over at any moment without breaking anything. Handoff to a human has to be immediate and must not reset the context.
5. Instrument the clock. Measure how many conversations get resolved inside the 24-hour window before you look at anything else. It is the one number that describes the service and the cost at the same time.
Escalation isn't the fallback — it's the design
What decides whether this works isn't what the agent answers. It's how it hands over what it can't.
When an agent escalates badly, the customer starts over: repeats the order number, repeats the problem, repeats what they already tried. That is the precise moment automation becomes more expensive than answering by hand, because the customer has now spent patience twice.
A good handoff gives the person receiving it three things: the full conversation, what the agent already verified, and why it stopped. The human arrives knowing more than the customer, not less. In xcale that is configurable escalation rules, a prioritised queue, and the complete thread history, with nobody having to reconstruct it.
Whether the agent remembers the customer between conversations — not just inside one — is a different and deeper problem, covered in persistent memory in AI agents.
What to measure, and why deflection rate is the wrong metric
Deflection rate comes from web chat and the call centre, where every contact costs an agent's minutes and the goal is to avoid the contact. On WhatsApp that arithmetic doesn't hold: inside the window the contact is free and the thread is permanent. A "deflected" customer who never writes again may be a lost customer rather than a saved cost.
Five numbers that do describe the operation:
- Time to first useful reply. Not time to first message — an automated acknowledgement is not an answer.
- Share of conversations resolved inside the 24-hour window. Service and cost in one figure.
- Escalation rate with its reason attached. A rate that falls with no reasons recorded isn't an improvement, it's a warning.
- Seven-day recontact on the same topic. It tells you whether the answer resolved anything or just closed the chat.
- Billed templates per reopened conversation. The direct cost of having replied late.
On benchmarks: there is no credible study of response or resolution times segmented by company size for SMBs in Latin America, so you will not find a reference figure here. What does exist is evidence about what customers will accept. In Kantar's study for Meta, The State of Business Messaging — 11,056 adults across 22 markets including Mexico and Colombia, fielded between April and September 2025 — 67.7% agree that getting an answer from an AI chatbot is useful, while only 42.9% believe AI would improve their messaging experience. That ~25-point gap is the distance between "they replied fast" and "they understood me", and it is exactly where this project is won or lost.
How to start this week
Take the five most repeated questions from your last two hundred chats, write the answers with the exceptions included, and decide which topics always escalate. That is an afternoon of work and it is the one part no vendor can do for you.
Then connect your number. The xcale support agent reads the question, answers from your knowledge base, escalates what you told it to escalate, and leaves the full context for whoever picks it up. There is a 7-day free trial — long enough to measure how many of this week's questions were resolved inside the 24-hour window, which is the question that matters. Plans and pricing are published.
Frequently asked questions
Replying inside the 24-hour customer service window the customer opens costs nothing: since 1 November 2024 Meta has made service conversations free for all businesses. Since 1 July 2025 Meta bills per delivered message, and outside that window you can only write using an approved template, which is billed. In practice: answering is free, interrupting has a price.


