CSAT (Customer Satisfaction Score) measures how satisfied a customer is with a specific interaction, usually via a quick post-conversation question like "How satisfied were you with this help?" on a 1-5 scale. Your CSAT score is the percentage of respondents choosing the top ratings (4 and 5) out of all responses.
How is CSAT calculated?
The standard formula:
CSAT = (number of satisfied responses, usually 4s and 5s / total responses) x 100
If 200 customers answer your post-chat survey and 168 pick 4 or 5, your CSAT is 84%. Three practical notes:
-
CSAT is transactional. It scores this conversation, not the brand overall (that is closer to NPS territory). That makes it the right instrument for judging support quality, human or AI.
-
Response bias is real. Angry and delighted customers answer more often than the satisfied middle. Track response rate alongside the score; a CSAT of 90% on a 5% response rate is a weaker signal than 85% on a 30% response rate.
-
Trends beat absolutes. Cross-company CSAT comparisons are noisy because question wording, scale, and timing differ. Your own week-over-week and per-intent trends are the actionable data.
How do you ask for CSAT in chat specifically?
Chat gives you the best survey placement in support: the customer is already in the thread, seconds after resolution. What works:
-
Ask immediately after resolution, in the same conversation, not by follow-up email hours later.
-
Keep it to one tap. A 1-5 scale, stars, or emoji options outperform multi-question forms in chat.
-
Ask in the customer's language and register. A customer who just chatted in Egyptian Arabic should get "قولنا رأيك، الخدمة عجبتك؟" (olna ra2yak, el khedma 3agabetak?, "tell us, did you like the service?"), not a formal English form. In MENA, survey language mismatch itself suppresses response rates.
-
Add one optional open field ("anything we could do better?") for the diagnostic texture the number lacks.
What actually moves CSAT in chat support?
Across chat support generally, the drivers are consistent, and they compound:
-
Speed to first meaningful reply. Waiting is the single most cited frustration in chat. This is where AI changes the game: on our platform, replies average under 14 seconds around the clock. Benchmarks per channel: Instagram and WhatsApp.
-
Resolution in the thread. Customers rate the conversation down when they are redirected to email, a portal, or a phone line. A chat that tracks the order, processes the return, or takes the COD order in place earns the top box.
-
Language fit. In MENA, replies in the customer's own dialect (Egyptian, Gulf, Levantine), matching their script (Arabic, English, or Arabizi) and gender, read as respect. Formal MSA boilerplate reads as a machine and drags scores; see Why MSA-Only Chatbots Fail Gulf Customers.
-
Clean handovers. When AI passes a hard case to a human, the customer should never repeat themselves. Re-asking for the order number is a reliable CSAT killer.
Should you measure AI conversations separately from human ones?
Yes, and per intent. Segment CSAT by:
-
Handled by AI vs handled by human vs AI-then-human: tells you whether automation is helping or hurting, and where the handover boundary should sit.
-
Intent (WISMO, sizing, returns, complaints): complaints will always score lower; do not let them mask a WISMO flow that quietly degraded.
-
Language/dialect: if Arabic-language conversations score below English ones, your localisation is the problem, not your policies.
A healthy pattern for automated support: AI-handled CSAT at or above human-handled CSAT on repetitive intents (because instant and correct beats slow and correct), with humans scoring higher on complaints, where empathy and discretion matter.
Where ReplAi fits
ReplAi handles the CSAT drivers directly: sub-14 second average replies 24/7, resolution inside the chat via Shopify actions (order tracking, returns, COD orders, catalogue search), and native dialect replies in Egyptian, Gulf, and Levantine Arabic, English, or Arabizi with gender-aware phrasing. Across our merchant cohort, 80% of messages are fully automated (Q1 2026), and the anecdotal ceiling is the Blaze Sportswear owner's line: "Customers don't even realise they're talking to AI until we tell them."
If you want to see what a top-box chat interaction feels like in Arabic, book a demo; plans including a free tier are on the pricing page.
Frequently asked questions
What is a good CSAT score for chat support?
Published figures vary so much by industry, question format, and response rate that a universal target is not honest. A practical approach: establish your own baseline over a month, then hold automation to the standard of matching or beating your human baseline on repetitive intents, and investigate any intent segment trending down two weeks in a row.
CSAT, NPS, CES: which should a small store track?
For support quality, CSAT first: it is transactional, cheap to collect in chat, and maps directly to fixable drivers (speed, resolution, language). NPS measures brand-level loyalty and belongs to a broader cadence; CES (effort score) is a useful complement if you suspect customers resolve issues but resent the friction.
Do customers actually rate AI conversations?
Yes, at comparable or higher rates, because the survey arrives seconds after an instant resolution while the customer is still present. The interesting finding across automated support is that customers rate outcomes, not the nature of the agent; fast, correct, native-language resolutions score well regardless of who typed them.
How many responses do I need before the score means anything?
Treat CSAT like any sampled metric: at low volume, single ratings swing the percentage wildly. As a rule of thumb, read weekly CSAT only once you collect a few dozen responses per segment you care about, and rely on 4-week rolling trends for decisions rather than day-to-day movement.