14–20 SEPTEMBER 2026 · MARKETING · SALES · CUSTOMER SERVICE
The agent finished. The customer did the extra work.
How to spot unnecessary customer effort, read a 3.2X sales claim before buying, and decide whose interests an AI recommendation should serve.
Hello, fellow laundry inspectors.
Imagine this support chat. It is an illustration, not a reported incident: the refund went through. The ticket closed. The dashboard turned green.
The customer also typed their order number three times and explained the problem twice.
The system recorded a successful refund. It did not record how much effort the customer spent getting it.
That gap is the thread running through this week’s wash. Completing a task, reporting more sales and making a recommendation can all look impressive. Each needs a different question before we call it useful.
Here are three stories to help you improve a customer journey, challenge a business case and make a better recommendation. Plus two short checks for the people handling customer data and agent instructions.
The refund went through. Why was it still hard?
There are two very different reasons a support agent might ask for your order number again. It may need to verify that it is changing the right order. Or it may have lost the information when it transferred you.
The first can protect you. The second makes you compensate for the system.
A new preprint, RideWay, makes a useful distinction between finishing a task and finishing it without unnecessary back-and-forth. It tests 24 models on 58 synthetic Chinese-language ride-hailing tasks, scoring successful completion alongside excess conversation and tool use.
Counting tool calls alone did not reliably predict which interaction would be preferred in the study. Including the conversation helped. In plain terms: knowing how often the software acted was not enough to judge the experience.
These were synthetic tasks, not live customer-service conversations. They do not establish that a shorter chat makes customers happier. They give us a useful question to bring to our own systems.
Did the customer get the refund?
Why did I have to do so much work to get it?
Try this with ten resolved conversations.
This is our suggested starting point, not a statistically representative study. For each conversation, note four things:
- Outcome: was the request actually resolved correctly?
- Repetition: what did the customer have to provide more than once?
- Reason: was that repetition needed for verification, or caused by missing information between systems?
- Next contact: did the customer come back about the same unresolved issue, within a follow-up period appropriate to that request?
If the same information keeps getting lost at a handoff, start there. For example, a human adviser could receive the verified order number, the problem already described and the action already attempted. Then review another comparable set of conversations after the change.
Do not reward the bot simply for asking fewer questions. Check that resolution stays correct and necessary identity checks remain intact.
The aim is not the shortest conversation. It is the least unnecessary work on the way to the right outcome.
Stain still visible: this is a recent preprint with synthetic tasks. We found no independent replication. Applying its insight to customer service is our proposed operating test, not a result established by the study.
Source: RideWay preprint
Before you put 3.2X in the business case.
HubSpot’s Fall Spotlight release reports 3.2X more won deals. That is the kind of number that can move from a launch announcement into a budget presentation without its footnote making the trip.
The comparison is between Pro and Enterprise customers using AI with high-quality context and Pro and Enterprise customers not using AI. The release’s note does not describe a randomised comparison.
Those are not interchangeable claims:
This group of customers won more deals.
Switch on this feature and your team will win 3.2X more deals.
Teams with better data might also have better sales processes, more resources or different customers. Without the comparison methodology, we cannot separate those possibilities from the effect of AI.
None of this makes the product unhelpful. HubSpot also announced self-updating CRM capabilities and a context completeness score. Reducing manual record maintenance is an understandable ambition. A complete field still needs to contain the right information.
What belongs in your business case instead?
Separate what you can measure from what you hope will happen. If the first use is updating CRM records after calls, estimate the time saved after corrections and review. Put a possible sales lift in a separate, explicitly unproven assumption.
For illustration, not a HubSpot result: a team saves 20 hours of manual entry in a month but spends eight hours checking and correcting it. The time benefit is 12 hours, not 20. The business case still needs subscription and usage costs, setup effort and a decision about what those 12 hours will be used for.
If the supplier’s sales multiplier is essential to getting the purchase approved, ask for the measurement period, group sizes and how differences between the groups were handled. If those answers are unavailable, do not treat the multiplier as your expected return.
The product may still be worth buying. It should not need an unsupported revenue promise to justify the first trial.
Stain still visible: this is a vendor-reported association, not independent evidence of a causal sales lift. The release does not provide enough cohort detail to resolve that question.
Source: HubSpot Fall Spotlight announcement and comparison note
Who is the recommendation working for?
Last week, we looked at AI shaping the buyer’s shortlist. This week’s question is more specific: what happens when the assistant represents the seller rather than the buyer?
In a controlled hotel-choice study, researchers changed the assistant’s assigned role: represent the traveller or represent the booking platform. Under the platform role, sponsored listings were penalised less in its choices. Replacing “Promoted” with the stronger “Sponsored” label did not eliminate that gap.
The hotel options were held fixed. The instructions about whom to represent changed.
This is a single-domain, single-turn study across multiple models, not evidence that a named live assistant is secretly steering purchases. Applying the finding to your own recommendation process is a question to investigate, not a proven result.
The useful decision is whose interests take priority.
Imagine an assistant recommending subscription plans. The customer needs two features. Both are in the cheaper plan. Your sales team would prefer the larger contract.
What should the assistant recommend? That is a commercial decision to make before writing instructions such as “be helpful and maximise conversion”. Those goals can conflict.
Would you still call it a good recommendation if the customer bought the cheaper option because it fitted their needs?
If the answer is yes, make that behaviour part of how you judge the assistant. A higher average order value alone would not tell you whether it was following the brief.
If you do not control the recommending system, you cannot assume you know its priorities. When an AI shortlist informs a purchase, ask what criteria produced it and check the important claims against the suppliers’ actual offers. Treat the assistant’s explanation as a starting point, not proof of how it reached the choice.
Source: Controlled study of assistant roles and sponsored recommendations
THE QUICK SPIN · 04–05
Two small checks worth keeping.
- 04The name was blacked out. Further down, it was still there.
An AWS redaction walkthrough found missed repetitions of personal information in narrative text. An additional matching step caught more of it, with a small increase in incorrect redactions. The test covered just 12 documents and 47 pages, so it is not a privacy guarantee.
Useful takeaway: test with a made-up customer whose name appears in a heading, a paragraph and a note. Check every occurrence in the exported file, not just the preview. This will not establish that the system catches everything, but it can expose a repeat-information failure before real customer data is involved.
- 05Your agent rewrote its instructions. Who checked the rewrite?
AWS describes using records of an agent’s past actions to improve its instructions. It also warns that chasing a better score can weaken rules. A support agent that closes more tickets by skipping refund approval is not an improvement.
Useful takeaway: when instructions change, test the rules that must stay fixed separately from the performance score. “Refuses refunds above its approval limit” should remain a pass-or-fail requirement, even if the new version is faster. This is a vendor engineering guide, not proof of a universal performance gain.
Which part of your automated process still makes customers repeat themselves?
Start with one repeated request: an order number, an explanation, a document already supplied. Find out whether it protects the customer or compensates for missing information between systems.
Keep the protection. Fix the unnecessary repetition. That is a useful result even if you do not add another AI feature this week.
AI Laundry is a weekly, evidence-led wash for marketing, sales and customer-service teams. Useful ideas, original sources and the limitations left visible. This edition draws on two new preprints and vendor disclosures. No model tests were rerun, and no comprehensive X coverage is claimed.