ChatGPT vs Claude vs Gemini: One of Them Invented a Refund Offer
I gave ChatGPT, Claude and Gemini the same scruffy internal note to turn into a customer email, with six hard rules attached. One of the rules was do not add any detail that is not in the paragraph below. All three came back polite, all three came back at the right length, and Gemini offered the customer a prepaid return label.
There is no prepaid return label in the source note. Gemini invented a shipping cost and committed the business to it inside a message that passed every other check I could automate. That is the difference between these three tools that I would actually want to know about, and it is not the one the comparisons usually lead with. Below: two tests I ran in September 2026, every output, and what I reach for now.
The short version
All three of these tools — ChatGPT (from OpenAI), Claude (from Anthropic), and Gemini (from Google) — can draft, rewrite, shorten, and proofread text well enough that most people won't be able to tell a machine helped. If you only ever use one, you'll be fine. The differences show up at the edges: how the writing sounds, how they fit into tools you already use, and how they handle longer or messier jobs.
| ChatGPT | Claude | Gemini | |
|---|---|---|---|
| Sounded most human in my test | Close second (“That timing would work much better for me” is filler) | Yes — the only one that didn't pad | Stiffest; turned “10” into “10:00 AM” unasked |
| Where it lives | Standalone app | Standalone app | Inside Gmail and Docs |
| Reach for it when | You want one dependable all-rounder | The wording itself matters | You write mostly in Gmail/Docs |
| Free plan | Yes | Yes | Yes |
| Invented a detail in my tests | No (2 runs) | No (2 runs) | Yes — offered a prepaid return label |
The same prompt in all three, September 2026
Run on a ChatGPT Plus account, a Claude Max account and Gemini's free tier, each with its default model selected, September 2026.
One prompt, sent to each tool the same afternoon: a three-sentence reply declining a Thursday meeting, suggesting Tuesday at 10, friendly, not apologetic, no exclamation marks.
Claude's is the one I'd send: it reads like a person and it is the only one that didn't pad. ChatGPT's is close, but “That timing would work much better for me” is filler. Gemini's is stiff (“I will be fully focused on meeting an upcoming deadline”) and it turned “10” into “10:00 AM” on its own — helpful if that is what I meant, wrong if it wasn't.
Second test: six rules, and the one that actually matters
The meeting reply above was about voice. This one is about obedience, because that is the part you cannot eyeball. I gave all three the same scruffy internal note and six hard rules, in a fresh conversation each, twice, on 22 September 2026:
Rewrite the paragraph below as a message to the customer. Hard rules: (1) Exactly 60 words. (2) No exclamation marks. (3) Do not use the words apologize, apologies, sorry, unfortunately, or delighted. (4) Keep the refund amount and the date exactly as written. (5) Do not add any detail that is not in the paragraph below. (6) End with a question.
Paragraph: ok so the customer wants their money back on the mixer, they paid 84.50 on 3 March, it’s past our 30 day window but they say it arrived broken and they only opened the box last week. we can refund but i want them to send it back first.
Five of those six rules are machine-checkable, so I checked them with a script rather than by eye.
| Exactly 60 words (run 1) | 60 ✓ | 60 ✓ | 60 ✓ | ||||||||||||||||||||||||
| Exactly 60 words (run 2) | 59 ✗ | 60 ✓ | 59 ✗ | ||||||||||||||||||||||||
| Avoided all five banned words | ✓ | ✓ | ✓ | ||||||||||||||||||||||||
| Kept 84.50 and 3 March unchanged | ✓ | ✓ | ✓ | ||||||||||||||||||||||||
| No exclamation marks | ✓ | ✓ | ✓ | ||||||||||||||||||||||||
| Ended with a question | ✓ | ✓ | ✓ | ||||||||||||||||||||||||
| Added nothing that was not in the source | ✓ | ✓ | ✗ run 1 |
Run on a ChatGPT Plus account, a Claude Max account and Gemini’s free tier (Flash), each on its default model, 22 September 2026. Two runs each, fresh conversation every time.
What the easy rules showed
Nothing, which is itself worth knowing. Across all six runs, every model avoided all five banned words, kept 84.50 and 3 March untouched, used no exclamation marks, and ended with a question. If your constraints are of that kind — don’t say X, keep Y exactly, end with Z — all three will do as they are told.
Exact word counts are not reliable, in any of them
Four of six runs landed on exactly 60. Two came in at 59: ChatGPT on its second run, Gemini on its second run. Same prompt, same account, minutes apart. Claude hit 60 both times, which on two runs is not enough to call it better — it is enough to say that none of the three can be trusted to count, and that if a word limit genuinely matters you have to check it yourself. Nothing in any of the three replies flagged that it had come up short.
The failure that no rule check catches
Rule five was the one that mattered, and it is the one you cannot script. On its first run, Gemini ended with this:
“Would you like a prepaid return label?”
There is no prepaid return label anywhere in the source note. The note says the opposite of generous — i want them to send it back first. Gemini invented a shipping cost and offered it to the customer on the company’s behalf, inside a message that was exactly 60 words, perfectly polite, and passed every other check I could automate. On its second run it asked “Would you like instructions on how to return it?” instead, which is fine. So it is not that Gemini always does this. It is that it did it once in two tries, and the version that did it looked exactly as correct as the version that didn’t.
ChatGPT and Claude added nothing across either run. The closest either came was ChatGPT opening with “Thank you for contacting us,” which the note does not state but does imply.
What I would take from this
The thing to check before you send an AI-drafted reply is not the tone and not the length. It is the promises. Scan for anything the draft offers on your behalf — a refund window, free shipping, a callback, a deadline — and confirm you actually said it. That is a ten-second check, it is the one failure here that would have cost real money, and it is the one the polish is most likely to hide.
ChatGPT: the safe default
In both tests ChatGPT did what was asked without drama. Its meeting reply was close to sendable — “That timing would work much better for me” is filler I would cut — and in the rules test it followed every instruction and invented nothing, missing only the exact word count on the second run.
What that reflects in ordinary use: it is forgiving of a vague request and it rarely leaves you with nothing. The cost is that its default voice leans eager and padded, so a first draft usually needs a line or two cut. That is a two-second fix, which is why it is a reasonable choice if you are only going to learn one of these.
Claude: the one I edited least
Claude’s meeting reply was the only one of the three I would have sent unchanged, and it was the only model that hit exactly 60 words on both runs of the rules test. It also added nothing to the refund message either time.
Two runs is not a benchmark and I am not claiming it counts better than the others. What I can say is narrower and still useful: across the tasks I ran, it padded least and stuck closest to what I gave it. The trade-off I hit is that it is more cautious — it will sometimes add a gentle caveat you did not ask for — and it is not wired into Gmail or Docs the way Gemini is.
Gemini: convenient, and the one to reread
Gemini is the only one of the three sitting inside Gmail, Docs and Sheets, and if you write most of your email there that convenience is real. It is also the only one that invented something in my tests: the prepaid return label above, and, in the earlier meeting test, turning “10” into “10:00 AM” without being asked. Both are small. Both are it deciding a detail on your behalf.
Its prose was the stiffest of the three — “I will be fully focused on meeting an upcoming deadline” — though that is a matter of taste and easily fixed. The invention is not a matter of taste. If you use Gemini for anything a customer or a colleague will act on, read it once for claims you did not make.
Four cautions that apply to all three
A few honest cautions that apply to all three, because no tool here is magic:
- They make things up. All of these can state something false with total confidence, especially names, dates, statistics, and quotes. Never let one write anything factual — a bio, a product spec, a price — without checking it yourself.
- Don't paste anything private. Customer data, passwords, medical details, contracts — assume anything you type could be seen or used to improve the service unless you've checked the privacy settings for your plan. When in doubt, leave it out.
- The default voice can be a giveaway. Readers are getting good at spotting AI writing. Always read the draft out loud and cut the filler; the tool gives you a first draft, not a finished one.
- Features and prices change constantly. Each has a free tier and paid plans, and the exact limits and costs shift often. Check the current details on each company's own site before you pay for anything — don't trust a number you read in an article, including this one.
Which one I reach for
Claude when the wording is the point and I do not want to edit — a message to a customer, a newsletter, anything where the tone has to land. It padded least across both tests.
Gemini when I am already inside Gmail or Docs and the writing is routine, or when the draft needs a current fact. With one habit attached: reread it for anything it has promised on my behalf.
ChatGPT when I want one tool that handles everything and I am not going to think about it. It is the safest default and the least likely to surprise me in either direction.
The honest caveat on all of that: this is two tasks and two runs each, on one afternoon, on whatever version each service was serving me. It is enough to show that a confident, correctly-formatted draft can still contain a promise you never made. It is not enough to rank three models, and you should distrust anyone who tells you otherwise from a sample this size. All three have free tiers, so the test that settles it for you is running your own real task through each of them and noticing which drafts you keep editing.
Related reading
- How to Write Better AI Prompts: A Simple Framework for Non-Techies — Role, Task, Context, Format — the four parts of a prompt that works.
- Can ChatGPT Read a Screenshot or a Photo of a Document? — what AI gets right and wrong reading a photo of a page.
- How to Use AI to Write and Tailor Your Resume and Cover Letter — tailoring to a posting without sounding like a chatbot.
Keep reading
Written by Mitch, a software analyst who tests software for a living. Every guide here comes from actually using the tool on a real task — including the parts where it falls over. Tested on ChatGPT Plus, Claude Max and Gemini (free). More about this site · Corrections: [email protected].
Get the next guide by email
A new tested how-to when there is one — usually a couple a month. No hype, unsubscribe any time.
Tested in real accounts. No affiliate links.