ChatGPT vs Claude: How to Choose for Your Own Work
Skip the benchmark noise. Run a five-task, blind test in about an hour to decide between ChatGPT and Claude for your own writing, coding and research.
In this article
Search for "ChatGPT vs Claude" and you will find a pile of confident winners. Most of them are out of date within a quarter, because both products ship new models and change their plans constantly. And even the careful ones answer the wrong question. Which assistant is better in general matters far less than which one is better at the three or four things you do every week.
So this guide does not crown a winner. It gives you a short, repeatable test that settles the question for your own work in about an hour, plus a script that does the tedious part for you. You will finish with evidence you can trust and re-run whenever a new model appears.
#What both tools have in common
It helps to start by removing the differences that don't exist. ChatGPT and Claude are both:
- Chat apps with a free tier and paid plans that raise the usage limits.
- Available through an API, so you can call them from your own code.
- Able to read documents and images you upload, write and explain code, and call external tools.
- Updated often enough that any feature list in a blog post is probably stale.
Because of that overlap, feature checklists rarely decide anything. What decides it is output quality on your tasks, how the tool fits your workflow, and what it costs at the volume you will actually use. We keep short, neutral summaries on the ChatGPT and Claude profile pages, but treat them as a starting point, not a verdict.
#Why most comparisons mislead you
Three problems show up again and again:
- They test generic prompts. "Write a poem about the ocean" tells you almost nothing about how a tool handles your messy customer emails or your 400-line legacy function.
- They run each prompt once. These models vary from run to run. One good or bad answer is an anecdote.
- The reviewer knows which answer came from which tool. Brand expectations leak into scoring even when people try to be fair.
The method below fixes all three: your tasks, repeated runs, and blind scoring.
#The five-task bake-off
#Step 1: Pick five real tasks
Choose work you actually repeat. Aim for a mix, because tools often trade strengths across task types. A good set:
| Task type | Example from real work |
|---|---|
| Writing | Rewrite a rough customer email so it is clear and kind |
| Summarising | Turn a long meeting transcript into decisions and action items |
| Coding | Explain and refactor a function you did not write |
| Data extraction | Pull names, dates and amounts from messy text into JSON |
| Reasoning | Compare two options and recommend one with trade-offs |
Write each as a complete prompt, with the real context included. Remove anything confidential first. If a prompt needs a long document, paste a shortened or anonymised version.
#Step 2: Run each prompt several times on both tools
You can do this by hand in the two chat apps, but you will get cleaner data with a script because it can randomise which answer is "A" and which is "B". The script below calls both APIs, runs every prompt three times, and writes the answers to a CSV in shuffled order. A separate file keeps the answer key.
You need an API key for each service, set as ANTHROPIC_API_KEY and OPENAI_API_KEY, and pip install anthropic openai.
import csv
import random
import anthropic
from openai import OpenAI
CLAUDE_MODEL = "claude-sonnet-5-5"
OPENAI_MODEL = "your-openai-model" # look up the current model name on OpenAI's models page
RUNS = 3
PROMPTS = {
"writing": "Rewrite this so it is clear and kind: 'Your refund is late because of our process. Wait.'",
"summary": "Summarise into decisions and action items:\n\n<paste anonymised transcript here>",
"coding": "Explain what this function does and refactor it:\n\n<paste function here>",
"extraction": "Return JSON with name, date and amount from: 'Paid $120 to Dana Reyes on 3 March.'",
"reasoning": "Option A costs less but takes 3 months. Option B costs 40% more and takes 1 month. Recommend one and explain.",
}
claude = anthropic.Anthropic()
openai = OpenAI()
def ask_claude(prompt):
r = claude.messages.create(
model=CLAUDE_MODEL, max_tokens=1200,
messages=[{"role": "user", "content": prompt}],
)
return r.content[0].text, r.usage.input_tokens, r.usage.output_tokens
def ask_openai(prompt):
r = openai.chat.completions.create(
model=OPENAI_MODEL,
messages=[{"role": "user", "content": prompt}],
)
return r.choices[0].message.content, r.usage.prompt_tokens, r.usage.completion_tokens
rows, key = [], []
for task, prompt in PROMPTS.items():
for run in range(1, RUNS + 1):
answers = {"claude": ask_claude(prompt), "chatgpt": ask_openai(prompt)}
order = list(answers)
random.shuffle(order) # hide which tool is A and which is B
item_id = f"{task}-{run}"
rows.append({
"id": item_id, "task": task, "prompt": prompt,
"answer_A": answers[order[0]][0], "answer_B": answers[order[1]][0],
"score_A": "", "score_B": "",
})
key.append({
"id": item_id, "A_is": order[0], "B_is": order[1],
"claude_in": answers["claude"][1], "claude_out": answers["claude"][2],
"chatgpt_in": answers["chatgpt"][1], "chatgpt_out": answers["chatgpt"][2],
})
with open("bakeoff_blind.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=rows[0].keys())
w.writeheader()
w.writerows(rows)
with open("bakeoff_key.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=key[0].keys())
w.writeheader()
w.writerows(key)
print(f"Wrote {len(rows)} blind items. Score bakeoff_blind.csv, then open bakeoff_key.csv.")#Step 3: Score blind
Open bakeoff_blind.csv and do not open bakeoff_key.csv yet. For each row, score answer_A and answer_B from 1 to 5 against the same few criteria:
- Correct: are the facts and the code right?
- Follows instructions: did it respect the format, length and tone you asked for?
- Useful as-is: how much would you edit before using it?
Keep the scale simple and write one line of notes when a score is 1 or 5. If you can, ask a colleague to score the same file separately; where you disagree is where your criteria need to be clearer.
#Step 4: Open the key and add up
Once every row has scores, join the two files. A few lines of Python will average the scores per tool and per task:
import csv
from collections import defaultdict
key = {r["id"]: r for r in csv.DictReader(open("bakeoff_key.csv", encoding="utf-8"))}
totals = defaultdict(list)
for row in csv.DictReader(open("bakeoff_blind.csv", encoding="utf-8")):
k = key[row["id"]]
for slot in ("A", "B"):
if row[f"score_{slot}"]:
totals[(row["task"], k[f"{slot}_is"])].append(int(row[f"score_{slot}"]))
for (task, tool), scores in sorted(totals.items()):
print(f"{task:12} {tool:8} mean {sum(scores) / len(scores):.2f} (n={len(scores)})")Look at the pattern by task, not just the grand total. A tool that wins four of five tasks but loses the one you do all day is the wrong choice for you.
#Step 5: Compare the cost at your real volume
Quality is half of the answer. The other half is what each tool costs for the amount you will use. The key file already records input and output tokens for every call. Multiply them by each provider's current per-token prices (check their pricing pages the same day, because they change) and scale up to your expected monthly volume.
Remember that subscriptions and API billing are separate products. If you mostly chat in the app, compare plan prices and usage limits. If you are building a product, compare token costs. Don't mix the two.
#Questions to ask beyond quality
Once you have scores, a handful of practical questions often break a tie:
- Where does your team already work? An assistant with a good integration for your editor, docs or chat tool saves more time than a slightly better answer in a separate tab.
- How long are your documents? If you routinely paste hundreds of pages, test that case specifically instead of assuming.
- How strict are your output formats? If a downstream script parses the reply, measure how often each tool returns valid, complete JSON.
- What are your data rules? Check each provider's current data-retention and training policies for your plan before you paste anything sensitive.
- Do you need to switch later? Avoid building your whole product around one vendor's special feature unless it gives you a real edge.
#You might not have to choose
Plenty of people use both. A simple pattern is to keep prompts portable and route by task: send one kind of work to the tool that scored best on it and everything else to the cheaper option. If you build with the APIs, put the model call behind one function, like ask(prompt, provider), so changing provider is a config change rather than a rewrite. The Claude API beginner's guide shows the Claude side of that function.
#Frequently asked questions
Is Claude better than ChatGPT? Neither is better at everything, and the answer changes as models update. Test on your own tasks with the method above.
Which one is cheaper? It depends on the plan or model you choose and how much you use it. Compare current prices against your measured token usage.
Can I switch later? Yes, if you keep your prompts and data in your own files rather than inside one app. Export what matters and keep the bake-off script so you can re-run it on the next model.
How often should I re-test? Whenever either vendor releases a new model you care about, or once a quarter, whichever comes first. The script makes it a ten-minute job.
#The takeaway
Choosing an AI assistant is a small engineering decision, so treat it like one: define the tasks, measure, and keep the harness. An hour of testing beats weeks of reading opinion pieces, and you can repeat it whenever the landscape moves. For more ways to put these tools to work, see web scraping with Python and AI.
Get the next workflow by email
One tested workflow every week. No filler, unsubscribe in one click.