Which AI is best for your work? Run your own fair test
Leaderboards test models on someone else’s questions. A fair test on your own tasks shows which assistant suits your work. Here’s how to set one up and score it.
In 30 seconds
- Pick 5 to 10 real tasks, and give every tool exactly the same prompt and files.
- Hide which tool wrote which answer, and score against criteria you set before you start.
- Run each task more than once, weigh cost and privacy too, and retest when models change.
It’s easy to stick with the first assistant you tried. That’s fine for quick questions. But if you rely on AI for real work, or you’re choosing between ChatGPT, Claude and Gemini for a team, a short, fair test on your own tasks tells you more than any review, leaderboard or demo.
Why your own test beats a leaderboard
Published benchmarks measure Jargon busterModel: What training produces: a very large set of numbers that captures patterns, used to make predictions. on standard questions, not your reports, emails and spreadsheets. And AI’s strengths are uneven, a pattern researchers call the Jargon busterJagged frontier: Researchers’ name for AI’s uneven abilities: excellent at some tasks, surprisingly poor at others that look similar.: the model that tops a coding chart may be ordinary at the board papers you write every month.
Answers also vary from run to run. Ask the same thing twice and you can get two different answers, and Anthropic notes that even developers using identical settings can get different outputs. One try proves very little.
Set up a fair test
- Choose 5 to 10 real tasksUse work you actually do, such as a summary, a tricky email or a data question. Include one that AI has got wrong before.
- Fix the promptWrite one Jargon busterPrompt: What you type or say to an AI to tell it what you want. per task, and use it with the same files in every tool. Test the plan you’d actually pay for, as free versions can use different models.
- Set the criteria firstDecide what a good answer must include before you read any, so the first impressive answer doesn’t set the bar.
- Run each task twiceRun every prompt at least twice in each tool, in a fresh or temporary chat, so earlier conversations don’t colour it and one lucky answer doesn’t decide it.
- Hide the namesPaste the answers as plain text into one document labelled A, B and C, and ask a colleague to shuffle them, making it a Jargon busterBlind test: A comparison where whoever judges doesn’t know which option is which, so expectations and brand loyalty can’t sway the scores..
- Score, then revealMark every answer against your criteria in a simple spreadsheet, then unmask the tools and add up the totals.
Score what matters
| Criterion | Ask yourself |
|---|---|
| Accuracy | Are the facts, figures and names right? Check them against your source. |
| Usefulness | Does it answer the real question, at the right length and level? |
| Edits needed | How long would it take you to make it good enough to send? |
| Tone | Does it sound like your organisation, and suit the reader? |
Score each from 1 to 5, and give accuracy extra weight if mistakes are costly. Note anything that would have caused real trouble: one invented figure can matter more than ten well-written paragraphs.
I’m going to test AI assistants on my work as a [your role]. My tasks are: [list 5 to 10]. For each task, write one clear prompt I can use unchanged in every tool, and list what a good answer must include. Then make a score sheet as a table with columns for task, answer label (A, B or C), run number, scores from 1 to 5 for accuracy, usefulness, edits needed and tone, and notes.
Look beyond the scores
- Speed. How long did each one take, and does that matter for this job?
- Cost and limits. At your volume, would you keep hitting a Jargon busterUsage limit: A cap on how much you can use an AI tool, such as messages or uploads, in a set period. Once you reach it, you wait for it to reset.? Price the plan for everyone who’ll use it.
- Privacy terms. Does the plan you’d buy train on your data, and what controls come with it? See how to vet an AI tool.
- Fit. Does it work with the apps you already use, such as your email, documents or spreadsheets?
If two tools finish close, pick the one with the better privacy terms and controls. Whichever wins, the habits in how to write a good prompt will get more out of it.
Check yourself
3 quick questions nothing is savedTools in this guide
Sources (3)
- GlossaryAnthropic
- Define success criteria and build evaluationsAnthropic
- Model deprecationsAnthropic
Spotted a mistake? Tell us and an editor will check it.