Methodology

How a verdict gets made

Short version: we judge each task against the same checklist, write the prompt we'd use ourselves, and say plainly where AI falls short.

The checklist

Every task is judged on four questions:

  1. Can a current AI app do the core of the job? ChatGPT, Claude and Gemini, including free tiers.
  2. How often does it get it wrong? Made-up facts, bad math and confident mistakes count against it.
  3. What does a mistake cost? A clunky packing list is fine. A wrong tax form or a missed legal deadline isn't.
  4. Can a regular person check the result? If you can't easily spot the error, the verdict gets stricter.

What the verdicts mean

Bot AI does the job well and mistakes are easy to catch.

Half-bot AI does most of the work, but a person has to finish, check or decide something.

Nah AI is wrong too often, or the stakes are too high. We still show how AI can help around the edges.

Why pages say “Reviewed”, not “Tested”

Verdicts are written by the Bot or Nah team with AI assistance, following the checklist above. We haven't yet logged a formal test run for every task, so we don't claim we have. Each page shows the date it was last reviewed. As we add logged test runs, those pages will say so.

Readers keep us honest

Every verdict has a “Did this work for you?” vote. We only show the tally once enough people have voted for it to mean something. If readers report a different result, we re-review the verdict.

Money never changes a verdict

Sponsors can buy labeled placements. They can't buy, change or remove a verdict, and a sponsor's own product gets the same checklist as everything else.

Not professional advice

Verdicts about money, law and health are general information, not advice. For anything with real stakes, check with a qualified professional.