Everything is marked out of 100 against the same six things, weighted the same way in every category. A 74 here means what a 74 means anywhere else on the site.
The rubric
| What | Weight | What we are really asking |
|---|---|---|
| Output quality | 35 | Is the result any good, judged blind where the category allows. Biggest weight, because it is the thing you are handing over money for. |
| Reliability | 15 | Does it do that again. A tool that is brilliant one run in three is not brilliant, it is a lottery. |
| Control | 15 | Can you steer it, or do you take what it gives you and fix it yourself afterwards. |
| Value | 15 | Quality against what it actually costs at real usage, including the moment you outgrow the starter tier and the bill doubles. |
| Workflow | 10 | Speed, exports, integrations. How much friction sits between the output and finished work. |
| Support | 10 | Is anyone there when it breaks. We contact support during every single test and write down what happened. |
What the number means
| Score | Translation |
|---|---|
| 75 to 100 | We would spend our own money on this. |
| 55 to 74 | Real tool, real limits. The review says exactly who should walk away. |
| Below 55 | We would not pay for it, and the review says why. |
Every score badge on this site is coloured straight from those thresholds. Green is always 75 or better. Red is always under 55. Nobody can override the colour, including us.
Three things worth exactly nothing
What the tool pays us. The person scoring does not know the rate. Commercial terms get attached after the score is locked.
Whether they paid to be tested early. Changes when we test. Never changes what we find.
How famous it is. A household name gets no benefit of the doubt. A tool nobody has heard of gets no sympathy.
Building the number
Each criterion is scored alone, before any tool is set against another. Then it is weighted and totalled.
Fail a task and you get zero for it, not a kind mark for effort. Half marks exist for genuinely partial results.
Scores go stale
A score is one version of one tool on one date, and both are printed on the review. We re-test on significant changes, on real price moves, and once a score passes a year old. When a number changes we say what moved it. The old one stays visible rather than quietly vanishing.
If you think a score is wrong
Wrong fact, tell us, we fix it and note the fix. Different conclusion, fair enough. The rubric and the task list are published precisely so you can find the point where your judgement and ours part company.