How we tested this category
Nine tasks, run in the same week, on paid accounts at list price. No demo accounts and no vendor walkthroughs. Each task is scored on its own before any tool meets another, and the weights come from the rubric.
Output quality, 35 points
1. The long brief, 15 points. A 1,500 word article from a real brief: topic, audience, angle, a fixed outline, tone notes and three points that must appear. We score whether it holds the outline, whether the argument survives to the end, and we time how long it takes to edit to publishable.
2. The rewrite, 10 points. A supplied 400 word technical passage, rewritten for a non-expert without losing accuracy. This tests comprehension rather than generation, and it separates the tools badly.
3. The fabrication check, 10 points. A 500 word piece on a deliberately niche topic, with five specific factual claims and sources. We verify every claim and open every source. Any invented source or made up statistic scores this task zero, whatever the prose was like.
Reliability, 15 points
4. Three runs, same brief. Task one, repeated three times in fresh sessions. We score the worst of the three, not the best. A tool that is brilliant one run in three is a lottery, and you cannot plan work around a lottery.
Control, 15 points
5. Brand voice, 8 points. We supply a 500 word writing sample and a five rule style guide, then ask for 300 words in that voice. We count rule violations.
6. Hard constraints, 7 points. Exactly 250 words give or take ten, three supplied phrases used verbatim, five banned words avoided, a required heading structure. Scored as constraints met out of the total, because a tool that ignores instructions is not a tool, it is a suggestion.
Value, 15 points
7. Cost per finished piece. Using the edit times from tasks one to six and the real plan price at a realistic monthly volume, we work out what one publishable 1,500 word piece actually costs. We also check what happens at the tier above, because that is where the bill usually doubles.
Workflow, 10 points
8. Getting it out. Export into WordPress and Google Docs with headings, links and formatting intact, plus any integrations the tool advertises. We count the manual fixes needed on the other side.
Support, 10 points
9. One real question. We contact support with a genuine, specific question during every test. We record the time to a human reply, whether it answered the question, and whether it was a bot loop pretending to be help.
Which one is yours
Rankings are useful up to a point, then the honest answer is that different people should buy different things. Once the scores above are live, this section names the pick for each case below.
- You publish weekly and edit everything anyway. Speed and export quality matter more to you than polish, because you were going to rewrite it regardless.
- You need one voice held across a team. Brand voice controls and constraint following are the whole game. Pay more for them.
- You write about regulated or technical subjects. Task three is your task. A tool that invents a source once will do it again on the piece you did not check.
- You write occasionally. Most of these are priced for people producing constantly. If you write twice a month, the honest answer is usually that none of them are worth the subscription.
We would rather tell you to buy nothing than sell you a subscription you will abandon in six weeks. It happens more often than the rest of this industry admits.