Every AI SEO content generator can produce 2,000 words on your category in under a minute. That stopped being the hard part some time ago. The hard part is telling, without reading all 2,000 words yourself, whether this particular draft is an asset or a liability, because the two look identical from a distance. Both have headings. Both are the right length. One of them contains a statistic that does not exist, and it is going out under your name. This piece is about what to look for in the output, and about the eleven rules we made blocking, because regenerating a draft costs a minute and apologising for one costs a customer.
What an AI SEO content generator actually gets right and wrong
Start with what is reliably good, because pretending otherwise wastes your attention. Structure is good: a model asked for an article with an H1, six H2s and a conclusion will deliver exactly that. Fluency is good. Coverage of the obvious subtopics is good. On-page mechanics are good, because they are rules and rules are easy.
What is unreliable is anything the model does not know and cannot check: your prices, your product names, your limits, your customers, and any number attributed to a source. Language models are trained to produce plausible text, and a plausible-sounding statistic is exactly as easy to produce as a true one. There is no internal signal that separates them.
So the shape of a good generator is not a better writer. It is a writer plus something that reads the output afterwards and knows what it is allowed to leave the building.
The four ways a draft becomes a liability
The invented statistic. A sentence beginning "according to" and ending in a precise percentage, with nothing behind it. This is the expensive one. It survives review because it reads like the most authoritative sentence in the article, and it is the one a customer will quote back at you.
The empty section. A heading that was planned and never written, or written in eleven words. Readers spot this in seconds, and it makes the entire piece look unattended.
The article about nothing. A piece that hits every structural target, repeats the target phrase in the title, the opening and an H2, and never once mentions what you actually sell. It scores well on any tool that measures form alone. It is worth nothing to a reader.
The stuffed article. The keyword repeated past the point where a person would notice. Once density crosses roughly two and a half percent of total words, it reads as manipulation to a reader and to a ranker.
Each of those has a checkable signature. That is the whole basis of the fifteen rules below.
The fifteen checks, grouped by what they protect
Structure: is this an article at all?
Word count has to land between 1,500 and 3,000. Below 1,500 there is not enough substance to compete for a commercial phrase; above 3,000 the piece is usually two articles. Heading hierarchy has to be exactly one H1, first, with no skipped levels on the way down, because an H4 under an H2 is a structural lie to a screen reader and a crawler alike. Both are blocking.
Placement: will a searcher and a crawler know what this is?
The keyword has to appear in the title, in the first 100 words of prose, and in at least one H2. Those three are where a searcher decides whether to stay and where a crawler decides what the page is about. Density has to stay at or below 2.5% of total words, counted as whole-phrase words so a three-word phrase is not treated as three times safer than a one-word phrase. All four are blocking.
Safety: could this sentence cost you something?
Your banned-phrase list is a hard rule. Whatever your legal team, your founder or your style guide will not allow, a draft containing it fails and does not publish.
The fabricated-specifics rule flags any sentence with the shape of a sourced claim: a percentage sitting alongside attribution language, an academic-style citation with a year, a reference to a dated study, or a quotation attributed to a person. A bare number is not flagged, because "takes fifteen minutes" is normal prose. The rule cannot verify a claim, so it does the next best thing and refuses the shape of a claim that would need verifying. It is blocking, and it is the rule we would keep if we could keep only one.
Substance: is this about your business?
Section substance fails any H2 whose prose, including its subsections, falls below 25 words. Subject coverage strips out every occurrence of the target keyword and then asks whether the article ever mentions what you sell: your category, your product names, the vocabulary a piece about your business is bound to use. An article that passes every other rule and fails this one is filler, and it fails hard.
Internal links have to number between two and five, pointing at your own article-like pages rather than at your terms page. Both are blocking when your site has enough pages to make them satisfiable.
Presentation: the advisory four
Title length over 60 characters gets truncated in results. A meta description outside 130 to 160 characters either wastes the snippet or gets cut. A reading grade above 12 usually means the sentences are too long. And when a title promises something specific, a price or a comparison, the body has to deliver a figure or name two things to compare.
These four lower the score and do not block, because each of them has a legitimate exception. A dense trade publication can write above grade 12 on purpose. A rule that misfires and eats a customer's article is worse than a nit that shows up on the scorecard.
What happens when a check fails
This is the part to ask any vendor about, because "we check quality" means nothing without it.
In Rankli, a hard failure regenerates the draft once, automatically, with the failure fed back in. If the second attempt passes, it publishes and the history keeps both versions. If it fails again, it stops and the reason is written on the article in words you can act on, not a red dot.
There is also a project setting: hold drafts that fail checks. Switch it on and nothing that fails a hard rule ever leaves the app, whatever the schedule says. For a regulated client that setting is the whole reason the pipeline is usable at all.
Where you want to change something the checker did not catch, edits are proposed rather than applied: each one is a diff you accept or reject, and the score re-runs as you accept. How that works is written up in the improvements documentation.
What no checker can do
Be clear about the limit, because a rule set that claims too much is its own kind of failure.
Nothing here judges whether the argument is any good. Nothing here knows whether your claim about your own product is true. Nothing here can tell a boring article from an interesting one. Those need a person, and the point of automating the other fifteen judgements is that the person spends their attention on the ones that need it.
What the rules do is remove the failures that are mechanical, frequent and expensive. That is a smaller claim than "AI writes your blog", and it is the one that survives contact with a month of output.
Reading a scorecard in thirty seconds
Once a generator hands you a report as well as a draft, the report is the faster read. Three things tell you almost everything.
Look at the hard failures first, and only at those. A draft with none of them is publishable in the mechanical sense, whatever the headline score says. A draft with one is a regeneration, not an edit.
Then look at the two substance rules. Subject coverage tells you whether the piece is about your business at all. Section substance tells you whether any heading was abandoned. Those two catch the defects that make an article embarrassing rather than merely imperfect.
Then look at the advisory list and decide which ones you actually care about. Plenty of good articles run long in the title or dense in the sentences. The score is a summary, not a verdict, and the binding fact is always the hard list.
What you should not do is compare headline scores across tools that use different rule sets. A 94 from one vendor and an 88 from another are not on the same scale, and neither number means anything until you can read what the rules were.
How to test a generator in twenty minutes
Take a keyword you know well. Generate one article in each tool you are considering. Then, before reading either, run both through a scorer and compare the reports rather than the prose.
Ask four questions of each draft. Does it name your products anywhere outside the keyword phrase? Does every heading have real prose under it? Is there a single number in it you could not source? And would you have published it without editing?
You can run exactly the fifteen checks described here on any draft, from any tool, free, in the SEO article checker. If a generator's output cannot pass a published rule set, that is worth knowing in twenty minutes rather than in month three. Our comparison pages put the same questions to each product in the category.
Rankli's $1 trial is three finished articles in your own voice, published live to your own site, each with its full scorecard attached. Fourteen days, cancel from the app any day, unused articles refunded. The plans are on the pricing page.