Skip to main content
← Blog

What a Book Readability Score Actually Measures

By SmartKDP

Every readability formula takes the same two inputs — sentence length in words per sentence and word length in syllables per word — and produces one readability score

Run a chapter through a readability checker and you get a grade level. It looks like a measurement of your book. It is not — and knowing precisely what it is turns a vanity number into one of the few genuinely useful checks you can run before publishing.

Not affiliated with Amazon. Nothing here is a KDP rule; KDP does not assess or require any reading level. This is about what these formulas do and how to use them.

The formulas measure two things

Every widely used readability formula — Flesch Reading Ease, Flesch–Kincaid Grade Level, Gunning Fog, SMOG, Coleman–Liau, the Automated Readability Index — is arithmetic over exactly two inputs:

  1. How long your sentences are (words per sentence)
  2. How long your words are (syllables per word, or characters per word)

That is the entire input. None of them reads vocabulary. None knows whether a word is common or obscure, whether a sentence is clear or nonsense, whether the subject is simple or hard. They are decades old — the oldest dates to 1948, the newest to the mid-1970s — and they were built for exactly this: a cheap proxy, computed by hand, over long samples of prose.

Two consequences follow immediately, and they are the ones a grade level hides.

A short sentence of rare words scores easy. "The quokka's pelage was xeric." Four short words and a full stop: a low grade level and an unreadable line for most people.

A long sentence of familiar words scores hard. Plain, warm, perfectly clear prose with a few subordinate clauses will be marked as demanding, because the formula counts the clauses and cannot see the warmth.

So a grade level is not a quality score and not a difficulty measurement. It is a description of your sentence and word rhythm, and that is a real thing worth knowing about your own writing.

Why there are six, and why they disagree

They were calibrated on different corpora for different readers — military manuals, health leaflets, school texts — and they weight the same two inputs differently. SMOG is stricter about polysyllables; Coleman–Liau uses characters rather than syllables, so it is unaffected by syllable-counting errors.

Reading all six together is more informative than trusting one, because their spread is itself the signal. Six numbers within a grade of each other means the text is consistent. A wide spread usually means the sample is short, or the prose is uneven — a dense opening paragraph followed by dialogue.

For the same reason, treat any single grade as approximate. Our checker reports grades with an explicit uncertainty of about a grade in each direction rather than a false-precision decimal.

Short samples produce nonsense

The formulas are averages, and averages over small samples are unstable. One long sentence in a forty-word excerpt moves the grade level by several years.

Our tool flags any sample under 100 words as thin and says so on the result, because a confident-looking grade printed over a paragraph is the most misleading thing a tool in this category can do. If you want a number that means something, paste a few thousand words from the actual body of the book — not the opening page, which is usually written more carefully than the rest.

The comparison that is actually useful

Here is the part almost nobody does, and the reason this tool has two modes.

Run your book and your Amazon description through the same pipeline and compare them.

A book description is not the book. It is a sales page a shopper skims on a phone, in competition with a dozen others. If your description reads several grades harder than your book, you have made the browsing experience harder than the reading experience — which is the wrong way round, because the description has to survive a few seconds of skimming while the book has the reader's commitment.

The mismatch is common and has an obvious cause: descriptions get written last, in marketing register, with long qualifying clauses stacked into single sentences. The prose in the book is looser and better.

That comparison is only meaningful if both numbers were produced identically — same formulas, same segmentation, same syllable counter. Comparing a grade from one site against a grade from another tells you about the two tools, not the two texts. The free Book Readability Checker runs both modes through the identical pipeline for exactly this reason, and it names each formula's original published reference on the page.

If the gap is large, the fix is in the description, not the book: shorten sentences, cut qualifiers, and put the promise in the first two lines — which is how to write a book description that converts in more detail.

What to do with the number

  • Children's and early-reader books. The one case where a target grade genuinely matters, because the audience is defined by reading level. Check against the age you are writing for.
  • Nonfiction and how-to. A high grade level is usually a symptom of stacked clauses, not of difficult ideas. Splitting sentences almost always improves it, and improves the prose too.
  • Fiction. Mostly ignore the absolute number. Use it comparatively — against your own earlier chapters, or against your description.
  • Any book. Watch the spread across the six formulas and any thin-sample warning before you take a number seriously.

One honest caveat about age bands specifically: converting a grade level to an age is a convention about school years, not a finding about readers. Our tool states its mapping as its own rather than borrowing someone else's authority for it — which is worth knowing before you present an age range to anybody as fact.

The short version

  1. Every formula measures sentence length and word length. That is all.
  2. A grade level is not a quality or difficulty score.
  3. Under about 100 words, the number is noise.
  4. Read all six — the spread tells you as much as any one score.
  5. The useful check is your description against your book, computed the same way.
  6. Absolute targets matter for children's books; for everything else the number is comparative.