The AI-detection panel and the writing-health panel are now one panel, called **Manuscript check**. Both of their scores are gone, and neither is coming back.

## What changed

- **One panel instead of two.** Manuscript check replaces both. Once the scores were removed, nothing was left to tell the two panels apart.
- **No AI-suspicion score.** The panel no longer estimates how much of your draft reads as AI-generated. In its place it carries a plain line: this cannot tell you whether text was written by AI, and neither can any other tool.
- **No writing-health grade, and no reading-ease number.** The passive voice, hedging, nominalization, and adverb warnings went with them.
- **Findings you can check yourself.** Spelling that drifts between two forms in the same document (analyse and analyze), hyphenation and capitalization that drift the same way, abbreviations used before they are defined, jargon with a plain equivalent, wordy phrases, and phrasing that reads as AI. Every finding points at a passage you can read and judge.
- **Typography is cosmetic, and says so.** Curly quotes and dashes are still listed, under Typography, described as what they are: punctuation your word processor inserted while you typed.
- **Spelling and grammar, in your browser.** New, and off by default. Turn it on and it runs entirely on your device, so your manuscript is never uploaded. One-time download of about 8 MB. For unpublished or embargoed work, that is the whole point.
- **Statistics, not grades.** Sentence length distribution and counts, described rather than scored.

## Why

Our own blog argues that AI-detection percentages are "probability estimates with documented bias, not lie detectors." The panel should agree with the blog.

The score was measuring typography, not writing. Watson and Crick's 1953 paper on the structure of DNA scores as low AI signal as they typed it. Paste the same words into Microsoft Word, let autocorrect turn the quotes and dashes into their curly forms, and the score climbs twenty points into "moderate AI signal." Not one word changed. The penalty was for using a word processor.

Document-level AI detection is not good enough to put a number on a scholar's page, and it fails hardest on the people who use Alfred Scholar. OpenAI withdrew its own classifier. A Stanford study measured a 61.3% false positive rate on writing by non-native English speakers. Detection degrades most on polished text, which is what careful academic writing is. The researchers behind the strongest work in the field say plainly that their method cannot determine whether a specific document used a language model.

Writing health was the more damaging of the two, because it gave advice. It told researchers that "may," "might," "suggests" and "likely" weaken their claims, and to state findings directly. That is backwards. Hedging is how a correlational study honestly declines to claim causation. The tool was instructing scholars to overclaim. Its reading-ease meter had the same problem from the other side: the 100 most-cited neuroimaging papers average 15.7 on that scale, which the old panel called "very difficult."

A score you cannot defend is not a signal. It is a number that feels like one. We took the numbers out and kept the findings you can verify.