On 21 July, Substack switched on a scanner. A reader can now ask the platform to check a post, a Note, a comment or a reply longer than a hundred words, provided it was published on or after that date, and a classifier called Pangram returns an estimate of how much of the text was written by hand and how much with AI assistance. Chris Best, Substack’s chief executive, gave the offence a name. Claudefishing: using AI to fake the human connection a reader thought they were entering.[1]
The worry underneath the feature is legitimate. Unmeant prose is in supply at a scale nothing before the models could reach, feeds are thick with sentences nobody stands behind, and reading has always been an act of trust. Reading is also an exchange of attention, mine for yours, and what makes slop offensive is less the machine’s involvement than the suspicion that nobody spent anything on the other side.
Best is more careful than many of the reactions to him. Not everything made with AI is slop, he writes, and not all slop is made with AI, and Pangram can detect possible machine involvement without being able to tell whether great care went into the work.[1] Nor is the limit hidden in the product. Substack lets writers add a “How I make this” statement, report a result they believe is wrong and now disable detection on an individual post or Note, and scan results are visible only to the person who requests them. Those are real safeguards. They do not make the design neutral.
The percentage arrives in one tap, numeric and ready to be screenshotted. The account asks to be read. A writer who disables the scan does not recover an unmarked space, because the reader is shown the words “AI detection unavailable” instead.[2] Substack presents the score and the account as complementary forms of evidence, and the interface has already ranked them. One arrives as measurement. The other arrives as defence.
Most of the argument since the launch has concerned whether Pangram works, and the history of the category gives people reason to ask. Early detectors misclassified human work, showed bias against some writers working in a second language and were placed inside disciplinary systems that gave a probability score more force than it could bear. Vanderbilt disabled Turnitin’s detector in 2023. Waterloo followed in 2025, and Curtin from January 2026, each directing attention instead towards evidence of learning, assessment design and transparent use.[3]
Pangram is a stronger instrument than those first systems. An independent University of Chicago study found near-zero false-positive and false-negative rates within its test corpus, and Pangram was the only detector evaluated that met the researchers’ strictest policy threshold.[4] The Atlantic later showed how much harder the result becomes to interpret when it leaves a benchmark and enters disputes about revision, humanising tools and reputational harm.[5]
But I do not need Pangram to fail. Grant it the strongest case.
A weak detector gets ignored. A strong one gets believed.
Suppose the scanner reads my next essay and returns a number that is, by its own lights, correct. I publish under a standing disclosure, so the number will not be flattering and it will not be wrong. What does the reader now know?
A few days after the launch, I pinned an essay called Most Evenings to the top of my homepage.
It was published at the end of March, which places it outside the scanner’s jurisdiction. The first thing a visitor to this newsletter now meets is a piece the detector cannot score, describing in some detail exactly the practice the detector exists to find.
Most Evenings is a room. A man on a sofa, leaning over a laptop on a coffee table too low for the purpose, a Korean zombie film muted on the television, his wife asleep, his daughter reading, the dog sighing at the far end. The man runs a prompt through one system, keeps a phrase, discards the rest. He runs it through another, keeps a connection he had not considered, discards that output too. Some nights the machines are only the thinking he needed to do before he could think. Other nights a section arrives better than anything he would have written. He keeps it, then works the surrounding prose until the seam disappears. His daughter once asked what he does on the laptop every night. He told her he was learning and could not say what.
That sentence about the seam is the one to come back to, because I no longer think it was finished.
Every piece of writing made with a machine has two seams running through it. They are not the same seam, and treating them as one is the mistake beneath the scanner.
The first seam is aesthetic. It runs between the paragraph the model offered and the paragraph I wrote, between the phrase kept from one system and the structure borrowed from another, between registers that would grate if left sitting against each other raw. Craft exists to close this seam. The reader should not have to step over the joins. The finished piece should read as one thing in one voice, because that is what finishing means.
The second seam is ethical. It runs between the work and its account of itself, between what the piece is and what the reader has been given to believe about how it came to be. This seam should remain visible. It belongs in the standing note that says a machine was part of the process and what part it played, in the willingness to describe the practice when asked, and in the record of choices for which a person remains answerable. Closing the first seam is craft. Closing the second is the lie.
There is an objection to this, and it is a good one. It says the seams cannot be separated as conveniently as I have just separated them, because closing the first does moral work whether I intend it or not. Prose persuades continuously, sentence by sentence, beneath the level of argument. A piece that reads as one voice gives the felt impression of a single mind at work. That impression is already a claim about origin. The disclosure arrives later, in smaller type, after the persuading is done. Reading is not an audit. On this account, the writer who smooths the join and then files a note has sold the impression and issued a receipt for it.
I cannot dismiss that, and I notice how much I would like to, being the one holding the receipt.
What I would say against it is that the receipt has to be specific enough to fail. “Written with AI assistance” is a receipt, and a nearly worthless one. It settles nothing, predicts nothing and exposes the writer to nothing. An account is a different object. It says what the machine did and what I did, what gets kept and what gets thrown away, where the balance tipped, and what standards the finished work claims to have met. It can be held against the writing and found wanting.
If I say I checked the study a paragraph leans on, a reader can check one. If I say fluency is the first thing I distrust, a reader can look for the places where I let a smooth line stand because it sounded good. If I say the model proposed the structure but not the conclusion, that claim can be compared with the drafts, or with the habits visible across the rest of the work. An account makes claims that can be caught. It does not close the objection. It gives the objection something to grip.
That is what I would put in the place where purity used to sit. Authorship survives the tools wherever there is someone who can be asked why a sentence is there and who will answer for what it does. Readers who want writing no model has touched are entitled to the preference, and it is a separate matter from whether a work is honestly presented, seriously made and owned by the name beneath it.
Now look at what the detector does with any of this. It reads the surface of the prose for patterns associated with machine involvement. It hunts for the first seam, the one a competent writer has spent the evening removing, then invites the result to stand as evidence about the second. The apparatus turns a fact about textual production into a fact about honesty.
Under that logic, a writer who closes the aesthetic seam well begins to look like someone with something to hide. A writer who leaves the prose lumpy begins to look innocent. Neither appearance tells us whether the study was checked, whether the claim is true, whether the metaphor clarified a relationship or quietly replaced the need to explain one, or whether anyone will answer when the piece causes harm.
The scanner never heard the sentences read aloud at eleven o’clock. It did not see the well-made line cut because it turned out not to be true. It has no opinion on whether an ending remained open because the subject had not earned closure or because the writer ran out of evening. Judgement can leave signs in prose, but it leaves no stable signature a classifier can isolate. The scan measures a fact about production. The ethical question is what that fact means inside the relationship between writer and reader.
The arms race makes the distinction clearer. Humanisers already exist to take model output and roughen or paraphrase it until a detector reads it as human. The Atlantic passed text from ChatGPT and Claude through one such service and found that Pangram then labelled the outputs human-written. Recent research has also shown that text from base models can appear strikingly human to commercial detectors, suggesting the systems may be tracking artefacts of model tuning rather than some permanent essence of machine authorship.[5][6]
Detector against humaniser, smoother against roughener, each side funding the other’s next release. The whole contest is fought along the aesthetic seam. The ethical seam belongs to neither. A classifier cannot locate it and a humaniser cannot settle it. A writer can lie about the process, of course. But the lie remains a claim made by a person, on the record, answerable to the work.
Then there is what the scanner does back to the writing, which nobody has to intend.
A number now exists, and readers can call for it, so the texture of prose has acquired a stake it did not have last month. Roughness begins to read as evidence. The odd sentence, the clumsy join, the paragraph that does not quite land, all of it gains a small forensic bonus. The writer knows this even when trying not to know it.
The launch interview contained a glimpse of that future before anyone had to imagine it. One of the hosts said he had largely stopped proofreading his daily essays after being accused of using AI. Passing the work through grammar correction gave it an “AI flavour”, so he left in stream-of-consciousness phrasing, strange turns and rough grammar. The roughness had become protective.[7]
I do not think many writers will sit down and consciously damage a sentence to beat a classifier. It will be quieter than that. A line will come out well and there will be a half-second of hesitation before it goes in. A join will be left proud. The revision that would have smoothed something will be skipped, and the reason will be narrated afterwards as restraint, or as letting the piece breathe, or as taste. The aesthetic seam will remain open, not for the reader but for the scanner.
In March I wrote that fluency had become the first thing I distrusted, because a response arriving too smoothly often meant I had asked the wrong question. That was a test about thinking. The scanner installs beside it a test about appearances, using the same evidence and pointing in the same direction. From inside the act of revision, the two will be difficult to tell apart. The pressure runs towards performing humanity rather than exercising it.
My own position in this is cushioned, and the cushioning belongs in the essay rather than in a reply to critics. I have disclosed the machine’s part in this work for years, from before doing so carried much reputational cost. A score that says assisted corroborates my account instead of contradicting it. I have a salary that does not move with my open rate, an archive long enough to establish a pattern and readers who arrived knowing the terms.
A score does not arrive into a vacuum. It lands on whatever standing a writer already has, and standing is distributed about as evenly as everything else. On a writer three months into a first newsletter, working in a second language, flagged by screenshot in a Notes thread by someone who dislikes the argument, the same number may become the entire body of evidence. The detector compounds the standing it encounters and gives suspicion a new object to circulate.
Readers who remain with a writer are good at testing an account over time. They notice evasions, repetitions, unexplained changes of register and claims that never survive contact with a source. A stranger carrying a screenshot into a feed is performing a different kind of judgement. The first has context. The second has a percentage.
Which points at what could have been built instead. A disclosure that travelled with every post by default, written by the writer, held in one place and readable across a whole archive, would hand a reader the material to test consistency over time, which is what they are actually trying to do when they reach for a scan. It would be slower. It would embarrass some people. It would produce no number at all, and it would be evidence about a person rather than about a paragraph.
This is why I want to be exact about what pinning Most Evenings does and does not do. It is not proof. A description of a practice can be fabricated as easily as a paragraph. What the pinned essay offers is a different genre of answer. A scan returns a verdict about origin. The essay gives an account of responsibility: here is the room, here is the method, here is what gets kept and what gets thrown away, here is a man reading his own sentences aloud in the dark and still not sure about them.
A verdict asks to be trusted. An account asks to be tested.
The coinage deserves a moment of its own. Best defines Claudefishing with some care, locating the deception in the mismatch between what a reader expects and what is actually there. Yet the name puts the offence inside a product. It pulls attention back towards the presence of Claude, when the fraud lives in the gap between what the writer invites the reader to believe and what the writer is willing to own.
A scanner can offer a clue about that gap and it cannot close it. Only an account, and conduct that goes on matching the account, can do that. Whether a model touched the sentences is a question with an answer. Whether anyone stands behind them is a question with a person at the end of it.
This piece, unlike the one pinned above it, qualifies for the scan. Point the thing at it, by all means.
Most evenings the room is the one it was in March. The laptop, the low table, something muted on the television, more language arriving than any one person needs. What is different is that I now write knowing a classifier can be turned on the result, and I cannot unknow it. I would like to tell you that the knowledge will stay outside the sentences. I do not believe it.
The number on this essay will say what it says, and the account sits above it, and between the two a reader has everything they are going to get from me. What I cannot tell you is what a year of writing under the scanner will do to the prose, or whether I will be able to tell when it has.
AI disclosure
This essay grew from my earlier essay Most Evenings and was developed through iterative work with Claude and ChatGPT. The systems helped test the two-seam distinction, surface objections, research the Substack and Pangram material and draft or rework parts of the prose. I chose the argument and final structure, revised the text, reviewed the cited evidence and take responsibility for the published version.
Notes
[1] Chris Best, “Against Claudefishing”, The Substack Post, 21 July 2026. Substack says the problem is an expectation mismatch, acknowledges that Pangram cannot assess human care, and explains the scanner and creator statement.
[2] Substack Help Center, “How can I detect AI on Substack?”, updated 26 July 2026. The page explains that creators can disable detection, after which readers see “AI detection unavailable”.
[3] Vanderbilt University, “Guidance on AI Detection and Why We’re Disabling Turnitin’s AI Detector”, 16 August 2023; University of Waterloo, “Discontinuing use of AI detection functionality in Turnitin”, effective September 2025; Curtin University, “Update on Turnitin AI-Detection Tool”, effective 1 January 2026.
[4] Brian Jabarian and Alex Imas, “Artificial Writing and Automated Detection”, Becker Friedman Institute for Economics, University of Chicago, Working Paper, 2 September 2025.
[5] Matteo Wong, “America Has a Pangram Problem”, The Atlantic, 30 May 2026.
[6] Yixuan Even Xu, Ziqian Zhong, Aditi Raghunathan, Fei Fang and J. Zico Kolter, “Base Models Look Human To AI Detectors”, arXiv, 19 May 2026.
[7] TBPN Digest, “Substack partners with Pangram to fight AI slop with content detection”, transcript of interview with Chris Best, 21 July 2026. The transcript is auto-generated and may contain errors.



Freddie deBoer posted this article challenging the false positive/negative metrics of Pangram. The primary purpose of the tool seems to be to enable trolls to dunk on authors with whom they disagree, à la Twitter/X. Will it come to the point authors need to display their AI assisted edits in “redline” format along with spelling and grammar correction, with bubble comments why they chose text A over text B?
The two-seams distinction is doing real work here, but I think the essay is sitting on top of an even blunter fact that doesn't need the apparatus to see: Best already told you the tool doesn't do the job it's being sold to do. He says outright that Pangram can't tell whether great care went into a piece. That's not a footnote or a limitation to be managed — that is the question a reader is actually asking when they reach for the scan. Nobody screenshots a percentage because they're curious about token distributions. They screenshot it because they want to know if the person on the other end meant it.
So the tool measures one thing and gets deployed to answer a completely different one, and the gap between those two isn't an oversight. It's the product. A detector that only reported stylistic origin, with no implied verdict on sincerity, wouldn't generate the anxiety that gets people checking their own posts and dropping screenshots into Notes threads about writers they've decided to distrust. The moral weight the number carries is doing commercial work that the number itself was never built to support.
Which means the honest response to "does Pangram work" isn't really about accuracy at all. Grant it perfect accuracy on its own terms — fine. It still can't see the thing everyone's using it to adjudicate. Best conceded that in the sentence introducing the feature. The confession is sitting right there in the launch copy, and the product shipped anyway.