Discussion about this post

User's avatar
Ken Hobbs's avatar

Freddie deBoer posted this article challenging the false positive/negative metrics of Pangram. The primary purpose of the tool seems to be to enable trolls to dunk on authors with whom they disagree, à la Twitter/X. Will it come to the point authors need to display their AI assisted edits in “redline” format along with spelling and grammar correction, with bubble comments why they chose text A over text B?

Peter Rex's avatar

The two-seams distinction is doing real work here, but I think the essay is sitting on top of an even blunter fact that doesn't need the apparatus to see: Best already told you the tool doesn't do the job it's being sold to do. He says outright that Pangram can't tell whether great care went into a piece. That's not a footnote or a limitation to be managed — that is the question a reader is actually asking when they reach for the scan. Nobody screenshots a percentage because they're curious about token distributions. They screenshot it because they want to know if the person on the other end meant it.

So the tool measures one thing and gets deployed to answer a completely different one, and the gap between those two isn't an oversight. It's the product. A detector that only reported stylistic origin, with no implied verdict on sincerity, wouldn't generate the anxiety that gets people checking their own posts and dropping screenshots into Notes threads about writers they've decided to distrust. The moral weight the number carries is doing commercial work that the number itself was never built to support.

Which means the honest response to "does Pangram work" isn't really about accuracy at all. Grant it perfect accuracy on its own terms — fine. It still can't see the thing everyone's using it to adjudicate. Best conceded that in the sentence introducing the feature. The confession is sitting right there in the launch copy, and the product shipped anyway.

5 more comments...

No posts

Ready for more?