AI Detection Is Not Authorship Detection: What Substack’s AI Score Actually Tells You

AI detectors are getting better. That does not mean they can tell who actually authored a piece of work. Substack’s Pangram integration exposes a deeper problem: detecting machine-generated prose is not the same thing as measuring human thought, research, creativity, or intellectual contribution.
A person working late at a desk with papers, books, a laptop, and a wall covered in notes, while a monitor in the foreground shows an 18% progress chart.
Contents

Here is the problem with AI content detection in 2026:

The detector can be right, and the accusation can still be wrong.

That is the part of the AI-writing debate we have not learned how to talk about.

In July 2026, Substack introduced a reader-facing “Scan for AI text” feature powered by Pangram. Readers can now ask Substack to analyze eligible posts, Notes, comments, and replies and receive an estimate of how much of the text appears human-written or AI-assisted. Writers can also add a “How I make this” statement explaining their process or disable detection on individual posts.

There is a legitimate problem Substack is trying to address. Fully automated accounts can mass-produce enormous amounts of low-effort material. People can present generated work as personal experience they never had. Students can submit assignments they do not understand. Publishers can flood search engines and social platforms with thousands of barely reviewed articles.

That deserves scrutiny.

But an AI detector does not actually observe any of those behaviors.

It sees the finished text.

It does not see the research session that preceded it. It does not know who developed the argument. It does not know who selected the sources, challenged the assumptions, rejected bad explanations, found contradictions, decided what mattered, or spent an hour interrogating evidence before asking an AI model to assemble the final prose.

And that creates a category error that becomes more dangerous—not less—as AI detection improves.

A percentage of AI-generated text is not a percentage of AI-generated thought.

That distinction should be obvious.

Increasingly, it is not.

The Old Criticism of AI Detectors Is No Longer Good Enough

Sherafy.com has criticized AI detection before, particularly its false positives, its tendency to stigmatize legitimate assistance, and the strange incentives created when writers become more concerned with looking “human” to a classifier than producing good work.

Some of those criticisms remain valid.

But the technology has changed.

It would be intellectually lazy to keep arguing against the weakest AI detectors of 2023 as though nothing has improved.

Pangram 4, released in July 2026, is considerably more sophisticated than the early-generation tools people associated with simplistic measures such as perplexity. Pangram describes its latest system as a deep-learning classifier capable of identifying human, AI-assisted, AI-generated, mixed-authorship, and deliberately “humanized” text. Its technical report claims a false-positive rate of 0.0041% on more than one million human-written English texts and an overall false-negative rate of 0.3396% in its primary evaluations.

Those are Pangram’s own results and should therefore be treated as vendor-reported evidence, not independent proof.

But independent research is also more favorable to Pangram than the blanket claim that “AI detectors don’t work.”

A 2025 University of Chicago working paper compared Pangram, OriginalityAI, GPTZero, and RoBERTa across a large benchmark of human- and AI-generated passages. The researchers found substantial differences among detectors and reported particularly strong performance from Pangram, including near-zero false-positive and false-negative rates within their test stimuli.

So this article is not going to pretend that Pangram is useless.

That would miss the much bigger problem.

Suppose Pangram becomes nearly perfect at detecting whether an AI model generated the words in front of you.

It still would not tell you who authored the work in the broader sense people actually care about.

What Does “90% AI” Actually Mean?

Imagine a detector analyzes an article and reports that 90% of its prose appears AI-generated.

Most readers will not interpret that sentence technically.

They will hear:

The person didn’t really write this.

They didn’t do much work.

The ideas probably came from ChatGPT.

They may not even understand what they published.

This is fake.

This is slop.

But none of those conclusions necessarily follows from the detector score.

The detector may have strong evidence about who or what rendered the final sentences.

That is not the same thing as knowing who created the intellectual substance those sentences express.

This distinction becomes obvious once you compare two hypothetical writers.

Writer A and Writer B

Writer A begins with an original question.

They spend an hour working with an AI system.

They propose a hypothesis.

The AI points out a weakness.

The writer changes it.

They ask for relevant research.

They open the studies.

They identify a contradiction between two papers.

They ask the AI to investigate the discrepancy.

They reject one explanation.

They develop another.

They upload a government report and ask the model to compare it against a dataset.

They discover something unexpected.

They change the article’s thesis.

They decide which evidence is strong enough to include.

They challenge several conclusions.

They decide what remains uncertain.

They organize their argument.

Then, after perhaps 60 or 90 minutes of research and reasoning, they say:

“Take everything we have established and write the final article.”

The AI generates almost every final sentence.

Now consider Writer B.

Writer B opens an AI chatbot and types:

“Write me an insightful 2,500-word article about this subject.”

The AI chooses the argument.

The AI finds the points.

The AI supplies the examples.

The AI creates the structure.

The AI reaches the conclusion.

Writer B then spends an hour manually rewriting every sentence so it sounds different.

Which writer contributed more intellectually?

The answer is not difficult.

Yet a text detector may have considerably more evidence of AI generation in Writer A’s article than Writer B’s.

That is not necessarily a detector failure.

It is an authorship-measurement failure.

We are asking the machine a question it was never equipped to answer.

Textual Provenance Is Not Intellectual Provenance

We need better language for this.

There are at least two separate things being discussed whenever someone asks whether something was “written by AI.”

Textual provenance

Where did the actual sequence of words originate?

Did a human type them?

Did an LLM generate them?

Did a human draft them before an AI rewrote them?

Did an AI draft them before a human rewrote them?

This is increasingly within the reach of sophisticated AI detection.

Intellectual provenance

Where did the substance originate?

Who chose the problem?

Who developed the hypothesis?

Who decided what research was necessary?

Who selected the evidence?

Who recognized contradictions?

Who challenged bad reasoning?

Who made the conceptual connections?

Who determined what conclusion was justified?

Who takes responsibility when the conclusion is wrong?

That is much harder.

And a post-hoc text detector simply does not have access to most of that history.

Researchers behind the 2026 Humanly project describe precisely this problem: a final document does not reveal whether it emerged from human typing, AI generation, or a mixed human-AI process. They argue that different workflows—human draft plus AI polish, AI draft plus human editing, translation, brainstorming, and other collaborative processes—can have very different meanings even when the finished text alone cannot reconstruct what happened.

That gets much closer to the actual problem.

A detector examines an artifact. Authorship is a process.

You cannot reconstruct an entire invisible process from an artifact and then pretend the resulting probability is a measurement of human contribution.

Pangram’s Own Technical Model Illustrates the Difference

Pangram deserves credit for explicitly trying to handle mixed authorship rather than pretending every document fits neatly into “human” or “AI.”

Its current technical framework distinguishes human, AI-assisted, and AI-generated material.

To construct training labels for edited text, Pangram compares transformed material against source material. Clauses closely retaining the source can be labeled human; substantial rewrites retaining the underlying meaning can be labeled AI-assisted; newly generated material can be labeled AI-generated. Its training framework then constructs a weighted AI fraction in which AI-assisted characters receive partial weight and AI-generated characters full weight.

That is a reasonable engineering strategy for training a text classifier.

But notice what the variable describes.

Characters. Clauses. Edits. Surface provenance.

It does not measure whether a human supplied 90% of the underlying reasoning.

It cannot know that the crucial argument in paragraph nine took 35 minutes of human-directed investigation to establish.

It cannot know that the model originally suggested the opposite conclusion and the human rejected it.

It cannot know that six sources were discarded because the human found them unreliable.

It cannot know that the human noticed a missing variable that fundamentally changed the analysis.

The final text does not contain a recording of the intellectual journey that produced it.

A machine cannot audit a history it never saw.

Mixed Human-AI Writing Is Already Exposing the Limits

Pangram 4’s strongest binary results should not be confused with perfect reconstruction of complicated collaborative writing.

On Pangram’s own evaluation set involving substantial AI editing of human essays, its latest model correctly classified the material as mixed in about 55% of cases. On another mixed-edit dataset created from real-world editing prompts, mixed-authorship recall was about 65%. Those results are major improvements over Pangram’s previous model—but they also demonstrate that nuanced assistance remains harder than separating clean human samples from clean AI samples.

Pangram performs much better on another type of mixed-authorship benchmark in which fully human and fully AI-generated sentence blocks are deliberately interleaved. Its overall token-level accuracy on that controlled benchmark reached about 92%. But that experiment is fundamentally cleaner than many real writing sessions because the test material contains known blocks of human and machine prose with constructed boundaries.

Real collaboration can be far messier.

A human may contribute a concept.

An AI may challenge it.

The human may find a paper.

The AI may summarize it.

The human may notice the summary missed something.

The AI may compare it against another paper.

The human may dictate a new argument.

The AI may rewrite it.

The human may replace a source.

The AI may restructure the section.

The human may make the final judgment.

Who “wrote” that paragraph?

There may not be a one-dimensional answer.

Even the Newest Research Is Running Into the Policy Problem

A separate August 2026 preprint examining AI detection in academic writing found a revealing mismatch between detection and misconduct.

The researchers tested Pangram 3.2—not the newer Pangram 4—and GPTZero against academic abstracts. They found that relatively light AI refinement could trigger detector flags at high rates while deliberate “humanization” of more heavily generated text could sharply reduce detection. The authors therefore argued that detector scores should not be used as standalone evidence of academic misconduct.

That paper should not be misrepresented as a direct test of Substack’s current Pangram 4 system. It isn’t.

But its underlying policy question survives every model upgrade:

What exactly are you trying to prove?

If a university’s rule is:

No AI may modify any part of this assignment.

Then detecting AI modification is directly relevant.

But if the university’s rule is:

The student must demonstrate their own understanding and reasoning.

Then AI-text detection is only indirect evidence.

Those are different rules.

The same applies to journalism, publishing, business, literature, and online commentary.

“The Detector Was Right” Does Not Mean “The Human Cheated”

This is the core distinction.

Suppose an article really is 95% AI-generated at the sentence level.

Fine.

The detector may be right.

Now what?

You still have to determine what that means.

Was somebody mass-producing 500 SEO articles about products they never researched?

Was a teenager outsourcing an essay they were explicitly required to write unaided?

Was a novelist secretly selling an entirely generated manuscript as handcrafted prose?

Or was an independent researcher using an AI system as a combination of research assistant, interlocutor, developmental editor, database navigator, fact-checking partner, and copywriter while personally directing the entire investigation?

Those cases are not ethically equivalent.

Treating them as equivalent because they can produce similar detector percentages is intellectually indefensible.

The detector did not tell you the meaning of the score.

You added the meaning yourself.

We Should Stop Pretending “AI Slop” Means “AI Was Used”

There is enormous amounts of AI slop online.

There is also enormous amounts of human slop online.

Humans invented spam.

Humans invented content farms.

Humans invented clickbait.

Humans invented plagiarism, fake reviews, astroturfing, propaganda, keyword stuffing, meaningless corporate copy, fabricated expertise, and confidently publishing things they had not bothered to understand.

Generative AI dramatically lowers the cost of scaling those behaviors.

That is a real problem.

But the defining feature of slop should not be that a transformer generated the syntax.

The useful distinction is closer to this:

Was automation used to replace human judgment, or to amplify it?

That tells us much more.

An AI can generate 2,000 words after a seven-word prompt.

An AI can also generate 2,000 words after a lengthy research process involving dozens of decisions made by a human who understands every conclusion and accepts responsibility for every factual claim.

Calling both processes simply “AI-generated content” may be technically convenient.

Intellectually, it is almost useless.

Human Authorship Has Multiple Layers

Instead of asking for one magical percentage, a serious assessment of authorship would consider several dimensions.

  1. Idea origination: Who identified the question, thesis, or problem?
  2. Research direction: Who decided what needed investigation and which evidence mattered?
  3. Reasoning and judgment: Who evaluated competing explanations, contradictions, uncertainty, and causality?
  4. Synthesis and structure: Who determined what the work ultimately argues and how its parts fit together?
  5. Surface expression: Who or what rendered the final words?
  6. Verification and accountability: Who checked the claims, approved the finished work, and accepts responsibility for errors?

Today’s AI detectors are heavily concentrated around number five.

People are using number five to make judgments about numbers one through six.

That is the mistake.

Why This Matters Most for People Without Large Teams

There is another part of this debate that becomes uncomfortable very quickly.

AI tools are not merely toys for people trying to avoid work.

They can also function as labor multipliers.

In a randomized experiment involving 444 college-educated professionals performing writing tasks, researchers Shakked Noy and Whitney Zhang found that ChatGPT reduced completion time while increasing average output quality. The technology also changed the composition of work: participants spent proportionally less effort rough-drafting and more effort on idea generation and editing.

Separate research involving thousands of customer-support workers found that access to a generative AI assistant increased productivity on average, with especially large gains among less-experienced and lower-skilled workers in that setting.

These studies do not prove that every use of AI improves work.

They show something more important:

AI can redistribute what humans spend their limited time doing.

That matters enormously for people without institutional resources.

A large newsroom can hire researchers, junior writers, editors, data analysts, SEO specialists, fact-checkers, designers, and production staff.

A small independent publication cannot.

A well-equipped academic institution can employ research assistants.

A freelance researcher cannot.

A corporation can distribute one project across fifteen salaries.

An individual creator has one body and 24 hours in a day.

AI changes that equation.

A single person can now perform workflows that previously required a small team—not because the human suddenly stopped contributing, but because parts of the execution layer became dramatically cheaper.

That democratizing effect should not be romanticized. AI also creates displacement, misinformation, concentration-of-power, training-data, labor, and quality-control problems.

But dismissing every AI-assisted creator as lazy ignores one of the technology’s most consequential effects:

People with less money can suddenly attempt work that used to require far more money.

That is not a trivial development.

SHERAFY Is an Example of Why the Binary Breaks Down

This publication is itself a useful case study.

SHERAFY.com has benefited enormously from AI systems.

There is no reason to hide that.

Its founder has worked with AI technology for years and uses both commercial and open-source tools throughout research and production workflows.

Sometimes AI helps identify a question worth investigating.

Sometimes the human brings the question first.

Sometimes AI searches.

Sometimes documents are supplied directly.

Sometimes a hypothesis survives scrutiny.

Sometimes it collapses.

Sometimes an argument changes repeatedly during research.

Sometimes the final article is substantially rendered by AI after the core reasoning has already been developed through an extended human-machine dialogue.

If someone takes only the finished article, uploads it to a detector, and receives a high AI probability, they may have learned something about the prose-generation layer.

They have not reconstructed the research process.

They have not measured who developed the thesis.

They have not measured who challenged the evidence.

They have not measured the intellectual novelty of the investigation.

They have not measured who was accountable for publication.

And they certainly have not obtained a magical percentage telling them how much of the article’s thought came from a machine.

If someone wants to criticize the reporting, there is a much better method:

Read it.

Check the sources.

Challenge the evidence.

Find the logical error.

Identify the unsupported inference.

Demonstrate that the conclusion is wrong.

That is criticism.

Running prose through a detector and announcing a percentage is not a substitute for engaging with the work.

Copyright Law Already Recognizes That “AI Was Used” Is Not the Whole Question

Copyright law is not an AI-detection standard, and legal authorship should not be confused with philosophical or editorial authorship.

But the U.S. Copyright Office’s current framework illustrates why binary thinking is inadequate.

Its 2025 report says that using AI as an assistive tool does not by itself destroy copyright protection. Human-authored expression can remain protected even when AI-generated material appears within a larger work, and human creative selection, arrangement, coordination, and modifications can also be protectable. At the same time, the Office says purely AI-generated material is not protected and that prompts alone, under currently generally available technology, ordinarily do not give the user enough control over expressive output to establish authorship of that output. Questions about sufficient human contribution are assessed case by case.

That is far more nuanced than either extreme.

It does not say:

“I prompted the AI, therefore I personally authored every word.”

But it also does not say:

“AI participated, therefore there is no human authorship.”

The law asks about human contribution and control.

The cultural conversation should probably learn something from that distinction.

There Are Cases Where AI Detection Is Completely Appropriate

None of this means AI detection has no legitimate purpose.

If a professor explicitly assigns an unaided writing exercise because the objective is to measure a student’s independent prose construction, AI generation is directly relevant.

If a literary competition requires submissions to be entirely human-written, detection may help enforce the rule.

If someone sells customers “100% personally written” material and secretly automates everything, provenance matters.

If a platform is fighting bot networks that create millions of fake comments, reviews, or propaganda posts, detection can be valuable.

And if someone claims personal experiences they never had by having an AI fabricate them, that deception matters regardless of literary quality.

The error is not detecting AI.

The error is failing to define why AI presence matters before interpreting the detection.

The policy has to come first.

The detector comes second.

Otherwise the score becomes a technological Rorschach test onto which people project whatever moral judgment they already wanted to make.

Substack Accidentally Built the Better Solution Next to the Detector

One of the more interesting parts of Substack’s implementation may not be Pangram at all.

It is the optional “How I make this” statement.

Substack lets creators explain how they make their work, and that statement can appear alongside a scan result.

That is much closer to the information readers actually need.

Imagine a straightforward disclosure:

Research and analysis are human-directed. AI tools are used for document analysis, research assistance, iterative critique, organization, editing, and final drafting. All claims, sources, conclusions, and publication decisions are reviewed and approved by the author.

Now the reader knows something meaningful.

Compare that with:

83% AI.

Which one actually describes how the work was made?

The percentage looks more scientific.

The disclosure contains more information.

That should tell us something.

The Future Should Be Process Provenance, Not AI Fortune-Telling

There is already research moving beyond post-hoc detection.

The Humanly project, for example, explores writing environments that preserve process evidence: activity logs, permitted AI interactions, drafting history, typing behavior, and other information that can help verify whether a writer followed a particular authorship policy. The researchers explicitly frame authenticity around whether the production history matches the claimed process rather than asking a classifier to infer everything from the final document.

This approach is not perfect either.

Continuous process logging raises privacy concerns.

Writers should not have to surrender complete surveillance records of every creative act simply to prove they are human.

And most ordinary publishing should not require forensic authorship certification at all.

But when authorship genuinely matters—academic examinations, competitions, legal declarations, certain research environments—process evidence is fundamentally more informative than post-hoc stylistic suspicion.

That is the direction worth developing.

Not better divination.

Better provenance.

Stop Asking the Wrong Question

The cultural debate keeps asking:

“Was this written by AI?”

That question is rapidly becoming inadequate.

Ask instead:

Who developed the idea?

Who directed the research?

Who evaluated the evidence?

Who made the important decisions?

Who understands the final argument?

What did the AI actually contribute?

Was that contribution disclosed where disclosure matters?

And who is accountable for the result?

Those questions can distinguish meaningful human-AI collaboration from automation masquerading as expertise.

A single AI percentage cannot.

The irony is that AI detection is becoming more sophisticated at exactly the moment when the concept it supposedly measures is becoming less binary.

People brainstorm with models.

They research with models.

They argue with models.

They upload PDFs.

They challenge outputs.

They dictate paragraphs.

They ask for rewrites.

They combine their own notes with machine summaries.

They reorganize generated drafts.

They replace generated claims with original findings.

They may spend hours directing a process whose final sentences take seconds to render.

What percentage of that is “AI”?

There is no scientifically honest answer unless you first define what you are measuring.

Characters?

Keystrokes?

Ideas?

Research decisions?

Semantic novelty?

Editorial control?

Time?

Responsibility?

All of those could produce different percentages.

That is why “90% AI-generated” should never be casually translated into “90% not human-authored.”

Those are different propositions.

The Detector Can Be Right and the Accusation Can Still Be Wrong

This is where the third generation of the AI-detection debate has to go.

The first question was whether AI-generated text could be detected.

The second was whether those detectors were reliable enough to trust.

The third is more important:

What exactly have we proven after the detector works?

Sometimes we have proven something important.

Sometimes we have merely identified the tool that generated a sequence of words.

Those are not the same thing.

Pangram may continue getting better.

Its false-positive rates may fall.

Its mixed-authorship classification may improve.

Future detectors may become astonishingly good at distinguishing human-rendered language from model-rendered language.

None of that solves the authorship problem.

Because authorship is not merely the physical production of sentences.

It can involve conception, investigation, judgment, synthesis, direction, revision, verification, and responsibility.

AI can participate in some of those layers.

Humans can participate in others.

And sometimes the human doing the deepest intellectual work will be the person who typed the fewest final words.

If platforms want to fight spam, deception, fake engagement, fabricated expertise, undisclosed automation, or genuinely low-effort AI slop, they should fight those things directly.

But they should stop encouraging the public to confuse text detection with mind reading.

The detector can inspect the words.

It cannot see the hour of thought that came before them.

And no percentage—no matter how many decimal places it has—can measure work the detector never observed.

AI detection is not authorship detection.

It is time we stopped pretending otherwise.

References and Further Reading

Substack and Pangram

  • Substack — How can I detect AI on Substack? — Substack’s official documentation for its reader-facing Pangram integration, including supported content, creator controls, the “How I make this” disclosure, and per-post detection settings.
  • Pangram 4 Technical Report — Pangram’s July 2026 technical paper describing its current classifier, training framework, human/AI-assisted/AI-generated categories, false-positive and false-negative evaluations, mixed-authorship testing, and humanization detection. Because Pangram researchers authored the report, its performance claims should be read alongside independent evaluations.
  • University of Chicago Becker Friedman Institute — Artificial Writing and Automated Detection — Independent working paper by Brian Jabarian and Alex Imas comparing Pangram, OriginalityAI, GPTZero, and RoBERTa and finding particularly strong binary-detection performance from Pangram within the study’s benchmark.

Mixed Authorship and Process Evidence

Human Contribution and Copyright

AI as a Productivity and Capability Multiplier

Earlier SHERAFY Analysis

Editorial currency note: AI-text detection is an unusually fast-moving technical field. Substack’s implementation, Pangram’s models, published benchmark results, and platform disclosure policies may change after publication. Detector performance figures in this article are tied to the specific model versions and research datasets identified above and should not be generalized to future versions without re-evaluation.

Cite this article

Published August 19, 2026

More to think on...