Someone commented on one of my Substack notes that I had “literally wrote this note with AI,” and that the title of the piece was “clearly generated by AI as well.”
I had used AI to format the note. I hadn’t used it for any of the content. And the title wasn’t mine to generate: I was amplifying a former colleague’s article, which I’d said in the note. The line she flagged as machine-sounding, “what gets rewarded gets repeated,” is close to a mantra in my field. The first line she objected to was a direct quote from the article I was pointing at.
I reached out through a direct message. We ended up in a long and generous exchange about it, and I’ve thought about it since, because almost nothing in the accusation was correct and I still understood exactly why she made it.
I’ve revisited that exchange multiple times with curiosity. Bestselling authors use ghostwriters. Executives run posts through comms teams before they go up. Academics thank research assistants in a footnote and put their own name on the paper. Speechwriters are a profession.
Nobody runs a detector on any of it. Nobody asks. The byline is treated as sufficient.
Use AI to help with a LinkedIn post and the question comes immediately: did you actually write this?
That asymmetry deserves more scrutiny than it gets. The usual explanation is that AI use signals an absence of thinking, that the objection is to work produced without a mind behind it. That explanation isn’t supported by the evidence. If that were the mechanism, disclosing a ghostwriter would draw the same suspicion. It doesn’t. Something else is being adjudicated, and it isn’t whether thinking occurred.
In an earlier piece, Colby Kennedy Nesbitt and I described the legitimacy threshold: the shared expectations an algorithmic judgment has to meet before people will accept it for high-stakes human consequences. We applied it to performance management, to what AI decides about people.
The same threshold governs a question we did not examine then. Not what AI decides about us, but what we permit it to help us make.
What the research finds
The disclosure penalty is well established and larger than most people assume. Schilke and Reimann (2025) ran thirteen preregistered experiments across more than 4,000 observations, and found that disclosing AI use reduces trust with every kind of evaluator they tested: students assessing a professor, hiring managers assessing an applicant, investors assessing a fund, legal professionals assessing a supervisor, tax filers assessing an advisor. The mechanism they identify is not accuracy. It is legitimacy.
That much has been widely reported. What hasn’t been is the comparison condition they built into the design.
Three of the thirteen studies included a human-help arm alongside the AI arm. A professor who disclosed that a graduate teaching assistant graded the papers. An applicant who disclosed that a career coach helped with the letter. A fund that disclosed an outside analyst wrote the ad.
Measured against saying nothing at all, admitting to the human helper barely registered. In all three cases the difference was small enough to be noise. Admitting to AI produced a large drop every time.
It’s tempting to read it as disclosing human help is free. That read is wrong, and the authors’ own follow-up says so. When they re-ran the career-coach study with about seven times as many participants, the human-help penalty did show up. But it was roughly half the size of the AI penalty. The earlier studies had too few people to detect an effect that small.
So the pattern is narrower and more interesting. Disclosing human help costs you something. It costs about half what disclosing AI costs.
Research by Reif, Larrick, and Soll (2025) approached it from a different direction. Across four experiments with 4,439 participants, they varied only the source of the assistance while describing the help itself identically. An attorney who “sometimes asks a paralegal to summarize information” was rated barely differently from one who asked no one. An attorney who “sometimes asks generative AI” to do the same thing was rated lazier, less competent, less diligent, less independent, and less self-assured.
Two research teams, different tasks, different populations, same pattern.
And then it reverses
Here is where the story stops being about AI.
Claessens, Veitch, and Everett (2026) ran a study with a third condition. Participants read about someone who either did a task themselves, “gets the AI tool ChatGPT to do it for them,” or “gets someone else to do it for them.” Twenty tasks: a love letter, an apology, wedding vows, a bereavement card, computer code, a dinner recipe.
Outsourcing to another human was judged more harshly than outsourcing to AI: less competent, less moral, less trustworthy, and lazier. The competence gap was the widest, close to a full point on a seven-point scale.
The penalties were largest for love letters and apologies, smallest for code and recipes.
Another study by Liu, Kang, and Wei (2024) found a compatible result for personal messages. When a close friend used help to write you a supportive note, getting help from another person was statistically indistinguishable from using AI on perceived effort, relationship satisfaction, and appropriateness. The authors’ explanation was that people don’t think a friend should use any third party, AI or another human.
In at least one domain the ordering flips entirely. Jago and Carroll (2024) found across four studies that producers received more credit for work when assisted by algorithms than when assisted by humans. The reason they identify is an assumption that algorithmic assistance requires more oversight from the producer.
Put these together and the pattern is not “people distrust AI.” At least not uniformly.
The actual variable
What predicts the penalty does not seem to be whether the helper was organic or inorganic. It is whether that kind of outsourcing has been ratified for that kind of task.
Where the human alternative is a sanctioned role-holder doing their job (a paralegal summarizing case law, a teaching assistant grading, a comms team polishing a statement), the arrangement is already legitimate. Everyone knows the role exists, what it covers, and who stays answerable. AI has no such standing yet, so it absorbs roughly twice the penalty.
Where the helper is an unspecified someone absorbing effort you personally owed, no arrangement is legitimate. Nobody has ratified having your apology written for you. Delegation is itself the violation, and a human delegate is judged as harshly or worse, because at least the machine was purpose-built to be used.
This is the legitimacy threshold, operating on authorship instead of assessment. It explains something the “AI can’t think” framing cannot: why the same person can use AI to draft a project update without comment and be accused of fraud for using it on a personal essay. The tool is constant. The ratification is not.
It also explains why the accusation stings. Being told your writing sounds like AI is not a claim about your intelligence. It is a claim that you used a form of help that is being perceived as illegitimate. That is a social charge, not a cognitive one, and social charges are harder to rebut, because there is no evidence you can produce.
But social ratification isn’t the whole story. I know because I’ve refused a ratified arrangement myself.
When I was building my own business, I dreamed about hiring someone to run my social media. Everyone told me to. It is about as legitimate an outsourcing arrangement as exists in professional life, and nobody would have thought twice about it. I tried a few times. I never could do it. Handing over my own voice felt wrong in a way I couldn’t argue myself out of, even though I could not have told you what principle I was defending.
So there are two thresholds. There is the social question of what a community will accept, which is what the research measures. And there is a personal question about where your voice stops being yours or what value you get out of a task that you’re unwilling to trade. No amount of convention settles that for you. The second one is why common practice never fully resolves the discomfort, and why watching someone else use AI comfortably doesn’t make you comfortable.
We already know how to ratify help
Institutions have been solving the first question separately and without much fanfare, and three of them drew the same line independently.
The EU AI Act requires disclosure of AI-generated text published to inform the public on matters of public interest. Then Article 50(4) carves out an exemption: the obligation does not apply “where the AI-generated content has undergone a process of human review or editorial control and where a natural or legal person holds editorial responsibility for the publication of the content.”
Not “where a human typed it.” Where a named party is answerable for it.
The US Copyright Office reached the same distinction through different reasoning (Copyright and Artificial Intelligence, Part 2, January 2025). Prompts alone do not confer authorship, and iteration doesn’t fix that. The Office’s phrasing is that revising prompts repeatedly amounts to “re-rolling the dice,” and that “no matter how many times a prompt is revised and resubmitted, the final output reflects the user’s acceptance of the AI system’s interpretation, rather than authorship of the expression it contains.”
The other half of the ruling gets far less attention. Copyright still protects what the person contributed, even when AI-generated material is mixed in: “Copyright protects the original expression in a work created by a human author, even if the work also includes AI-generated material.” How you choose and arrange what the AI produced can count too.
One artist drew an image by hand and used it as her input. Her drawing was still visible in what the AI returned, so the Office registered that part of the work.
Accepting a machine’s interpretation is not authorship. Putting your own expression through a machine is. The law already distinguishes these.
Amazon’s publishing platform draws the line in operational terms. Content counts as AI-generated, and must be disclosed, “even if you applied substantial edits afterwards.” Content counts as AI-assisted, and requires no disclosure, if “you created the content yourself, and used AI-based tools to edit, refine, error-check, or otherwise improve that content,” including using AI to brainstorm.
Three bodies, three distinct traditions, a similar conclusion: generation triggers disclosure; assistance does not. But all three leave the same case unsettled: the idea is yours, you worked it out by going back and forth with the model, and a lot of the words on the page came from the model.
Why the tools we are building miss this
The instruments being built now do not measure the thing that turns out to matter.
Anthropic began watermarking Claude’s output in August 2026, in response to the EU requirement. The company’s own documentation is unusually candid about the limits. A watermark “can only determine that Claude was likely involved with the content at some point. It cannot distinguish ‘Claude wrote this’ from ‘Claude heavily edited this.’” And: “The watermark only applies to words Claude chooses. When Claude proofreads text written by a person… there’s very little (if anything) for the watermark to attach to.”
Read those two sentences together. The mark is densest where the model generated freely and sparsest where a person did the thinking and used the model as a tool. It tags typing, not authorship. That is backwards from what almost anyone wants the signal to mean. Someone who talks through a half-formed idea and iterates until ninety percent of the words are their own gets marked. Someone who has a model produce a clean factual summary may not.
The test also only works in one direction. Finding a watermark tells you something. Not finding one tells you almost nothing. The text might be human. It might come from a different AI, or from a version of Claude released before August. It might be too short to carry the mark, or too factual. Or someone rewrote it.
LinkedIn’s contribution is a “Seems like AI slop” control, added in July 2026. What it actually does is feed the ranking system and train classifiers. It is not a report category under the platform’s policies, it does not label content, and it does not remove it. The company’s own product leadership has been explicit that “AI and slop are not the same thing” and that many people refine their thinking with AI. The target is mass-produced low-effort content. But the button’s label invites users to read it as a verdict on authorship, and users will.
What ratification would require
If the constraint is legitimacy and not detection, the work is different, and much less technical.
First, change the incentives before expecting people to be upfront. Tell people you used AI and it costs you. Get caught and it costs about twice as much. Say nothing and it costs the least. When those are the choices, people hide it. That is why there is a market for tools that strip the AI markers out of writing. As more tools emerge to catch it, others will emerge to circumvent them.
I saw this stated plainly in that same exchange. The person who had flagged my note described her own practice: she uses AI as a final spelling and grammar check, and otherwise limits herself to “private usage or just ways that I feel are reasonably undetectable.” She was not being evasive. She was being honest about a policy many careful people have settled into without ever discussing it, which is to use the tool where nobody can tell. That is what the math looks like from the inside. No norm will take hold while honesty is the expensive option.
The other thing I noticed: she edited her comment afterward so the negativity wouldn’t sit in my comment section. I appreciated it. But the retraction was private and the accusation had been public, which is the shape most of these take.
Careful wording does not solve it either. Schilke and Reimann tested six different ways of phrasing the disclosure, including “a human has reviewed and revised the work” and “AI was used only for proofreading.” All six reduced trust compared with saying nothing.
Second, show the difference between using AI and delegating the job to it. In a second Claessens study, someone who said plainly that they used AI as a tool, not as a replacement, was rated more moral and more trustworthy than someone who used no AI at all. Not just excused. Better. What people punish is handing over the whole task. Once they can see you didn’t, the penalty goes away.
That is the most actionable finding in this literature, and it aligns with the line that the EU, the Copyright Office, and Amazon already drew.
Third, say what a byline covers. This sounds bureaucratic. But this is the only item on this list an organization can settle on its own. A byline already covers a comms team, an editor, a research assistant, a fact-checker. Nobody wrote that rule down. We know what those roles do, and we know the person named on the piece is still responsible. Nobody has done that for AI. So everyone works it out on their own, in public, while being second-guessed.
The questions are concrete. Does our byline cover drafting? Structuring? Summarizing our own prior work? Who stays answerable when it is wrong? What has to be disclosed, and to whom? Does that disclosure go to readers, or to the institution, as it does at Amazon and at the Copyright Office?
Fourth, stop treating this as a detection problem. People identify AI-written text at roughly chance levels: 50 to 52 percent across six experiments with 4,600 participants (Jakesch, Hancock, and Naaman, PNAS, 2023). They also agree with each other about which texts look suspicious. When people agree with each other and are still wrong, they are all using the same bad heuristic. Researchers then tuned AI text to hit those cues. Readers judged it human 65.7 percent of the time. They judged actual human writing human only 51.7 percent of the time.
The heuristics are worse than useless. Grammatical errors read as human and were less likely to be AI. Contractions read as human and were more likely to be AI. First-person pronouns and family references, which people lean on heavily, carry no diagnostic signal at all.
So the defensive behaviors backfire. Deleting em-dashes, introducing typos, roughening your own prose: all of it optimizes against a model of detection that does not describe how people actually judge. It reads as not caring. Across eight studies with 5,306 participants, people who used texting abbreviations were rated less sincere, and the reason was that abbreviating made them look like they had put in less effort (Fang, Zhang, & Maglio, 2025).
And when these tools are wrong, they are wrong about the same people over and over. Researchers ran student essays through seven detectors. Essays by writers whose first language was not English were flagged as AI more than 60 percent of the time. Essays by American eighth-graders were almost never flagged (Liang et al., 2023).
The person who commented on my note told me her sister wrote a paper herself and had it flagged by her university’s AI checker. Everybody has one of these stories now. That is the system we are building.
The uncomfortable part
One finding underneath all of this reframes the anxiety, and it has nothing to do with AI.
Researchers had people pick out gifts and hand them to each other (Zhang & Epley, 2012). Half the givers were told to choose carefully. Half were told to choose at random. The receivers noticed the difference. But it made no difference at all with respect to how much they liked the gift. None. Givers felt closer to the people they had chosen carefully for. Receivers did not feel closer back. Givers were sure they would.
Effort counts most for the person expending it.
Hold that next to the experience of spending a week on something and being told it sounds like AI. Your effort was always clearer to you than to anyone reading it. What changed is not that readers stopped noticing care. It is that there is now a cheap explanation for why writing looks polished, and polish never carried as much information as we believed.
We have been treating the finished piece as evidence of the work behind it. It was never good evidence. It held up through a century of shortcuts. People cited a citation instead of reading the source, or skimmed an abstract instead of the study. Those shortcuts still cost real time. They were cheaper than the work, but not by much. So a polished piece still told you something, as long as faking one took nearly as much work as writing one.
That is no longer true, and no watermark will make it true again. That threshold has been crossed, and no watermark is going to uncross it.
What we are actually deciding
The question people think they are asking is whether a human wrote this.
The question they are actually asking is whether anyone is answerable for it, and whether the help behind it is the kind we have agreed to accept.
The first half we already know how to institutionalize. It is what a byline is. It is what editorial responsibility means in Article 50(4), what the Copyright Office tests when it asks whose expression is visible, what a reputation is for.
The second half is unsettled, and better detection will not settle it, because it was never a technical question. It is a question about which arrangements we are prepared to recognize, and we have answered it before: for ghostwriters, for editors, for research assistants, for every other form of help that now passes without comment.
It’s worth remembering when the demand is that contribution be separable, as it is when we question what percentage of a text that AI contributed. I have done my best work with other people, and the mark of the collaborations I’ve valued most is that at some point we lost the thread of who contributed what. The argument stopped being mine or theirs. The result was better than either of us would have produced alone, and neither of us could have drawn the line afterward if you’d asked.
Nobody audits those. We don’t ask co-authors to mark their sentences, or demand a coauthor’s paragraphs be shaded a different color, and if someone did we’d recognize it as a misunderstanding of what collaboration is for. Separability was never the standard for work we consider legitimate. It’s the standard we invented for work we haven’t decided about yet.
Separability was never the standard for work we consider legitimate. It’s the standard we invented for work we haven’t decided about yet.
We answered the earlier questions slowly, through repeated practice and visible sanction, over decades.
Whether that process can run at the speed the tools are changing is the open question. The Economist compared 1.2 million words of model output against its own archive and found the markers had already flipped: the models that once overused the em-dash now use fewer of them than professional writers do (”How to Spot AI Writing,” 2026). Norms form through repeated interaction. If the ground shifts faster than the interaction cycle, you don’t get new norms.
So the question is not how to spot AI writing. What does our byline cover? Who stays answerable when the work is wrong? And how do we say so out loud, before the next person has to guess?
You get an argument that never ends. That is what is in the comments. Not people disagreeing about the rules. People finding out there aren’t any yet.
References
Amazon Kindle Direct Publishing. (n.d.). Content guidelines: Artificial intelligence (AI) content. Retrieved August 17, 2026, from https://kdp.amazon.com/en_US/help/topic/G200672390
Anthropic. (2026, August 14). How Claude’s text watermark works. https://www.anthropic.com/news/claude-text-watermark
Claessens, S., Veitch, P., & Everett, J. A. C. (2026). Negative perceptions of outsourcing to artificial intelligence. Computers in Human Behavior, 177, Article 108894. https://doi.org/10.1016/j.chb.2025.108894
Fang, D., Zhang, Y. (E.), & Maglio, S. J. (2025). Shortcuts to insincerity: Texting abbreviations seem insincere and not worth answering. Journal of Experimental Psychology: General, 154(1), 39–57. https://doi.org/10.1037/xge0001684
How to spot AI writing. (2026, July 30). The Economist. https://www.economist.com/culture/2026/07/30/how-to-spot-ai-writing
Jago, A. S., & Carroll, G. R. (2024). Who made this? Algorithms and authorship credit. Personality and Social Psychology Bulletin, 50(5), 793–806. https://doi.org/10.1177/01461672221149815
Jakesch, M., Hancock, J. T., & Naaman, M. (2023). Human heuristics for AI-generated language are flawed. Proceedings of the National Academy of Sciences, 120(11), Article e2208839120. https://doi.org/10.1073/pnas.2208839120
Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), Article 100779. https://doi.org/10.1016/j.patter.2023.100779
Liu, B., Kang, J., & Wei, L. (2024). Artificial intelligence and perceived effort in relationship maintenance: Effects on relationship satisfaction and uncertainty. Journal of Social and Personal Relationships, 41(5), 1232–1252. https://doi.org/10.1177/02654075231189899
Perez, S. (2026, July 30). LinkedIn adds a button to report AI-generated ‘slop’. TechCrunch. https://techcrunch.com/2026/07/30/linkedin-adds-a-button-to-report-ai-generated-slop/
Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), 2024 O.J. (L 2024/1689). https://eur-lex.europa.eu/eli/reg/2024/1689/oj
Reif, J. A., Larrick, R. P., & Soll, J. B. (2025). Evidence of a social evaluation penalty for using AI. Proceedings of the National Academy of Sciences, 122(19), Article e2426766122. https://doi.org/10.1073/pnas.2426766122
Schilke, O., & Reimann, M. (2025). The transparency dilemma: How AI disclosure erodes trust. Organizational Behavior and Human Decision Processes, 188, Article 104405. https://doi.org/10.1016/j.obhdp.2025.104405
U.S. Copyright Office. (2025, January). Copyright and artificial intelligence, part 2: Copyrightability. https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf
Waters, S., & Nesbitt, C. K. (2026, February 10). Performance management in the age of algorithms series: Part II — How social systems, psychological contracts, and market tolerance shape algorithmic evaluation. Fractional Insights AI. https://fractionalinsightsai.substack.com/p/performance-management-in-the-age
Zhang, Y., & Epley, N. (2012). Exaggerated, mispredicted, and misplaced: When “it’s the thought that counts” in gift exchanges. Journal of Experimental Psychology: General, 141(4), 667–681. https://doi.org/10.1037/a0029223




