Skip to main content
← Blog
ArticlesBy Cameron Knight

Make AI Writing Sound Human: What the Research Says + Free Claude Skill

Surreal collage of a figure with a boxy robot head typing on a vintage typewriter in a Mediterranean wildflower field, with a screen of garbled code floating beside them

New research shows AI writing is identifiable from what it chooses to say, that expert readers reject it and that the gap closes once it carries a real writer's material. Here is what that means, and the free humanizer skill for Claude we built from it.

Most advice on fixing AI copy starts with a list of banned words. Take out "delve", swap "leverage" for "use", lose the em dashes and the copy is supposed to read as human.

New research says that barely touches the problem. In a study of more than 61,000 human and AI stories published this year, a classifier that never looked at word choice or sentence style still told the AI stories apart 93% of the time. When the AI stories were edited to strip out clichés and purple prose, it still found them. What gives AI writing away is less about the words and more about what the writing decides to say.

We use AI in our writing every day, from client emails to this blog, so we took that research and turned it into an editing skill. Below is what the studies found, what it changes if you write with ChatGPT or Claude, and the skill itself, which is free to download.

The words were only ever the surface

The vocabulary tells are real, and they are measurable. Dmitry Kobak and colleagues counted words across more than 15 million PubMed abstracts and watched a handful of them take off once ChatGPT arrived. In 2022, "delves" appeared in 0.07 of every 1,000 abstracts. By 2024 it appeared in 3.57, about 48 times as often. "Underscores" and "showcasing" rose about 14 times.

Bar chart showing how much more common six words were in PubMed abstracts in 2024 than in 2022: delves 48 times, showcasing and underscores 14 times, intricate 7 times, pivotal and notably 3 times

From those shifts they estimate that at least 13.5% of 2024 abstracts were processed with an LLM, and as many as 40% in some groups of papers.

So the banned-word lists aren't wrong. They just expire. Each model release brings new habits and drops old ones; the StoryScope authors point out that GPT-5.4 already uses far fewer em dashes than its predecessors. Chasing vocabulary means chasing a target that moves every few months, and it leaves the deeper pattern untouched.

The real difference is what gets said

StoryScope, published in 2026 by Jenna Russell, Mohit Iyyer and colleagues, is the study that makes this clear. They took 10,272 writing prompts, each drawn from a story a human had written, and gave the same prompts to five models: Claude, DeepSeek, Gemini, GPT and Kimi. That produced 61,608 stories of around 5,000 words. Then they described every story by 304 narrative features. Who drives the plot? Does time run in order? How much does the narrator explain?

Plotted on those features, the five models land almost on top of each other. The human stories sit somewhere else, and they are spread much wider.

Scatter plot from the StoryScope paper projecting stories onto narrative features: the five AI models cluster together while human-written stories occupy a separate, wider region

Figure 2 from Russell et al. (2026), StoryScope, arXiv:2604.03136, reproduced under CC0.

That separation is what the classifier picks up. Using narrative features alone it scored 93.2% (macro-F1) at telling human from AI. Using only style features, the sentence rhythm and figurative language most humanising advice targets, it scored 85.8%. The researchers then ran 278 Gemini stories through an editing framework built from professional writers' fixes for cliché, redundant exposition and purple prose. Detection moved from 95.5% to 93.9%.

Bar chart of classifier accuracy: style features only 85.8%, narrative features only 93.2%, narrative features after surface edits 93.9%, both together 96.0%

So what are the models doing differently? The paper answers in unusual detail.

Dot plot comparing AI and human stories on nine narrative choices, such as the narrator spelling out the theme in 77% of AI stories and 52% of human stories

The narrator spells out the theme in 77% of AI stories and 52% of human ones. AI shows emotion through the body (a tightening chest, cold sweat) in 81% of stories against 38% for people, while people simply name the feeling far more often, 29% to 8%. Most AI stories have no subplot at all. Humans name real books and authors twice as often, and address the reader directly four times as often. The authors sum it up as "over-determination: AI spells out meaning rather than trusting the reader to infer it."

Each model has its own accent too. Claude's stories have notably flat escalation, GPT leans on gossip to move the plot and Gemini describes characters from the outside.

This is fiction, and a services page isn't a short story. But we see the same habits in AI business copy constantly: a paragraph that ends by explaining why it mattered, every section given equal weight, a closing line that turns the page into a lesson, and nothing a competitor couldn't have published word for word. None of that shows up in a banned-word list.

Do readers actually care?

Some of them care a great deal.

Tuhin Chakrabarty, Jane Ginsburg and Paramveer Dhillon asked MFA-trained writers and three frontier models (ChatGPT, Claude and Gemini) to write passages of up to 450 words in the style of 50 award-winning authors. Then 28 MFA-trained readers and 516 general readers compared them blind.

Odds ratios: below 1 means readers preferred the human writer, above 1 the AI. Source: Chakrabarty, Ginsburg and Dhillon (2026), arXiv:2510.13939.
Who was judgingAI simply promptedAI fine-tuned on the author's books
Expert readers, quality0.131.87
Expert readers, faithful to the author0.168.16
General readers, quality1.825.42
General readers, faithful to the author1.0616.65

General readers were happy with ordinary prompted AI and even slightly preferred it. The people who read closely for a living were not, and rejected it by a wide margin. In business those close readers are often the ones deciding: the editor, the client, the buyer comparing three agencies' proposals.

The gap closed once the model had the writer's real material. Fine-tuned on each author's complete works, the AI was preferred by both groups, and detectors flagged it 3% of the time against 97% for the prompted version. The researchers' analysis traces the change to one cause: fine-tuning removed the quirks readers were penalising.

Most businesses won't fine-tune a model on their back catalogue. The practical version is simpler. Give the model what only you have, and edit out the quirks it adds.

Everyone gets better and everyone sounds the same

There's a cost that doesn't show up in any single piece of copy.

Anil Doshi and Oliver Hauser had 293 people write short stories, some with AI-generated ideas to start from. With five AI ideas, stories were judged 8.1% more novel. For the least creative writers the lift was larger: novelty up 10.7% and "well written" up as much as 26.6%. But the AI-assisted stories were measurably more similar to each other than the ones people wrote alone. Each writer did better, and the pool got narrower.

A 2026 follow-up found the same squeeze and one useful exception. When AI supplied the ideas, the diversity of what a group produced shrank. When AI only refined ideas the writers already had, the diversity survived.

For a business, sameness is the real cost. If your competitor asks the same model the same question, you both publish the same page.

What this means if you write with AI

Keep using AI. Just move the editing up a level, to what the copy says and how much of it.

  1. Bring what the model can't know. A real number, a client's own words, the approach that failed first. It's the one signal no model can fake, and it's what closed the gap in the fine-tuning study.
  2. Decide what matters before you prompt. Give the important point room and let the secondary ones go. AI treats every idea as equally worth a paragraph.
  3. Cut the explanation after the point lands. If a sentence tells the reader what the last one meant, delete it. That's the StoryScope finding as an editing rule.
  4. Leave real tensions open. When the evidence is mixed, say so and stop. A neat resolution reads as manufactured.
  5. Use AI to refine your ideas, not to supply them. That's the version of the workflow that kept writing distinct.

Swap the words last. Doing it first is why so much "humanised" copy still reads like a model.

The Human Copywriting skill

We turned all of this into a skill: a set of instructions an AI assistant loads when it writes or edits prose. It works in Claude, Claude Code, Cursor and as a ChatGPT project file.

It started as Cameron's ChatGPT project skill, built on Stephen Offer's human-voice project. For this version we added the pattern catalogue from Siqi Chen's humanizer, which traces back to Wikipedia's Signs of AI writing guide, and the research above. Both upstream projects are MIT licensed, and so is ours.

What makes it different from a banned-word list:

  • It edits in order. Intent, shape, substance and selectivity come first. Rhythm and syntax follow. Word choice is nearly last.
  • Tells are ranked by strength. A "not X but Y" contrast justifies an edit on sight. A single "crucial" doesn't, unless other tells keep it company.
  • Facts are locked. In a rewrite, numbers, names, links, caveats and the strength of every claim stay exactly as they were. It never invents a statistic or an anecdote to sound more human.
  • It guards against overcorrection. No fake typos, random slang or choppy fragments, and no chasing detector scores.
  • A small script flags the mechanical tells: dashes, stock phrases, contrast formulas, repeated openers and sentences stuck in the same length band.

We run a condensed version as the base layer for every email drafted in Conductor, our internal AI platform, with each founder's own style guide layered on top. This article went through the full version.

It won't make copy undetectable, and it doesn't try to. It makes copy read as if someone chose every sentence, which is the part people notice.

How it compares to other humanizer skills

The best-known is Siqi Chen's humanizer, with more than 52,000 stars on GitHub. It's very good at what it targets: 25 patterns ranked by strength, rewrites that keep every fact and voice matching from a sample of your writing. Our catalogue is built on it, so if sentence-level cleanup is all you need, install humanizer.

Ours adds the layer the research points to. Before it touches a phrase, it checks what the draft says: whether the important point gets the most room, whether a paragraph ends by explaining itself, whether anything in it could only have come from you. That's where StoryScope found the widest gap between AI and human writing, and pattern lists don't reach it. It also comes with a checker script for the mechanical tells and a single-file version for ChatGPT Projects.

Where it fits in our blog workflow

The skill decides how an article reads. What goes into it (research, evidence, structure and SEO) is the job of the SEO content workflow we published in August. That workflow now hands every draft to the Human Copywriting skill before its 90-point quality check, and the download includes it.

This article was the first to go through both: a research ledger for every figure, the quality gate and the skill.

If you'd like help setting up a writing workflow like this for your team, talk to us about AI consulting.

← All posts
  • #AI copywriting
  • #Human copywriting
  • #AI writing
  • #ChatGPT
  • #Claude
  • #Content quality
  • #AI research