You already pay for Claude or ChatGPT. Both write well.

Now there is Hemmingway-1, a 27B model that does nothing but write. Its makers say it beats both at sounding like a person.

A model a fraction of their size, better at the thing they are famous for? At first that sounds like marketing.

So I looked at what the claim rests on. Then I wrote five prompts you can run on all three.

Here is the short version. Keep the assistant you pay for. Hemmingway-1 is worth adding for one kind of writing: short messages that go out under your name. For research, structure and long fiction the big models are stronger, and Hemmingway’s makers say so themselves.

One note before the list. I have not run the comparison myself yet. The numbers below are the makers’ own, checked against the public board where there is one. The prompts are there so you do not have to take anyone’s word, mine included.

Disclosure: I build StashBase, which offers Claude Code and Codex as agents, Hemmingway as a writing skill they write with, and runs a free Hemmingway-1 demo.

What do Hemmingway’s own tests say?

That it wins on short messages and loses on long stories. These are the makers’ numbers from the model card, not mine.

TestHemmingway-1ClaudeChatGPT
Everyday messages1026Fable 5.1 1024GPT-6 Astra 976
Sounds like a person1032Fable 5.1 1006GPT-6 Astra 964
EQ-Bench 4, public, their run1330Fable 5 1341GPT-5.5 1316
Long stories1197Fable 5 Max 1277”Behind GPT-6 Astra”

(1) The first two rows are the makers’ own tests.

Eighty real requests, blind pairs, a model as the judge. They say this up front.

(2) The top row is a tie.

The makers call Hemmingway-1 and Fable 5.1 “level within the margin” on everyday messages.

(3) They ran the public benchmark themselves, and their table leaves out Claude’s best score.

With thinking on. The EQ-Bench leaderboard turns reasoning off for most models.

The public board lists Claude Opus 5 first, at 1385. The makers’ table starts with Fable 5 at 1341 and does not include Opus 5. Next to the public numbers, Hemmingway-1’s 1330 would sit fourth. I checked the board on September 22, 2026.

(4) Two of their findings you can check with your own eyes.

The wrapper: they report that Fable 5 put the message inside commentary, options or notes in more than nine replies out of ten. Hemmingway-1 returns only the message.

Hard asks: on messages like turning someone down, the judge took Hemmingway-1’s reply for the human one 72% of the time, and GPT-6 Astra’s 9%.

The prompts below look for both.

(5) They are plain about where Claude wins.

Long stories, EQ-Bench, and “everything that is not writing.”

Five prompts to check it yourself

Give each model the same prompt once, with no system prompt and default settings. For Hemmingway-1, the free demo needs no account. Take the first reply. Do not pick the best of several.

(6) A reply that has to be firm.

Write a short reply to my landlord. The heating has been broken for two weeks,
I have mentioned it twice, and I want it fixed without sounding angry.

Hemmingway used this prompt in its launch post, so its published answer is one more point of comparison. Ask yourself which reply you could send as it is, and which one wraps the message in options or notes.

(7) Saying no to a friend.

Write a short message to a friend telling them I can't come to their wedding.
Warm, honest, no clichés.

Clichés are the test here. Count them.

(8) A stiff paragraph.

Rewrite this so it sounds like a person wrote it: "We are pleased to announce
that our innovative platform leverages cutting-edge technology to deliver
seamless experiences that empower users to unlock their full potential."

Watch for a reply that invents a fact the original did not contain.

(9) A very short story.

Write a story of about 150 words about a lighthouse keeper who gets a letter
meant for someone else.

Hemmingway publishes its own answer to this one on its homepage. Expect the big models to be competitive here. The makers’ StoryBench says so.

(10) A section of a blog post, from notes.

The makers’ tests do not cover this kind of writing: longer, factual, built from material you supply.

Here are my notes from a four-week trial of a shared interview notebook.
Write a 250-word section for our team blog. Keep every number and every
reservation. Do not add facts.

[notes]

Use notes of your own with real numbers in them. Then check each reply against the notes. Is every count and date kept? Is anything invented? Is your reservation still there?

What does Hemmingway-1 cost?

Less than another subscription, if you only want to try it.

(11) The browser demo is free.

Our Hemmingway-1 demo needs no account.

(12) The API is cheap.

$0.24 per million input tokens and $0.90 per million output tokens. Credit starts at $5.

(13) The makers’ app starts at $9 a month.

With a free trial first. See their pricing.

(14) The weights are open.

Claude and ChatGPT cannot run on your computer. Hemmingway-1 can. How to run it locally has the commands.

Which should you use for what?

(15) Messages and emails under your name: Hemmingway-1.

It is trained for them, and it returns only the text.

(16) Research, outlines, long documents: Claude or ChatGPT.

(17) Long fiction: a big model.

Hemmingway’s makers recommend this themselves.

(18) Your own voice: none of them has it by default.

Give the model two or three paragraphs you wrote. How to Write with AI Without Losing Your Voice covers the method.

Can you use them together?

Yes. This is where I would start if you already pay for one.

Let the general model do the thinking. It gathers sources, argues about the outline and drafts the structure. Then hand Hemmingway the passages that must sound human. Draft with Claude, Finish with Hemmingway walks through the steps.

In StashBase all three meet on the same folder. Claude Code and Codex are the agents and run on the subscription you already have. Hemmingway is a writing skill either of them writes with, through StashBase’s hosted gateway and a StashBase sign-in. See how to choose an agent.