Product Launch2 min read

Cheap AI Reading Means Checking Everything, Not a Sample

By , Senior AI ConsultantPublished

Anthropic's new small model makes reading a thousand long reports cost about a dollar, which turns checking a sample into checking everything.

Anthropic has released Claude Haiku 5.5, the small model in its Claude family, built for high-volume routine work: summaries, sorting messages into categories, database lookups and live customer chat. For prompts up to 100,000 tokens, which covers most requests, it costs $0.10 per million input tokens and $0.50 per million output tokens. That is a tenth of what Haiku 4.5 charged. OpenAI's small model, GPT-6 Luna, has cost the same since September, so this is now the going price for small models from both companies.

A thousand documents for about a dollar

Prices per million tokens are hard to feel, so turn them into documents. A long report is about 5,000 tokens, and a short summary of it about 500. On Haiku 5.5 that is $0.0005 to read the report and $0.00025 to write the summary, so a thousand reports cost about a dollar. On Haiku 4.5 the same job cost about $7.50. The new model uses slightly more tokens for the same work, so the real bill is a little higher than that, but the order of magnitude holds.

At that price, cost stops deciding what gets read. Most organizations ration attention: a quality team listens to a few conversations a week, an auditor pulls a sample of invoices, a manager reads the summary of a summary. Nobody chose sampling because it is a good method. It is what people do when every item costs a person's time.

What it looks like in a small business

For example, a support lead at an online shop gets 5,000 customer messages a month. Now she reads the angry ones and a handful of others. With a small model, every message can be read, tagged as a refund request, a late delivery, a damaged item or a product question, and given a drafted reply. At about 1,500 tokens in and 300 out per message, that costs around $2 a month. At the end of the month she can see something sampling would have hidden, such as one product behind a fifth of the complaints.

She still makes the decisions that matter: which drafted replies go out as written, which go to a person, and what to do about that product. The same pattern fits a finance team that checks every invoice against its purchase order instead of one in twenty.

Where the cheap model stops

Haiku 5.5 is a cheap reader, not the best thinker. On a hard coding benchmark it scores 39%, against 71% for Sonnet 5.5, Anthropic's larger model. So the sensible setup has two tiers: Haiku reads everything and flags what looks odd, and a bigger model or a person handles the flagged cases.

Two limits change the bill. Prompts over 100,000 tokens cost five times as much, so long files are better split than fed in whole. And the best scores were measured at the maximum effort setting, which makes the model think longer and use more tokens; the default medium setting scores lower.

The work moves to the question

After Gmail, nobody got better at filing email. They got better at searching it, because keeping everything had stopped costing anything. Reading everything works the same way: the skill moves from deciding what deserves attention to deciding what to ask of all of it, and from reading to checking what comes back.

With two companies now at the same price, the cost of reading will keep falling, and a spot check will start to look the way an annual email purge looks now: a habit from when reading was expensive.

Share this

STAY INFORMED

Get AI intelligence like this delivered to your inbox.

Free forever · Unsubscribe anytime


You May Also Find Valuable