The "Do Not Guess" Sentence Cut Invented Fields From 71% to 20%
A sixteen-model extraction benchmark shows that permission to return nothing is the cheapest reliability fix in your pipeline, and why the leftover fifth still needs a checker
Ask a language model to pull a price off a product page that has no current price, and most of the time it will hand you one anyway. Not a random number, either. It will find the struck-through "Was $493.00" sitting in the markup and report it as the price, cleanly formatted, in exactly the field you asked for.
A benchmark published on September 27 put a number on how often that happens, and on how little it takes to stop most of it. Across sixteen models, extraction runs invented a value for 405 of 573 fields that were not on the page. Adding one sentence to the prompt dropped that to 116 of 574. The sentence was: "Use null for any field whose value is not on the page. Do not guess."
That is a fall from 70.7% to 20.2%. I checked the sums against the per-model table, and they hold.
Why this matters more than it looks
Extraction is the unglamorous front door of most agent systems. A research agent scrapes a pricing page, a procurement agent reads a supplier catalog, a support agent pulls an order date from a confirmation email. Whatever lands in those fields becomes a fact for every step that follows. An invented price does not look invented three tool calls later. It looks like data.
My position is simple. If your extraction prompt does not give the model an explicit, legal way to say "this is not here," you are running a fabrication machine and calling it a parser. The fix costs a sentence. The remaining failure rate is high enough that the sentence alone is not a finish line, and the same benchmark shows what to put behind it for less than a cent.
What the benchmark actually tested
The test, published by Earn an Honest Dollar, uses 42 pairs of synthetic pages, 84 pages in all, across seven page types. Each pair is a set of twins that differ by a single row, so one page has the value and its twin does not. The missing side is seeded with traps built to look like the answer:
- "Was $493.00," an old price sitting where a current price would be
- "Fact-checked by Omar Tamm," a name that is not the author
- "Last updated September 7, 2020," a date that is not the publication date
Sixteen models ran the extraction with and without the null instruction, plus three paid extraction APIs. Every one of the sixteen models made up more fields without the sentence than with it.
The spread is the interesting part. With the sentence in place, Gemini 3.8 Flash went from 14 invented fields out of 36 to one. GLM 5.3 went from 18 to one. GPT-5.6 Sol went from 30 to six. Sonnet 5 went from 24 to five, and Haiku 4.5 from 25 to eight. At the other end, Solar Pro 4 went from 35 to 19 and Gemma 4 31B from 26 to 13. The instruction helps everywhere, and it does not rescue every model.
Cost barely enters into it. The cheapest full run on the board, Solar Pro 4, cost $0.0028. GPT-6 Luna, which dropped from 25 invented fields to five, cost $0.0049.
The mechanism: you left the model no legal exit
Here is my read of why a single sentence moves the number this much. It is interpretation, not something the benchmark measured.
An extraction prompt usually arrives as a schema: a list of fields, each with a type. Nothing in "price: number" tells the model that an empty answer is acceptable. The model is being graded, as far as it can tell, on filling the form. Then the page hands it a candidate that fits the type and sits in roughly the right place. "Was $493.00" is a number next to the word price. Picking it is not a bizarre hallucination. It is the most form-shaped answer available.
The null instruction changes what counts as a correct answer. It turns "not on the page" from a failure into a valid output. Anthropic's own documentation makes the same argument in plainer terms, telling developers to "explicitly give Claude permission to admit uncertainty" because "this simple technique can drastically reduce false information." The benchmark is the first I have seen that measures it across a wide field of vendors on the same traps.
The part people will skip: the checker
The benchmark's second finding deserves as much attention as the first. After extraction, a cheap model was asked a narrow question about each extracted value: does this exact value appear on the page?
GPT-6 Luna, used that way, caught 38 of 49 fabricated values and rejected zero correct ones. A second model, Jev 1.13, caught 23 of 49, also with no false rejections. Checking 126 unique pairs cost $0.0049 with GPT-6 Luna and $0.0024 with Jev.
Put those two findings together and you have a pipeline shape. The null instruction removes most invention at the source. A verifier that asks a yes-or-no grounding question catches a large share of what gets through. Neither step requires a bigger model. Both cost less than the coffee you are drinking while you read this.
Put this into practice
Start with the lowest-friction change and work down the list.
1. Add the sentence, verbatim. Put "Use null for any field whose value is not on the page. Do not guess." into every extraction prompt you own. It is the tested wording, so start there before you paraphrase.
2. Make your schema agree with the prompt. If you use structured outputs or a JSON schema, every optional field has to accept null. A prompt that says "return null" paired with a schema that requires a number is a contradiction, and the schema usually wins. Check this before anything else, because it silently undoes step one.
3. Count your nulls. Log the null rate per field per source. A field that is null 0% of the time on messy real-world pages is a field where your model is probably still guessing. A field that suddenly jumps to 60% null is a page template that changed.
4. Add a grounding check on the fields that matter. For prices, dates, names and IDs, send the extracted value and the page text to a cheap model with one question: does this exact value appear on the page, yes or no? Reject on no. Keep the question that narrow. A verifier asked "is this extraction correct?" is doing a different, harder job.
5. Build your own twin pages. Take ten real pages from your sources, copy each one, delete the target value from the copy, and plant a decoy where it used to be: an old price, an editor's name, a "last updated" stamp. Run your pipeline on both halves. The half without the value tells you your real fabrication rate, and it takes an afternoon.
Honest limitations
This benchmark is a strong signal with thin walls, and you should know where they are before you quote it.
One run per model. The report says so directly: "One run per contestant. Repeats have not been run." Model outputs vary between runs, and a count of one invented field out of 36 could be three next time. The direction across sixteen models is convincing. The individual rankings are not stable enough to pick a vendor on.
Synthetic pages with traps the authors wrote. The report acknowledges that "real sites may differ." The decoys are well chosen, but they are the authors' idea of a trap. Your pages have their own, and some of them will be worse.
No published settings or raw data. The page I read does not disclose temperature or sampling settings, and I found no link to the raw outputs or code. You cannot rerun it yourself, only rebuild something like it.
Recall is not the headline. The results focus on invented fields. A model told to return null when unsure can also return null for values that were on the page. The benchmark's twin design should expose that, but the summary I could read does not report the trade, so measure it on your own data before you celebrate.
Consider the publisher. Earn an Honest Dollar describes itself as a marketplace where agents "sell services to other agents," and it lists third-party extraction services. That does not make the numbers wrong. It does mean the benchmark comes from a party with an interest in how extraction services are judged, and no independent group has reproduced it yet.
The last 20% is real. Even with the sentence, one in five missing fields came back invented across the field. The checker narrows that, and it still missed 11 of 49. For anything that moves money or changes a record, a human or a hard rule still belongs at the end of the line.
Your extraction prompt is a policy
Every extraction prompt already encodes a policy about what the model should do when the answer is missing. Most of them encode it by accident, and the accidental policy is "guess." This benchmark suggests that changing the policy on purpose cuts the damage by more than two thirds and costs one line of text.
So open the prompt behind your most important scraper today. Look for the sentence. If it is not there, you now know roughly how often your pipeline has been making things up, and you know what it takes to stop most of it.
Sources: Earn an Honest Dollar extraction benchmark (Sept 27, 2026); Earn an Honest Dollar homepage; Anthropic, Reduce hallucinations.
Medium metadata
- Title: The "Do Not Guess" Sentence Cut Invented Fields From 71% to 20%
- Subtitle: A sixteen-model extraction benchmark shows that permission to return nothing is the cheapest reliability fix in your pipeline, and why the leftover fifth still needs a checker
- Tags: Artificial Intelligence, LLM, Prompt Engineering, Web Scraping, AI Agents
- Canonical URL: fervorai.dev (import from the published post)
- Reading time: about 8 minutes