ScreenshotNeo

BlogEngineering

What Is AI Hallucination, and Can It Be Fixed?

AI hallucinations are confident, unsupported errors. Learn why they happen, what reduces them, and how to evaluate claims without mistaking fluency for fact.

By the ScreenshotNeo team1 October 20269 min read

AI hallucination is a confident output that is false, unsupported, internally inconsistent, or unrelated to the prompt. NIST uses confabulation for this behavior and notes that hallucination and fabrication are common names for it. Hallucinations can be reduced with evidence, retrieval, verification, calibrated uncertainty, abstention, and human review. Current evidence does not show a universal way to eliminate them.

Fluent wording is not proof. A model can invent a citation, merge two real facts, use outdated information, or answer an ambiguous question with an unjustified guess.

What is an AI hallucination?

NIST defines confabulation as a generative AI system confidently presenting erroneous or false content. The definition also covers responses that diverge from the prompt or contradict earlier statements in the same context. NIST explains that these behaviors are a natural consequence of how generative models are designed.

Large language models approximate patterns in their training data. During generation, they predict likely tokens in context. That process can produce accurate, consistent text, but it does not consult a built-in truth database or guarantee that each claim is correct. Open-ended prompts and questions requiring current, specialized, or contextual knowledge create more opportunities for error.

Do not label every fictional or creative output a hallucination. In a creative task, invention may be the intended result. The term matters when readers reasonably expect factual accuracy.

Why do language models make things up?

Prediction is different from verification

A model is optimized to continue a sequence plausibly. It may know patterns associated with a person, law, API, or historical event without having a reliable, item-by-item record of the facts. A plausible continuation can therefore contain a wrong date, nonexistent function, or fabricated source.

The prompt may not contain enough information

Some questions are ambiguous, unanswerable from the available context, or depend on information that changed after training. If the system is pushed to answer every question, it may fill the gap instead of asking for clarification or declining.

Evaluation can reward guessing

OpenAI’s analysis of hallucinations argues that evaluations focused mainly on accuracy can make guessing look better than admitting uncertainty. A system that says “I don’t know” receives no credit on some tests, while a confident guess has a chance of being marked correct. Better evaluations penalize confident errors and give credit for appropriate abstention.

Long contexts create consistency failures

In a long conversation or document, the model can lose track of which detail came from the user, which was inferred, and which was generated earlier. This produces contradictions, unsupported summaries, and answers that quietly change assumptions.

Common types of hallucination

Type Example Useful check
Factual error A wrong date, number, name, or definition Compare the claim with an authoritative source
Fabricated citation A paper, URL, quotation, or court case that does not exist Open the source and verify that it contains the cited claim
Prompt divergence An answer that ignores a required format or constraint Check every requirement against the output
Internal contradiction The answer gives two incompatible values Extract claims and compare them pairwise
Unsupported inference A conclusion that does not follow from supplied evidence Separate observed facts from assumptions and conclusions
Stale information Outdated pricing, API behavior, law, or personnel Check the publication date and current primary documentation

Can AI hallucinations be fixed?

They can be reduced, but “fixed” is too strong if it means guaranteed elimination. The practical goal is to lower the rate and impact of errors, detect uncertainty, and prevent unsupported output from reaching a consequential decision.

Mitigation effectiveness depends on the model, task, available tools, prompt, evidence quality, and evaluation method. Retrieval can provide relevant documents, but it does not prove that the answer follows from them. A model can misread a correct source or cite the wrong passage.

A practical hallucination-reduction workflow

  1. Define the claim. Break a broad request into atomic statements that can be checked independently.
  2. Supply authoritative evidence. Provide the relevant policy, specification, dataset, or primary source. For current facts, allow a controlled lookup.
  3. Require evidence mapping. Ask the system to attach each material claim to a quoted passage, document identifier, or source URL.
  4. Permit abstention. Instruct it to say that evidence is insufficient and ask a clarifying question when appropriate.
  5. Verify independently. Check citations, calculations, dates, and assumptions with deterministic code, a second source, or a qualified reviewer.
  6. Use human review for impact. Healthcare, finance, legal, safety, identity, and security decisions need review matched to the consequences of an error.
  7. Log and evaluate. Keep prompts, retrieved context, model version, tool state, output, and reviewer result so failures can be reproduced.

Grounding and retrieval

Grounding means giving the model relevant external context instead of relying only on learned parameters. OpenAI’s GPT-4 technical report lists additional context and human review among precautions for reliability-sensitive uses. Retrieval improves access to evidence, but rankers can return irrelevant passages and documents can be incomplete or outdated.

Tool use and current information

Browsing, database queries, calculators, and code execution can reduce errors caused by stale knowledge or arithmetic. Results still require validation: a tool can return the wrong record, a search can surface an unreliable page, and a model can misinterpret a correct response.

Uncertainty and abstention

Ask for a confidence rationale tied to evidence, not a made-up percentage. Define when the system must stop, request clarification, or escalate. OpenAI’s Model Spec guidance says it is better to indicate uncertainty or ask for clarification than provide confident information that may be incorrect.

Claim-level checking

Whole-response ratings hide partial failures. Extract claims, classify them as supported, contradicted, unverifiable, or irrelevant, and record whether the system abstained. This makes a long answer measurable and exposes one wrong detail that would otherwise be buried in fluent prose.

How to evaluate claims about hallucination reduction

Do not repeat a “reduced hallucinations” claim without the test context. Compare:

  • the unit counted: claims, answers, or complete responses;
  • the error definition and severity;
  • abstention and refusal rates;
  • whether browsing or retrieval was enabled;
  • task type: short factual questions, open-ended writing, biography, or domain-specific work;
  • the grader: fixed reference, model grader, expert, or claim-by-claim review;
  • model version, prompt, date, and comparator.

OpenAI’s SimpleQA contains 4,326 short-answer questions designed to have one indisputable answer. Its dataset-development work estimated approximately 3% inherent error after additional review; that estimate applies to the dataset process, not to all AI outputs.

OpenAI’s 2025 explainer reports one SimpleQA comparison in which gpt-5-thinking-mini had 52% abstention, 22% accuracy, and 26% error, while o4-mini had 1% abstention, 24% accuracy, and 75% error. Those figures describe that model pair and test setup; they are not general real-world rates.

The GPT-5 system card reports claim-level evaluations in which GPT-5 main had a 26% smaller hallucination rate than GPT-4o and GPT-5 thinking had a 65% smaller rate than o3, under the specified prompts, grader, and comparison. It also reports 75% human agreement with the factuality grader. These numbers do not guarantee correctness for an individual answer.

Verification code patterns

For production systems, keep the model responsible for drafting and use deterministic code for checks the computer can perform. This Python example rejects a response when required fields are missing and records evidence identifiers for human review:

from dataclasses import dataclass

@dataclass
class Claim:
    text: str
    evidence_ids: list[str]


def validate_claims(claims: list[Claim], required_ids: set[str]) -> list[str]:
    errors = []
    for i, claim in enumerate(claims, 1):
        if not claim.text.strip():
            errors.append(f"claim {i}: empty text")
        if not set(claim.evidence_ids) & required_ids:
            errors.append(f"claim {i}: no approved evidence")
    return errors

claims = [
    Claim("The policy requires two reviewers.", ["policy-17", "policy-18"]),
    Claim("The deadline is Friday.", ["notes-4"]),
]
errors = validate_claims(claims, {"policy-17", "policy-18"})
if errors:
    raise ValueError("\n".join(errors))
print("Claims passed structural checks; verify meaning with a reviewer.")

This check verifies structure and approved evidence identifiers; it cannot determine whether a passage truly supports the claim. Use source-specific parsers, tests, and expert review for semantic verification.

Troubleshooting hallucination controls

Symptom Likely cause Fix
The answer cites pages that do not exist The model was asked for citations without source access Provide an allowlisted corpus or browsing tool and require links to be opened
Citations exist but do not support the sentence Evidence was retrieved but not mapped claim by claim Require passage IDs and have a separate checker compare claim and passage
The model refuses too often Abstention threshold or prompt is overly strict Define answerable scope, provide better context, and measure useful abstentions separately
Answers are accurate only on short questions Long-form generation introduces more claims and consistency risk Generate sections separately, validate each claim, then assemble
Results changed after a model update Model version, system prompt, retrieval, or tool behavior changed Pin versions where possible and rerun a regression set before release
Current facts are wrong Training data is stale or retrieval returned an old page Use dated primary sources, freshness filters, and an explicit as-of date

Performance, reliability, and cost considerations

Every extra retrieval, verification, or human-review step adds latency and operational cost. A staged design keeps routine requests fast: validate format first, retrieve only when a claim needs evidence, and escalate high-impact or low-confidence cases. Cache stable documents with their version and retrieval date, but never use a cache when freshness is part of the requirement.

Reliability improves when you log the complete execution path: model and prompt versions, retrieved passages, tool responses, abstentions, and final decisions. Monitor false positives as well as false negatives; a system that refuses everything can look safe while failing its users.

For budgeting, report cost per completed answer and cost per verified claim, not only tokens. Include retrieval, tool calls, retries, reviewers, and failed attempts. A lower error rate may be worth more than a lower per-request price when mistakes have material consequences.

When AI agents need visual evidence

Text retrieval cannot tell an agent whether a page actually rendered, whether a consent banner covers the content, or whether a responsive layout broke. ScreenshotNeo supplies visual evidence through a website screenshot API and MCP server. It can capture PNG, JPEG, WebP, or PDF output after handling consent banners and removing more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

Use the API when an evaluator or agent needs to inspect the rendered page:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for the full option set.

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Other controls include full-page and selector capture, device presets, dark mode, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, geolocation, caching, signed links, asynchronous jobs, bulk capture, and usage reporting.

Start with 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Are hallucinations unique to large language models?

No. The term is used broadly for generative systems that confidently produce false or unsupported content, including systems that generate images, audio, or video. The risk and evaluation method depend on the modality and task.

Does asking a model to “be accurate” prevent hallucinations?

It may change behavior slightly, but it does not provide missing evidence or a verification mechanism. Grounding, abstention rules, and independent checks are stronger controls.

Is a citation enough to trust an answer?

No. Open the source, confirm it is authoritative and current, and check that it supports the exact claim rather than a nearby statement.

Should a second AI review the first AI?

A second model can find some errors, but it can repeat the same mistake or invent a plausible critique. Combine automated review with deterministic checks and human review where the consequences justify it.

What is the safest default for an unknown question?

Ask for clarification or state that the available evidence is insufficient. A calibrated abstention is safer than a confident guess.

Primary sources