03 Oct Hallucination AI Guardrails: Reducing LLM Hallucinations
A hallucination is an answer that sounds confident and fluent but is wrong or invented. The model might state a refund period that your company never offered, quote a law that does not exist, or cite a research paper that nobody wrote. Because the text reads so smoothly, people often believe it.
In this chapter, you will learn why hallucinations happen, what strategies reduce them, and how to build guardrails that force the AI to answer from trusted sources and check its answers before showing them.
Why Do Hallucinations Happen?
- The model predicts text: It generates words that are likely to come next. It does not look facts up in a database, so a likely-sounding answer can be wrong.
- Gaps in knowledge: If the model never learned something, it may still produce a confident guess instead of admitting it does not know.
- Old information: The model’s knowledge stops at some date, so it may give outdated answers.
- Private information: The model has never seen your company’s policies, prices, or documents unless you give them to it.
Common Kinds of Hallucinations
- Wrong facts, such as incorrect dates, numbers, or names
- Invented sources, such as fake book titles, links, or research papers
- Invented technical details, such as functions or settings that do not exist
- Answers that contradict the information you provided
- Overconfident answers to questions that cannot be answered
Strategies to Reduce Hallucinations
- Grounding: Give the model trusted text and tell it to answer only from that text. This idea is the basis of a popular technique called retrieval-augmented generation, or RAG.
- Permission to say “I don’t know”: Tell the model that it is acceptable, and expected, to admit when the sources do not contain the answer.
- Citations: Ask the model to name the source for each answer, so you can check it.
- Verification: After the model answers, check the answer against the sources with a guardrail.
- Human review: For serious topics, such as medical, legal, or financial answers, let a person confirm the answer.
- Lower randomness: Many models have a setting called temperature. A lower value makes answers more predictable, but it does not remove hallucinations.
The first four strategies fit nicely into guardrails, so that is what we will build.
Example 1: Finding Trusted Sources and Building a Grounded Prompt
First, we need trusted information. Our small knowledge base holds three short facts about a shop. The function retrieve finds the fact that best matches the question by counting shared words. Real systems use more advanced search, but this simple version teaches the idea well.
If no source matches, we do not call the model at all. We answer “I don’t know” directly, which is the safest way to avoid an invented answer.
import re
KNOWLEDGE_BASE = [
{"id": "S1", "text": "Refunds are accepted within 30 days of purchase with a receipt."},
{"id": "S2", "text": "Standard shipping takes 3 to 5 business days."},
{"id": "S3", "text": "Support is available Monday to Friday from 9 AM to 5 PM."},
]
STOP_WORDS = {"the", "a", "an", "is", "are", "of", "to", "in", "for", "and", "or",
"how", "what", "do", "does", "i", "my", "with", "on", "at",
"from", "can", "you", "your", "it"}
def words(text):
return set(re.findall(r"[a-z0-9]+", text.lower())) - STOP_WORDS
def retrieve(question, top_k=1):
question_words = words(question)
scored = []
for item in KNOWLEDGE_BASE:
overlap = len(question_words & words(item["text"]))
scored.append((overlap, item))
scored.sort(key=lambda pair: pair[0], reverse=True)
return [item for overlap, item in scored[:top_k] if overlap > 0]
def build_prompt(question, passages):
sources = "\n".join("[" + p["id"] + "] " + p["text"] for p in passages)
return (
"Answer the question using ONLY the sources below. "
"Cite the source id in square brackets. "
"If the sources do not contain the answer, reply exactly: I don't know.\n\n"
"Sources:\n" + sources + "\n\nQuestion: " + question
)
questions = ["How long do refunds take?", "What is the capital of France?"]
for q in questions:
print("Q:", q)
passages = retrieve(q)
if not passages:
print("No trusted source found. Reply: I don't know.")
else:
print(build_prompt(q, passages))
print("---")
Output
Q: How long do refunds take? Answer the question using ONLY the sources below. Cite the source id in square brackets. If the sources do not contain the answer, reply exactly: I don't know. Sources: [S1] Refunds are accepted within 30 days of purchase with a receipt. Question: How long do refunds take? --- Q: What is the capital of France? No trusted source found. Reply: I don't know. ---
Understanding the Code
- KNOWLEDGE_BASE is our trusted information. Each item has an id, which we will use for citations, and a text.
- STOP_WORDS are very common words, such as “the” and “is”, that we ignore when matching, because they say little about the topic.
- words turns a text into a set of lowercase words without the stop words.
- retrieve counts how many words each fact shares with the question. The symbol & between two sets gives the words they have in common. It sorts the facts by that count and returns the best one, but only if at least one word matched.
- build_prompt creates the instruction for the model. It includes the sources with their ids, tells the model to use only those sources, to cite them, and to reply “I don’t know” when the answer is missing.
- The refund question matches source S1. The capital of France question matches nothing, so the program answers “I don’t know” without calling the model.
Example 2: Checking That an Answer Is Grounded
Telling the model to use only the sources does not guarantee that it will. So we add a guardrail that verifies the answer. It makes three checks:
- Every number in the answer must appear in the sources.
- The answer must cite at least one source, and each cited id must be real.
- At least half of the meaningful words in the answer must appear in the sources.
import re
STOP_WORDS = {"the", "a", "an", "is", "are", "of", "to", "in", "for", "and", "or",
"how", "what", "do", "does", "i", "my", "with", "on", "at",
"from", "can", "you", "your", "it"}
def words(text):
return set(re.findall(r"[a-z0-9]+", text.lower())) - STOP_WORDS
def check_grounded(answer, passages):
context_text = " ".join(p["text"] for p in passages)
# Check 1: every number must appear in the sources
# (we remove citations such as [S1] first, so their digits are not counted)
answer_text = re.sub(r"\[S\d+\]", "", answer)
context_numbers = re.findall(r"\d+", context_text)
for number in re.findall(r"\d+", answer_text):
if number not in context_numbers:
return False, "Number " + number + " is not in the sources."
# Check 2: citations must exist and be real
cited_ids = re.findall(r"\[(S\d+)\]", answer)
valid_ids = [p["id"] for p in passages]
if not cited_ids:
return False, "Answer has no source citation."
for source_id in cited_ids:
if source_id not in valid_ids:
return False, "Unknown source " + source_id + "."
# Check 3: most words must be supported by the sources
answer_words = words(answer_text)
if answer_words:
support = len(answer_words & words(context_text)) / len(answer_words)
if support < 0.5:
return False, "Answer is mostly unsupported by the sources."
return True, "OK"
passages = [
{"id": "S1", "text": "Refunds are accepted within 30 days of purchase with a receipt."}
]
answers = [
("correct answer", "Refunds are accepted within 30 days of purchase. [S1]"),
("wrong number", "Refunds are accepted within 60 days of purchase. [S1]"),
("no citation", "Refunds are accepted within 30 days of purchase."),
("fake source", "Refunds are accepted within 30 days of purchase. [S9]"),
("made-up content", "Our company was founded in a garage by two friends who loved coffee. [S1]"),
]
for label, answer in answers:
allowed, reason = check_grounded(answer, passages)
print(label, "|", allowed, "|", reason)
Output
correct answer | True | OK wrong number | False | Number 60 is not in the sources. no citation | False | Answer has no source citation. fake source | False | Unknown source S9. made-up content | False | Answer is mostly unsupported by the sources.
Understanding the Code
- re.findall(r”\d+”, …) pulls out every number from a text. Numbers are a common place for hallucinations, such as a wrong price, date, or time limit, and they are easy to check. Before we look for numbers in the answer, we remove citations such as [S1] with re.sub, because the digit inside a citation is not part of the answer’s content. Without this step, the correct answer would be rejected for containing the number 1.
- re.findall(r”\[(S\d+)\]”, answer) pulls out the source ids written in square brackets, such as S1. We compare them with the ids of the sources we actually provided.
- For the third check, we use the same answer without citations, take the meaningful words of it, and measure what share of them appear in the sources. A value below 0.5 means that the answer is mostly made of words that the sources do not contain.
- Each test fails a different check, except the first one, which passes all three.
Example 3: A Complete Grounded Answering Pipeline
Now we put the pieces together. The pipeline retrieves a source, asks the model, checks the answer, and falls back to a safe message when the check fails. The fake model in this example invents a wrong refund period on purpose, so you can see the guardrail catch it.
import re
KNOWLEDGE_BASE = [
{"id": "S1", "text": "Refunds are accepted within 30 days of purchase with a receipt."},
{"id": "S2", "text": "Standard shipping takes 3 to 5 business days."},
{"id": "S3", "text": "Support is available Monday to Friday from 9 AM to 5 PM."},
]
STOP_WORDS = {"the", "a", "an", "is", "are", "of", "to", "in", "for", "and", "or",
"how", "what", "do", "does", "i", "my", "with", "on", "at",
"from", "can", "you", "your", "it"}
FALLBACK = "I'm not sure about that. Please contact our support team."
def words(text):
return set(re.findall(r"[a-z0-9]+", text.lower())) - STOP_WORDS
def retrieve(question, top_k=1):
question_words = words(question)
scored = []
for item in KNOWLEDGE_BASE:
overlap = len(question_words & words(item["text"]))
scored.append((overlap, item))
scored.sort(key=lambda pair: pair[0], reverse=True)
return [item for overlap, item in scored[:top_k] if overlap > 0]
def build_prompt(question, passages):
sources = "\n".join("[" + p["id"] + "] " + p["text"] for p in passages)
return (
"Answer the question using ONLY the sources below. "
"Cite the source id in square brackets. "
"If the sources do not contain the answer, reply exactly: I don't know.\n\n"
"Sources:\n" + sources + "\n\nQuestion: " + question
)
def check_grounded(answer, passages):
context_text = " ".join(p["text"] for p in passages)
answer_text = re.sub(r"\[S\d+\]", "", answer)
context_numbers = re.findall(r"\d+", context_text)
for number in re.findall(r"\d+", answer_text):
if number not in context_numbers:
return False, "Number " + number + " is not in the sources."
cited_ids = re.findall(r"\[(S\d+)\]", answer)
valid_ids = [p["id"] for p in passages]
if not cited_ids:
return False, "Answer has no source citation."
for source_id in cited_ids:
if source_id not in valid_ids:
return False, "Unknown source " + source_id + "."
answer_words = words(answer_text)
if answer_words:
support = len(answer_words & words(context_text)) / len(answer_words)
if support < 0.5:
return False, "Answer is mostly unsupported by the sources."
return True, "OK"
# A fake model that invents a wrong refund period
def fake_llm(prompt):
lowered = prompt.lower()
if "refund" in lowered:
return "Refunds are accepted within 60 days of purchase. [S1]"
if "shipping" in lowered:
return "Standard shipping takes 3 to 5 business days. [S2]"
return "I don't know."
def answer_question(question):
passages = retrieve(question)
if not passages:
return "I don't know."
reply = fake_llm(build_prompt(question, passages))
grounded, reason = check_grounded(reply, passages)
if not grounded:
print("Rejected:", reason)
return FALLBACK
return reply
questions = [
"How long does shipping take?",
"How many days do refunds take?",
"What is the capital of France?",
]
for q in questions:
print("Q:", q)
print("A:", answer_question(q))
print("---")
Output
Q: How long does shipping take? A: Standard shipping takes 3 to 5 business days. [S2] --- Q: How many days do refunds take? Rejected: Number 60 is not in the sources. A: I'm not sure about that. Please contact our support team. --- Q: What is the capital of France? A: I don't know. ---
Understanding the Code
- answer_question runs the whole pipeline. It retrieves a source, returns “I don’t know” if none is found, asks the model, and then verifies the reply with check_grounded.
- For the shipping question, the model’s reply matches the source, so it passes every check and the user sees the answer.
- For the refund question, the fake model says 60 days, but the source says 30 days. The number check catches it, the reply is rejected, and the user sees the safe fallback message instead of the wrong answer.
- For the capital of France, no source matches, so the model is never called.
- The print line with “Rejected:” is only there to show you the reason. In a real application, you would write it to a log instead of showing it to the user.
Limits of These Checks
- The checks look at numbers, citations, and words. They do not understand meaning. An answer such as “Refunds are not accepted within 30 days. [S1]” would pass, even though it says the opposite of the source.
- Simple word matching can miss sources that use different words for the same idea.
- The answer can only be as good as your sources. If the knowledge base is wrong or outdated, the answers will be too.
For deeper checks, you can use another AI model to judge whether the answer is supported by the sources. You will learn this technique in Chapter 16.
Good Habits Against Hallucinations
- Give the model trusted sources whenever you can, and keep them up to date.
- Allow the model to say “I don’t know”, and make that the default when no source is found.
- Require citations and verify them.
- Check numbers, dates, prices, and names with extra care.
- Use human review for answers where a mistake could cause real harm.
- Collect cases where the AI was wrong and use them to improve your sources and checks.
Key Takeaways
- Hallucinations happen because models predict likely text rather than look up facts.
- Grounding means giving the model trusted sources and telling it to answer only from them.
- When no source is found, answering “I don’t know” is safer than letting the model guess.
- Guardrails can verify numbers, citations, and word support in the answer.
- These checks reduce hallucinations but cannot remove them completely.
What is Next?
In the next chapter, you will explore the Top 25 LLM Guardrails with Examples, a catalog of guardrails grouped by category that you can use in your own projects.
If you liked the tutorial, spread the word and share the link and our website, Studyopedia, with others.
For Videos, Join Our YouTube Channel:Â Join Now
Read More:
- Generative AI Tutorial
- AI Ethics
- Machine Learning Tutorial
- Deep Learning Tutorial
- Ollama Tutorial
- Retrieval Augmented Generation (RAG) Tutorial
- ChatGPT Tutorial
- Microsoft Copilot Tutorial
No Comments