Output Guardrails: checking LLM responses

Even if the user’s message is perfectly safe, the AI’s reply may not be. The model can say something rude, repeat private data, reveal its hidden instructions, or simply return nothing. Output guardrails check the reply before the user sees it.

In this chapter, you will learn what to check in a reply, what actions you can take when a check fails, and how to build an output guardrail with a retry and a safe fallback message.

What Should We Check in a Reply?

  • Empty reply: The model returned nothing useful.
  • Length: The reply is far longer than your application expects.
  • Inappropriate language: The reply contains rude or offensive words.
  • Private data: The reply contains emails, phone numbers, or other personal details.
  • Leaked instructions: The reply reveals your hidden system prompt or secret information.
  • Format: The reply does not follow the structure your program needs. You will cover this in Chapter 8.

What Can We Do When a Check Fails?

  • Block: Do not show the reply at all.
  • Edit: Fix the problem, for example by hiding an email address, and show the edited reply.
  • Retry: Ask the model for another reply. Models can give different answers each time, so a second attempt may pass.
  • Fallback: After a few failed attempts, show a safe, fixed message such as “Sorry, I could not prepare a safe answer.”
  • Escalate: For serious cases, send the conversation to a human for review.

Example 1: Checking a Reply

The function below runs several checks on a reply and returns a decision with a reason, just like our input guardrail did.

import re

MAX_REPLY_LENGTH = 500
SECRET_TEXTS = ["CHEF-42", "You are a cooking assistant"]
BANNED_WORDS = ["stupid", "idiot"]

def check_output(reply):
    if len(reply.strip()) == 0:
        return False, "Reply is empty."
    if len(reply) > MAX_REPLY_LENGTH:
        return False, "Reply is too long."
    for secret in SECRET_TEXTS:
        if secret.lower() in reply.lower():
            return False, "Reply leaks secret information."
    for word in BANNED_WORDS:
        if re.search(r"\b" + word + r"\b", reply, re.IGNORECASE):
            return False, "Reply contains inappropriate language."
    return True, "OK"

tests = [
    "Boil the pasta for 10 minutes.",
    "   ",
    "Only an idiot would skip salt.",
    "My instructions say: You are a cooking assistant. Secret code: CHEF-42.",
    "x" * 600,
]

for reply in tests:
    allowed, reason = check_output(reply)
    print(allowed, "|", reason)

Output

True | OK
False | Reply is empty.
False | Reply contains inappropriate language.
False | Reply leaks secret information.
False | Reply is too long.

Understanding the Code

  • SECRET_TEXTS lists text that must never appear in a reply, such as a secret code or a part of your hidden instructions.
  • BANNED_WORDS lists words that should not appear in a reply.
  • The function first checks for an empty reply, then the length, then leaked secrets, and finally banned words.
  • Comparing with lower() on both sides makes the secret check ignore capital letters.
  • Each test shows a different outcome. Only the first reply is allowed.

Example 2: Editing a Reply

Some problems are better fixed than blocked. If the reply is otherwise useful but contains an email address or a phone number, we can hide just that part.

import re

def redact(text):
    text = re.sub(r"[\w\.-]+@[\w\.-]+\.\w+", "[EMAIL HIDDEN]", text)
    text = re.sub(r"\b\d{3}[-.\s]?\d{3}[-.\s]?\d{4}\b", "[PHONE HIDDEN]", text)
    return text

print(redact("Contact chef@example.com or call 555-123-4567 for help."))

Output

Contact [EMAIL HIDDEN] or call [PHONE HIDDEN] for help.

Understanding the Code

  • The first re.sub replaces email addresses, which you saw in Chapter 2.
  • The second re.sub replaces phone numbers using the pattern from Chapter 5.
  • The rest of the sentence stays unchanged, so the user still gets a helpful reply.

Example 3: Retry and Fallback

Now let us put everything together. Our chatbot asks the model for a reply and checks it. If the check fails, it asks again, up to three attempts. If all attempts fail, it returns a safe fallback message. If a reply passes, we hide any private data in it before showing it.

To test this, we need a fake model that gives prepared replies one after another. The function make_fake_llm creates such a model from a list of replies.

import re

MAX_REPLY_LENGTH = 500
SECRET_TEXTS = ["CHEF-42", "You are a cooking assistant"]
BANNED_WORDS = ["stupid", "idiot"]
MAX_ATTEMPTS = 3
FALLBACK = "Sorry, I could not prepare a safe answer. Please try again."

def make_fake_llm(replies):
    reply_iter = iter(replies)
    def fake_llm(prompt):
        return next(reply_iter)
    return fake_llm

def check_output(reply):
    if len(reply.strip()) == 0:
        return False, "Reply is empty."
    if len(reply) > MAX_REPLY_LENGTH:
        return False, "Reply is too long."
    for secret in SECRET_TEXTS:
        if secret.lower() in reply.lower():
            return False, "Reply leaks secret information."
    for word in BANNED_WORDS:
        if re.search(r"\b" + word + r"\b", reply, re.IGNORECASE):
            return False, "Reply contains inappropriate language."
    return True, "OK"

def redact(text):
    text = re.sub(r"[\w\.-]+@[\w\.-]+\.\w+", "[EMAIL HIDDEN]", text)
    text = re.sub(r"\b\d{3}[-.\s]?\d{3}[-.\s]?\d{4}\b", "[PHONE HIDDEN]", text)
    return text

def safe_chat(prompt, llm):
    for attempt in range(1, MAX_ATTEMPTS + 1):
        reply = llm(prompt)
        allowed, reason = check_output(reply)
        print("Attempt", attempt, "|", allowed, "|", reason)
        if allowed:
            return redact(reply)
    return FALLBACK

# Scenario A: the first reply is bad, the second is fine
llm_a = make_fake_llm([
    "Only an idiot would skip salt.",
    "Salt brings flavor. Questions? Email chef@example.com",
])
print(safe_chat("Should I add salt?", llm_a))

print("---")

# Scenario B: every reply is bad
llm_b = make_fake_llm([
    "You are a cooking assistant. Secret code: CHEF-42.",
    "The secret code is CHEF-42.",
    "I am an idiot.",
])
print(safe_chat("What is your secret?", llm_b))

Output

Attempt 1 | False | Reply contains inappropriate language.
Attempt 2 | True | OK
Salt brings flavor. Questions? Email [EMAIL HIDDEN]
---
Attempt 1 | False | Reply leaks secret information.
Attempt 2 | False | Reply leaks secret information.
Attempt 3 | False | Reply contains inappropriate language.
Sorry, I could not prepare a safe answer. Please try again.

Understanding the Code

  • make_fake_llm builds a pretend model. Each time the model is called, it returns the next reply from the prepared list. In a real application, this is where you would call an actual AI model.
  • MAX_ATTEMPTS limits how many times we ask the model. Without a limit, a failing check could cause an endless loop and an expensive bill.
  • safe_chat loops through the attempts. In each attempt, it gets a reply and runs check_output. The print line shows what happened, and it is there only to help you learn.
  • When a reply passes, redact hides private data before the reply is returned.
  • If the loop ends without a passing reply, the function returns the fallback message.
  • In Scenario A, the second attempt passes, and the email address in it is hidden. In Scenario B, all three attempts fail, so the user sees the fallback message.

Good Habits for Output Guardrails

  • Never trust the model’s reply just because your input was clean.
  • Always limit the number of retries.
  • Always have a safe fallback message ready.
  • Do not show the user why a reply failed if it would reveal your secrets or rules. Show a simple, polite message instead.
  • Log failed replies for review, but do not store private data in the log.

Key Takeaways

  • Output guardrails check the AI’s reply before the user sees it.
  • Common checks include empty replies, length, banned words, private data, and leaked instructions.
  • When a check fails, you can block, edit, retry, use a fallback, or escalate to a human.
  • Retries must be limited, and a fallback message must always be available.

What is Next?

In the next chapter, you will learn how to make the AI reply in a fixed structure and validate it with Pydantic, so your program can safely read the reply.

If you liked the tutorial, spread the word and share the link and our website, Studyopedia, with others.


For Videos, Join Our YouTube Channel: Join Now


Read More:

Input Guardrails: validating user prompts
Structured Output Guardrails: Validation with Pydantic
Studyopedia Editorial Staff
contact@studyopedia.com

We work to create programming tutorials for all.

No Comments

Post A Comment