04 Oct Types of AI Guardrails
Not all guardrails do the same job. Some check what the user sends, some check what the AI replies, and some protect against attacks. Knowing the types helps you decide which ones your application needs.
In this chapter, you will learn the main types of guardrails and see how several of them work together in one small program.
Types by Position
The first way to group guardrails is by where they run.
Input Guardrails
Input guardrails check the user’s message before it reaches the AI model. They can reject a message, clean it, or flag it for review.
- Message is too long or empty
- Message contains abusive language
- Message contains private data that should not be sent to the model
- Message looks like a prompt injection attempt
Output Guardrails
Output guardrails check the AI’s reply before the user sees it. They can block the reply, edit it, or ask the model to try again.
- Reply contains emails, phone numbers, or other private data
- Reply contains harmful or offensive content
- Reply is not in the required format
- Reply contains claims that cannot be supported
Types by Purpose
The second way to group guardrails is by what they protect.
Topical Guardrails
These keep the AI inside its intended subject. A cooking assistant should talk about cooking, and a banking assistant should talk about banking.
Safety Guardrails
These stop harmful, hateful, violent, or otherwise unsafe content from being accepted or produced.
Security Guardrails
These defend against attacks, such as prompt injection and jailbreak attempts, where someone tries to make the AI break its own rules.
Privacy Guardrails
These detect and hide personal information such as names, emails, phone numbers, and card numbers.
Format Guardrails
These check that the output has the structure your program expects, for example valid JSON with the right fields.
Accuracy Guardrails
These try to reduce hallucinations, for example by checking that the answer is supported by trusted documents.
Types by Method
The third way to group guardrails is by how they make their decision.
- Rule-based guardrails use fixed rules such as keyword lists, length limits, and regular expressions. They are fast, cheap, and predictable, but they can miss tricky cases.
- Model-based guardrails use another AI model or a classifier to judge the text. They understand meaning better, but they are slower, cost more, and can also make mistakes.
- Hybrid guardrails combine both. A common approach is to run quick rules first and use a model only when the rules are not enough.
Putting Several Types Together
Real applications use more than one guardrail. In the example below, we combine a length check and a topic check on the input, and an email-hiding check on the output.
import re
# A fake AI model that accidentally includes an email address
def fake_llm(prompt):
return "Try this: boil pasta for 10 minutes. Questions? Email chef@example.com"
# Input guardrail 1: length check
def length_check(prompt):
return len(prompt) <= 200
# Input guardrail 2: topical check
def topic_check(prompt):
keywords = ["recipe", "cook", "bake", "ingredient", "pasta"]
text = prompt.lower()
for word in keywords:
if word in text:
return True
return False
# Output guardrail: privacy check
def mask_emails(text):
return re.sub(r"[\w\.-]+@[\w\.-]+\.\w+", "[EMAIL HIDDEN]", text)
# The complete chatbot
def safe_chat(prompt):
# Input guardrails
if not length_check(prompt):
return "Your message is too long."
if not topic_check(prompt):
return "I can only help with cooking questions."
# Call the model
reply = fake_llm(prompt)
# Output guardrail
return mask_emails(reply)
print(safe_chat("How do I cook pasta?"))
print(safe_chat("Tell me a joke about politics"))
print(safe_chat("pasta " * 100))
Output
Try this: boil pasta for 10 minutes. Questions? Email [EMAIL HIDDEN] I can only help with cooking questions. Your message is too long.
Understanding the Code
- length_check is an input guardrail. It returns True only if the message has 200 characters or fewer.
- topic_check is a topical guardrail. It returns True only if the message contains one of the cooking keywords.
- mask_emails is a privacy guardrail on the output. It replaces any email address in the reply.
- safe_chat runs the checks in order. If an input check fails, it returns a message right away and the model is never called. If both pass, it calls the model and then cleans the reply.
- The first message passes every check. The second fails the topic check. The third repeats the word “pasta” 100 times, so it is longer than 200 characters and fails the length check.
Choosing the Right Guardrails
You do not need every type in every project. A good way to decide is to ask what could go wrong in your application.
- If the AI handles customer data, add privacy guardrails.
- If your program reads the AI’s reply, add format guardrails.
- If the chatbot is public, add safety and security guardrails.
- If wrong answers are costly, add accuracy guardrails.
- If the AI has a specific job, add topical guardrails.
Key Takeaways
- Guardrails can be grouped by position (input and output), by purpose (topical, safety, security, privacy, format, accuracy), and by method (rule-based, model-based, hybrid).
- Real applications combine several guardrails in a pipeline.
- Choose guardrails based on the risks of your own application.
What is Next?
In the next chapter, you will set up your Python environment so you can run all the examples in this tutorial on your own computer.
If you liked the tutorial, spread the word and share the link and our website, Studyopedia, with others.
For Videos, Join Our YouTube Channel:Â Join Now
Read More:
- Generative AI Tutorial
- AI Ethics
- Machine Learning Tutorial
- Deep Learning Tutorial
- Ollama Tutorial
- Retrieval Augmented Generation (RAG) Tutorial
- ChatGPT Tutorial
- Microsoft Copilot Tutorial
No Comments