04 Oct Your First AI Guardrail: keyword and regex rules
The simplest guardrails are rule-based. They look for specific words or patterns in the text and decide what to do. They are fast, free, easy to understand, and a great starting point for beginners.
In this chapter, you will learn how to check for blocked words correctly, how to detect patterns such as phone numbers using regular expressions, and how to build a reusable input check that explains why a message was rejected.
Blocklist and Allowlist
There are two common ways to write keyword rules.
- Blocklist: A list of words or patterns that are not allowed. Anything that matches is rejected. Use this when most messages are fine and only a few things are dangerous.
- Allowlist: A list of words or patterns that are required or permitted. Anything else is rejected. Use this when your application has a narrow purpose, like the cooking assistant from the earlier chapters.
The Problem with Simple Keyword Matching
In Chapter 1, our guardrail used the Python expression “hack” in text. This finds the letters h-a-c-k anywhere, even inside other words. So a harmless message about a “hackathon” was blocked. We need to match whole words only.
Matching Whole Words with Regular Expressions
A regular expression, or regex, is a pattern that describes text. The special symbol \b means “word boundary”, which is the edge between a letter and a space or punctuation mark. Putting \b on both sides of a word makes sure we match the complete word.
Example 1: Whole-Word Blocklist
import re
BLOCKED_WORDS = ["hack", "password", "exploit"]
def contains_blocked_word(text):
for word in BLOCKED_WORDS:
pattern = r"\b" + re.escape(word) + r"\b"
if re.search(pattern, text, re.IGNORECASE):
return True
return False
print(contains_blocked_word("Tell me about a hackathon event"))
print(contains_blocked_word("How to hack a website"))
print(contains_blocked_word("Reset my PASSWORD"))
Output
False True True
Understanding the Code
- BLOCKED_WORDS is our blocklist.
- re.escape(word) makes sure that special characters in a word are treated as normal text. It is a good habit even when your words are simple.
- r”\b” + … + r”\b” builds a pattern that matches the word only when it stands alone.
- re.search looks for the pattern anywhere in the text and returns a match, or None if nothing is found.
- re.IGNORECASE makes the search ignore capital letters, so “PASSWORD” and “password” are treated the same.
- The first message contains “hackathon”, not the whole word “hack”, so it is allowed. The other two messages contain blocked words.
Detecting Patterns: Phone Numbers
Some things cannot be listed as words, such as phone numbers, because there are endless possible numbers. For these, we describe the shape of the data using a regex. Here are the pieces we will use:
- \d matches one digit.
- \d{3} matches exactly three digits.
- [-.\s]? matches one optional separator: a dash, a dot, or a space. The question mark means “zero or one time”.
Example 2: Finding Phone Numbers
import re
def find_phone_numbers(text):
pattern = r"\b\d{3}[-.\s]?\d{3}[-.\s]?\d{4}\b"
return re.findall(pattern, text)
print(find_phone_numbers("Call me at 555-123-4567 or 555.987.6543"))
print(find_phone_numbers("My order number is 42"))
Output
['555-123-4567', '555.987.6543'] []
Understanding the Code
- The pattern means: three digits, an optional separator, three more digits, an optional separator, and four final digits.
- re.findall returns a list of every piece of text that matches the pattern.
- The first message has two phone numbers in different styles, and both are found.
- The second message has only the number 42, which is too short to match, so the list is empty.
This pattern fits phone numbers written in a ten-digit style. Phone number formats differ across countries, so for your own website you may need to adjust the pattern.
Building a Reusable Input Check
When a guardrail rejects a message, it is useful to know the reason. This helps you show a clear reply to the user and also helps you debug problems. Let us combine our rules into one function that returns two things: whether the message is allowed, and a reason.
Example 3: Input Check with Reasons
import re
BLOCKED_WORDS = ["hack", "password", "exploit"]
def contains_blocked_word(text):
for word in BLOCKED_WORDS:
pattern = r"\b" + re.escape(word) + r"\b"
if re.search(pattern, text, re.IGNORECASE):
return True
return False
def find_phone_numbers(text):
pattern = r"\b\d{3}[-.\s]?\d{3}[-.\s]?\d{4}\b"
return re.findall(pattern, text)
def check_input(text):
if len(text.strip()) == 0:
return False, "Message is empty."
if contains_blocked_word(text):
return False, "Message contains a blocked word."
if find_phone_numbers(text):
return False, "Message contains a phone number."
return True, "OK"
tests = [
"",
"How to hack a website",
"Call 555-123-4567",
"What is a guardrail?",
]
for t in tests:
allowed, reason = check_input(t)
print(repr(t), "|", allowed, "|", reason)
Output
'' | False | Message is empty. 'How to hack a website' | False | Message contains a blocked word. 'Call 555-123-4567' | False | Message contains a phone number. 'What is a guardrail?' | True | OK
Understanding the Code
- check_input runs the rules one by one and stops at the first one that fails.
- text.strip() removes spaces from both ends, so a message with only spaces is treated as empty.
- The function returns a pair of values. Python lets us unpack them into two variables with allowed, reason = check_input(t).
- repr(t) prints the text with quotes, which makes the empty message easy to see in the output.
- Each of the first three messages fails a different rule. The last message passes every rule.
Limits of Rule-Based Guardrails
Rule-based guardrails are useful, but you should know where they fall short.
- Tricks: A user can write “h a c k” or “h4ck” to slip past a word list.
- Meaning: Rules look at words, not meaning. A message can be harmful without using any blocked word.
- False alarms: A rule can block innocent messages, such as a student asking how security teams protect a password.
- Maintenance: Lists need to be updated as new problems appear.
For these reasons, rule-based checks are best used as the first layer, combined with smarter methods that you will learn in later chapters.
Key Takeaways
- A blocklist rejects listed items. An allowlist accepts only listed items.
- Use \b around a word in a regex to match whole words only.
- Use regex patterns to detect data with a fixed shape, such as phone numbers.
- Return both a decision and a reason from your guardrail functions.
- Rule-based guardrails are fast and simple, but they can be tricked and can miss meaning.
What is Next?
In the next chapter, you will go deeper into input guardrails and learn how to validate and clean user prompts before they reach the AI model.
If you liked the tutorial, spread the word and share the link and our website, Studyopedia, with others.
For Videos, Join Our YouTube Channel:Â Join Now
Read More:
- Generative AI Tutorial
- AI Ethics
- Machine Learning Tutorial
- Deep Learning Tutorial
- Ollama Tutorial
- Retrieval Augmented Generation (RAG) Tutorial
- ChatGPT Tutorial
- Microsoft Copilot Tutorial
No Comments