04 Oct Content Moderation AI Guardrails: Blocking Toxic and Harmful Content
If your chatbot is open to the public, some users will send rude, hateful, or threatening messages. Sometimes the AI model itself may produce harmful text. Content moderation guardrails detect such content and decide how to respond.
In this chapter, you will learn the common categories of harmful content, build a simple rule-based moderation check, choose a different action for each category, and understand how score-based moderation with a trained model works.
Common Categories of Harmful Content
- Insults and harassment: Rude or abusive language aimed at a person or group.
- Hate speech: Content that attacks people because of who they are.
- Threats and violence: Messages that threaten or encourage harm to others.
- Self-harm: Messages from someone who may be thinking about hurting themselves.
- Sexual content: Explicit material that does not belong in your application.
- Illegal or dangerous requests: Requests for help with serious wrongdoing.
Which categories you need depends on your application. A children’s learning website needs strict rules, while a tool for security researchers may need different ones.
Not Every Category Needs the Same Response
It is tempting to block everything. But the best action depends on the category:
- A mild insult can get a polite warning.
- A threat should be blocked, and you may want to log it for review.
- A message about self-harm should not just be blocked. The person may be in a difficult moment, so a caring reply that encourages them to seek help is much better than a cold refusal.
Example 1: A Simple Rule-Based Moderation Check
The function below checks a message against phrase lists, one list for each category. It returns the categories it found and the phrases that matched. To keep the example simple, the lists are very short. Real lists are much longer.
import re
CATEGORIES = {
"INSULT": ["idiot", "stupid", "moron", "loser"],
"THREAT": ["i will hurt you", "i will find you", "you will regret this"],
"SELF_HARM": ["hurt myself", "end my life"],
}
def moderate(text):
lowered = text.lower()
found = {}
for category, phrases in CATEGORIES.items():
for phrase in phrases:
if re.search(r"\b" + re.escape(phrase) + r"\b", lowered):
found.setdefault(category, []).append(phrase)
return found
tests = [
"You are an idiot and a loser",
"I will hurt you",
"I feel like I might hurt myself",
"What is the capital of France?",
"This movie was stupid",
]
for text in tests:
print(moderate(text))
Output
{'INSULT': ['idiot', 'loser']}
{'THREAT': ['i will hurt you']}
{'SELF_HARM': ['hurt myself']}
{}
{'INSULT': ['stupid']}
Understanding the Code
- CATEGORIES maps each category name to a list of phrases.
- moderate lowercases the text and checks each phrase as a whole word or whole phrase, using the \b technique from Chapter 5.
- found.setdefault(category, []).append(phrase) creates an empty list for a category the first time it is needed, then adds the matched phrase to it.
- The function returns a dictionary. An empty dictionary means that nothing was found.
- Look at the last test. The message “This movie was stupid” is about a movie, not an attack on a person, yet it was flagged. This is a false positive, and it shows the main weakness of word lists: they do not understand context.
Example 2: Choosing an Action for Each Category
Now let us decide what to do. We assign an action to each category and set a priority order, so that when a message matches several categories, the most serious one decides the response.
import re
CATEGORIES = {
"INSULT": ["idiot", "stupid", "moron", "loser"],
"THREAT": ["i will hurt you", "i will find you", "you will regret this"],
"SELF_HARM": ["hurt myself", "end my life"],
}
PRIORITY = ["SELF_HARM", "THREAT", "INSULT"]
ACTIONS = {
"INSULT": "warn",
"THREAT": "block",
"SELF_HARM": "support",
}
RESPONSES = {
"warn": "Please keep the conversation respectful.",
"block": "Sorry, I can't continue with threatening messages.",
"support": "I'm really sorry you're feeling this way. You deserve support. Please consider contacting a local emergency number or a trusted person, and look for a helpline in your country.",
}
def moderate(text):
lowered = text.lower()
found = {}
for category, phrases in CATEGORIES.items():
for phrase in phrases:
if re.search(r"\b" + re.escape(phrase) + r"\b", lowered):
found.setdefault(category, []).append(phrase)
return found
def handle_message(text):
found = moderate(text)
for category in PRIORITY:
if category in found:
return RESPONSES[ACTIONS[category]]
return "OK: message passed to the AI model."
tests = [
"What is the capital of France?",
"You are an idiot",
"You idiot, I will find you",
"I feel like I might hurt myself",
]
for text in tests:
print(text, "|", handle_message(text))
Output
What is the capital of France? | OK: message passed to the AI model. You are an idiot | Please keep the conversation respectful. You idiot, I will find you | Sorry, I can't continue with threatening messages. I feel like I might hurt myself | I'm really sorry you're feeling this way. You deserve support. Please consider contacting a local emergency number or a trusted person, and look for a helpline in your country.
Understanding the Code
- PRIORITY lists the categories from most serious to least serious. The loop in handle_message checks them in this order and responds to the first one it finds.
- ACTIONS connects each category to an action name, and RESPONSES connects each action name to the text shown to the user.
- The third message contains both an insult and a threat. Because THREAT comes before INSULT in the priority list, the user gets the threat response.
- The fourth message gets a caring reply instead of a block. In a real application, you should write this message carefully, include a helpline suitable for your users’ country, and consider letting a human review such conversations.
Score-Based Moderation with a Trained Model
Word lists miss many harmful messages and flag many harmless ones. A better method is to use a trained moderation model, also called a classifier. You send it a piece of text, and it returns a score between 0 and 1 for each category. A higher score means the model is more confident that the text belongs to that category.
Moderation models are available in several forms, such as moderation endpoints offered by AI providers and open-source toxicity classifiers. Because they vary, we will simulate one here with prepared scores. This lets you learn the most important idea, which is how to turn scores into decisions with thresholds.
Example 3: Using Scores and Thresholds
# Pretend scores, as if they came from a trained moderation model
FAKE_SCORES = {
"Have a nice day": {"toxicity": 0.01, "threat": 0.00},
"This movie was stupid": {"toxicity": 0.35, "threat": 0.01},
"You never listen to anything": {"toxicity": 0.55, "threat": 0.02},
"You are a worthless idiot": {"toxicity": 0.92, "threat": 0.05},
}
BLOCK_AT = 0.80
REVIEW_AT = 0.50
def decide(scores):
highest = max(scores.values())
if highest >= BLOCK_AT:
return "block"
if highest >= REVIEW_AT:
return "send to human review"
return "allow"
for text, scores in FAKE_SCORES.items():
print(text, "|", decide(scores))
Output
Have a nice day | allow This movie was stupid | allow You never listen to anything | send to human review You are a worthless idiot | block
Understanding the Code
- FAKE_SCORES stands in for the answers of a real moderation model. Each message has a score for toxicity and a score for threats.
- BLOCK_AT and REVIEW_AT are thresholds. A score at or above 0.80 is blocked, and a score from 0.50 up to 0.80 is sent to a person for review.
- decide takes the highest score among all categories and compares it with the thresholds.
- The movie review, which was flagged by the word list in Example 1, now gets a low score and is allowed. This is the benefit of a model that looks at the whole sentence.
- The middle band is useful, because uncertain cases can be checked by a human instead of being wrongly blocked or wrongly allowed.
Choosing Thresholds
Thresholds are a balance between two kinds of mistakes.
- A low threshold blocks more harmful content, but it also blocks more harmless messages (false positives).
- A high threshold allows more harmless messages, but it lets more harmful content through (false negatives).
Test your thresholds on real examples from your own application and adjust them over time. A children’s website will choose lower thresholds than a general discussion forum.
Moderate Both Directions
Use moderation on the user’s message before it reaches the model, and also on the model’s reply before it reaches the user. The same moderate or decide function can be used in both places, just like the input and output guardrails from earlier chapters.
Good Habits for Moderation
- Start with rules for obvious cases, and add a trained model for better accuracy.
- Use different actions for different categories.
- Handle self-harm messages with care, not with a cold refusal.
- Allow a human review path for uncertain cases.
- Review blocked and allowed examples regularly to find mistakes.
- Tell users clearly what is not allowed in your application.
Key Takeaways
- Moderation guardrails detect harmful categories such as insults, hate, threats, and self-harm.
- Word lists are easy to build but do not understand context, so they cause false positives and false negatives.
- Trained moderation models return scores, and thresholds turn scores into allow, review, or block decisions.
- Choose the response by category. Warn for mild cases, block threats, and respond with care to self-harm.
- Apply moderation to both the input and the output.
What is Next?
In the next chapter, you will learn about prompt injection and jailbreak attacks, and how to defend your AI application against them.
If you liked the tutorial, spread the word and share the link and our website, Studyopedia, with others.
For Videos, Join Our YouTube Channel:Â Join Now
Read More:
- Generative AI Tutorial
- AI Ethics
- Machine Learning Tutorial
- Deep Learning Tutorial
- Ollama Tutorial
- Retrieval Augmented Generation (RAG) Tutorial
- ChatGPT Tutorial
- Microsoft Copilot Tutorial
No Comments