Account & AI SafetyVerified Module

AI Governance & Safeguards (MCP)

Real-time AI guardrails: automatic response evaluation, content filtering, and human escalation.

AI Governance & Safeguards (MCP) is our built-in safety gatekeeper for Orbion Agents. Before an AI Agent replies to a customer on WhatsApp, Instagram, or Email, the MCP engine automatically evaluates the message against strict safety rules. It decides whether to ALLOW safe responses, BLOCK harmful or abusive content, or ESCALATE sensitive customer issues (like refunds, complaints, or high-value leads) directly to a human team member in the Omni-Channel Inbox.

Why use it?

Core Capabilities & Value

Purpose-built operational advantages engineered for high-throughput teams.

1

Automated Response Safety (Allow / Block / Escalate)

Zero-Risk Automation

Every AI reply is evaluated before reaching the customer. Safe inquiries receive instant autonomous answers, while inappropriate or harmful messages are blocked immediately.

2

Human Escalation for High-Value & Sensitive Inquiries

Human-in-the-Loop

When a customer requests a refund, files a major complaint, or asks about custom enterprise pricing, the AI pauses and immediately escalates the conversation to your team in the Omni-Channel Inbox.

3

Blocked Keywords & Spam Filtering

Brand Safety Protection

Filter out competitor names, prohibited words, or spam. If a user attempts prompt injection tricks or sends abusive messages, the AI safely rejects the input and stays on topic.

4

Accurate Knowledge Base Answers (Confidence Scores)

No Hallucinations

The AI only answers when it is confident in the information retrieved from your business Knowledge Base (RAG). If confidence is low, it connects the user to a human agent instead of guessing.

FAQ & Troubleshooting

Feature Troubleshooting & Common Questions

Clear answers and solutions to common questions when connecting and managing AI Governance & Safeguards (MCP).

What does the AI Safeguard do when it cannot find an answer in my Knowledge Base?

Answer: If the AI's confidence score is low because the answer is not in your uploaded documents, the safeguard prevents the bot from guessing or making up false details. Instead, it politely informs the customer and escalates the chat to a human team member.

How does the system handle refund requests or angry customer messages?

Answer: The MCP safeguard detects sensitive keywords (like 'refund', 'cancel', or complaint phrases) and flags the chat as an 'Escalation'. An alert banner appears on the conversation in your Omni-Channel Inbox so a human agent can step in immediately.

Can I prevent the AI Agent from discussing certain topics or competitors?

Answer: Yes. You can configure blocked keywords in your workspace settings. If a user asks about a blocked topic or competitor, the safeguard stops the AI from promoting or discussing those terms.

How do human agents take over an escalated conversation?

Answer: When a chat is escalated, your team sees an 'Escalated by AI Safeguard' badge in the Omni-Channel Inbox. Any human agent can simply click into the chat and start typing. The AI pauses automatically until handed back.

Does the AI safeguard delay responses to customers on WhatsApp or Instagram?

Answer: No. The safeguard evaluates messages in real time in memory within a fraction of a second, so customers experience seamless, instant replies.