AI Governance & Safeguards (MCP)
Real-time AI guardrails: automatic response evaluation, content filtering, and human escalation.
AI Governance & Safeguards (MCP) is our built-in safety gatekeeper for Orbion Agents. Before an AI Agent replies to a customer on WhatsApp, Instagram, or Email, the MCP engine automatically evaluates the message against strict safety rules. It decides whether to ALLOW safe responses, BLOCK harmful or abusive content, or ESCALATE sensitive customer issues (like refunds, complaints, or high-value leads) directly to a human team member in the Omni-Channel Inbox.
Core Capabilities & Value
Purpose-built operational advantages engineered for high-throughput teams.
Automated Response Safety (Allow / Block / Escalate)
Zero-Risk AutomationEvery AI reply is evaluated before reaching the customer. Safe inquiries receive instant autonomous answers, while inappropriate or harmful messages are blocked immediately.
Human Escalation for High-Value & Sensitive Inquiries
Human-in-the-LoopWhen a customer requests a refund, files a major complaint, or asks about custom enterprise pricing, the AI pauses and immediately escalates the conversation to your team in the Omni-Channel Inbox.
Blocked Keywords & Spam Filtering
Brand Safety ProtectionFilter out competitor names, prohibited words, or spam. If a user attempts prompt injection tricks or sends abusive messages, the AI safely rejects the input and stays on topic.
Accurate Knowledge Base Answers (Confidence Scores)
No HallucinationsThe AI only answers when it is confident in the information retrieved from your business Knowledge Base (RAG). If confidence is low, it connects the user to a human agent instead of guessing.
Feature Troubleshooting & Common Questions
Clear answers and solutions to common questions when connecting and managing AI Governance & Safeguards (MCP).
Answer: If the AI's confidence score is low because the answer is not in your uploaded documents, the safeguard prevents the bot from guessing or making up false details. Instead, it politely informs the customer and escalates the chat to a human team member.
Answer: The MCP safeguard detects sensitive keywords (like 'refund', 'cancel', or complaint phrases) and flags the chat as an 'Escalation'. An alert banner appears on the conversation in your Omni-Channel Inbox so a human agent can step in immediately.
Answer: Yes. You can configure blocked keywords in your workspace settings. If a user asks about a blocked topic or competitor, the safeguard stops the AI from promoting or discussing those terms.
Answer: When a chat is escalated, your team sees an 'Escalated by AI Safeguard' badge in the Omni-Channel Inbox. Any human agent can simply click into the chat and start typing. The AI pauses automatically until handed back.
Answer: No. The safeguard evaluates messages in real time in memory within a fraction of a second, so customers experience seamless, instant replies.