For the complete documentation index, see llms.txt. This page is also available as Markdown.

Create Guardrail Policies

Overview

The Create Guardrail Policies page allows you to define enforcement rules that evaluate agent requests and responses against security, safety, and compliance criteria.

Each guardrail policy combines metadata, content filters, detection logic, and deployment scope into a single configuration that applies consistently across selected agents and servers.

Access Guardrail Policies

You can access guardrail policy configuration from the Akto console.

  • Navigate to the Agentic Security product.

  • Select Agentic Guardrails → Guardrail Policies.

The guardrail policies list displays existing policies and provides access to policy creation.

Create a Guardrail Policy

Access the Create Guardrail Form

  1. Locate the Create Guardrail button in the top-right corner of the Guardrail Policies page.

  2. Select Create Guardrail to open the guardrail configuration form.

Fill the Configuration Form

The configuration form is organised into multiple sections, each targeting a specific enforcement layer.

1. Provide Guardrail Policy Details (Mandatory)

This section defines identifying metadata and user-facing enforcement messages.

  • Enter a policy name and description.

  • Select the severity level to classify the importance of violations generated by the guardrail policy.

  • Define the blocked message shown when a request is denied.

  • Enable the option to apply the same blocked message to responses, if consistent messaging is required across requests and responses.

This section is required to create a guardrail policy.

2. Content & Policy Guardrails

Configure content moderation and policy-based control:

Prompt Injection Attack Filter

  • Enable the prompt attack filter to detect jailbreaks and manipulation attempts.

  • Adjust detection strength using the horizontal slider..

Context Poisoning Attacks

  • Enable the context poisoning filter to detect attempts to manipulate agent memory or context.

Add Denied Topics

Denied topics block specific concepts in user inputs or model responses.

  • Select Add Denied Topic to open the topic configuration form.

  • Provide a topic name, definition, and optional sample phrases.

  • Use + Add to include additional topics.

  • Select Save Topic after completing topic configuration.

  • Alternatively, select Add Akto Default Topics to instantly add Akto's predefined denied topics, grouped by category:

    • Safety-critical: Weapons & Firearms, Explosives & Bombs, Self-harm & Suicide, Illegal Drugs.

    • Professional advice: Medical / Health Advice, Financial / Investment Advice, Legal Advice.

You can add up to 30 denied topics per guardrail policy.

Map Topics to Compliance Frameworks

Akto auto-tags each custom denied topic against supported compliance frameworks, including SOC 2, HIPAA, GDPR, and ISO 27001. Every topic maps to the controls it enforces, so you can see at a glance which compliance requirements a guardrail policy helps satisfy.

  • Review the frameworks suggested under Compliance frameworks supported by this guardrail.

  • Select + Add to map the topic to additional frameworks.

  • Select the x on a framework tag to remove that mapping.

Harmful Categories Filter

  • Enable the harmful content filter using the category checkbox.

  • Configure enforcement strength for hate, insults, sexual content, violence, and misconduct by adjusting the category sliders.

  • Enable the option to apply category filtering to responses, if required.

Agent Intent Verification

Enable Intent verification to validate whether agent-generated requests align with expected intent.

  • Set Confidence Threshold (0–1). Higher values require more confidence to block content.

3. Language Safety and Abuse Guardrails

Gibberish Detection

Identify nonsensical or meaningless inputs that may disrupt agent processing or indicate misuse.

  • Enable gibberish detection and define a confidence threshold.

Sentiment Detection

You can just enable sentiment detection to evaluates inputs for negative, toxic, or inappropriate sentimentand configure a confidence threshold.

Profanity

Word filters enforce keyword- and phrase-based restrictions.

  • Enable profanity filtering to redact offensive language.

  • Add custom words or phrases:

    • Enter a word or phrase and select Add Word.

Akto support for up to 10,000 custom entries.

4. Sensitive Information Guardrails

Configure controls to detect, block, or anonymise sensitive data.

Personally Identifiable Information Types

  • Select predefined PII types to detect and block.

  • For each selected PII type, configure the guardrail behavior:

    • Block to deny the request or response containing the detected PII.

    • Mask to redact the detected PII before processing or returning the content.

Regex Pattern

Akto detects sensitive data based on defined patterns.

  • Add up to 10 custom regex patterns for structured data detection.

Secrets Detection

Akto detects API keys, passwords, and similar sensitive values.

  • Just enable secrets detection and set the confidence threshold.

Sensitive Data Anonymisation

Akto replaces detected sensitive data with placeholders while preserving original values securely for controlled restoration.

5. Advanced Code Detection Filters

You configure detection and blocking of programming code and code injection attempts.

Code Detection Filter

Akto detects code patterns across supported programming languages and blocks content based on the configured level.

  • Enable the code detection filter and set the Code Detection Level.

Ban Code Detection

Enable ban code detection to block all code regardless of programming language.

  • Set the Confidence Threshold (0–1) to control detection strictness.

Behaviour

  • Lower threshold values enforce stricter blocking.

  • Higher threshold values allow more permissive detection.

6. Custom Guardrails

LLM prompt based rule

This rule uses an LLM to classify content against a custom prompt.

  • Define the evaluation prompt.

  • Configure a confidence score threshold.

  • Content is blocked when the model confidence exceeds the configured threshold.

LLM based redaction

This rule uses an LLM to mask sensitive content described in plain language, instead of blocking the request or response outright.

  • Enable LLM based redaction.

  • Define the redaction instruction describing what to mask, for example "Redact customer full names and home addresses." Be specific: only content matching the instruction is masked, and anything not described is left untouched. To redact a second, unrelated category, create another policy.

  • Matching text is replaced in place and the request is allowed through, instead of being blocked.

External model based evaluation

External evaluation allows integration with third-party or internal scoring systems.

  • Provide the external evaluation endpoint URL.

  • Configure a confidence threshold that determines enforcement actions based on the external response.

7. Usage Based Guardrails

Token Limit Detection

Akto can evaluate input size and detects requests that exceed acceptable token limits.

  • Enable token limit detection and set the confidence threshold.

Behaviour

  • Higher threshold values allow larger inputs.

  • Lower threshold values enforce stricter limits and block oversized requests.

8. Anomaly Detection

Configure anomaly detection rules to identify unusual patterns in agent behaviour, tool usage, and system metrics.

Anomaly detection guardrails are tagged with the OWASP Agentic risk categories they help mitigate, for example ASI07 – Insecure Inter-Agent Communication, ASI08 – Cascading Failures, and ASI10 – Rogue Agents.

Statistical Anomalies

  • Enable anomaly detection to detect and alert on abnormal tool call counts and error counts per session.

  • Set the Tool call limit (per session): total tool calls allowed per session before triggering an anomaly.

  • Set the Error limit (per session): total errors (4xx/5xx) per session before triggering an anomaly.

Structural Anomalies

Detects irregularities in request structure or interaction flows, with configuration options rolling out.

Behavioral Anomalies

Identifies unexpected agent actions or tool usage patterns, with configuration options rolling out.

9. Enterprise License Compliance Guardrails

Block prompts and responses that would breach your LLM provider's acceptable-use policy, keeping your usage compliant so your enterprise license isn't put at risk.

Enable the categories you want enforced:

  • Child Sexual Abuse Material (CSAM): Blocks content that sexualizes, exploits, grooms, or endangers minors.

  • Malicious Code & Cyberattacks: Blocks requests to create or deploy malware, ransomware, exploits, phishing, hacking, DDoS, or other cyberattacks.

  • Weapons of Mass Destruction (CBRN): Blocks assistance with chemical, biological, radiological, or nuclear weapons, including synthesis, weaponization, delivery, or procurement.

  • Violent Extremism & Terrorism: Blocks content that facilitates, promotes, or incites terrorism, mass violence, genocide, or extremist attacks, including planning and recruitment.

  • Hate Speech & Discrimination: Blocks content that dehumanizes or incites hatred against individuals or groups based on race, ethnicity, religion, gender, sexual orientation, disability, or national origin.

  • Human Trafficking & Sexual Exploitation: Blocks content facilitating human trafficking, forced labor, sexual exploitation, or modern slavery, including recruitment scripts and coercion methods.

  • Non-consensual Surveillance & Tracking: Blocks assistance building covert tracking tools, stalkerware, or unauthorized monitoring systems targeting individuals without consent.

Select the info icon next to a category for its full detection scope.

10. Tool Guardrails

Tool Misuse

Evaluate tool invocation patterns and blocks suspicious activity.

  • Enable tool misuse detection to identify unauthorized or unsafe tool usage by agents.

Detect Malicious Tool

Block tools that exhibit unsafe or exploitative characteristics.

  • Enable malicious tool detection to identify tools with harmful behavior or intent.

Detect Tool Name and Description Mismatch

Detect cases where tool behavior does not align with declared metada

  • Enable mismatch detection to identify inconsistencies between tool name and description.

11. Access Restrictions

Configure controls to restrict which hosts, paths, and account types can interact with your agents.

Block host / path

Block outbound traffic by host or path pattern to prevent agents from reaching unauthorised external services or endpoints.

  • Enter a host or path pattern in the Host or path pattern field.

  • Select Add to include the pattern in the blocked list.

The Blocked patterns list displays all configured entries. You can review and remove patterns from this list at any time.

Akto sits between your agents and their outbound traffic. When a request is made, Akto checks the destination host or path against the blocked patterns and denies any match before it reaches the target.

Block personal accounts

Prevent users with personal or consumer email accounts from accessing the AI agent. Enterprise accounts using company email domains are allowed through.

  • Enable personal account blocking to restrict access to organisation-managed accounts only.

Personal account blocking is enforced through the browser extension, and currently supports the following browser LLMs: chatgpt.com, gemini.google.com, claude.ai, copilot.microsoft.com, and grok.com. Requires Akto browser extension v1.0.69 or later; it is currently supported only through the browser extension, with AI Endpoint Shield Agent support coming soon.

12. Exceptions

Configure phrases that this policy's own detectors should treat as safe and skip, without affecting how other guardrail policies evaluate the same traffic.

Ignore Phrases

  • Enter a phrase in the Ignore phrases field (e.g. your product name or sample test data). Phrases are matched as a whole word by default.

  • Enable Regex to match the entry as a regular expression instead of literal text.

  • Enable Case sensitive to match the phrase's casing exactly.

  • Select Add phrase to include it in the ignored list.

The Ignored phrases list displays a count of configured entries and lets you remove any phrase at any time.

Ignore phrases only affect this policy's own detectors — other guardrail policies evaluated on the same traffic still see the real text.

13. Scope

Configure which servers the guardrail should be applied to and specify whether it applies to requests, responses, or both. The coverage controls differ by product.

Akto Atlas: Agentic Assets

Choose which assets the guardrail is deployed across:

  • Apply to all (Recommended): Applies the guardrail to all assets. The panel shows the current counts, for example 46 Agents, 95 MCP Servers, and 22 LLMs, as links you can select to view the full list.

  • Select Agentic Assets: Target specific assets using a condition builder.

    • Under Where my, choose an asset type such as Agents, MCP Servers, or LLMs.

    • Use the are field to search and select the matching values.

    • Select Add condition to layer additional conditions, or Clear all to reset the configuration.

See how many MCP servers and agent servers a policy touches before you deploy it. Apply to All now shows the full scope up front.

Akto Atlas: Device Tags & Users

Choose which users the guardrail applies to:

  • Apply to all (Recommended): Applies the guardrail to all users in your organisation, for example 101 Users.

  • Select Device Tags & Users: Target specific users using the same condition builder.

    • Under Where my, choose an attribute such as Groups.

    • Use the are field to search and select the matching values.

    • Select Add condition to layer additional conditions, or Clear all to reset the configuration.

    • The number of matching users updates as you build the condition, for example Applies to 1 user with the current selection.

Akto Argus: Coverage

Akto Argus shows a single Coverage section instead of separate Agentic Assets and Device Tags & Users sections:

  • Apply to all (Recommended): Applies the guardrail to all discovered assets. The panel shows the current counts, for example 60 Agents & 31 MCP Servers, as links you can select to view the full list.

  • Advance Configuration: Build a condition to target specific assets.

    • Under Where my, choose an asset type such as AI Agents or MCP Servers.

    • Use the are field to search and select the matching values.

    • Select Add condition to layer additional conditions, or Clear all to reset the configuration.

Some agentic assets may show as disabled in Advance Configuration. Block rule behaviour requires the target server to run in inline (sync) mode.

Rule Behaviour

Choose how Akto responds when a guardrail condition is triggered:

  • Block: Stops the request or response when the condition is met.

  • Alert: Generates an alert for review without blocking content.

  • Human Approval: Blocks the triggering attempt and sends it to the Needs Approval tab in Guardrail Activity for a reviewer to approve. See Manage Human Approval Requests for the full workflow.

Application Settings

Specify whether the guardrail should be applied to responses and/or requests:

  • Apply guardrail to requests: Evaluates user inputs before processing.

  • Apply guardrail to responses: Evaluates model outputs before delivery.

You can enable either option independently or both together based on enforcement requirements.

Browser LLM Enforcement

For Browser LLMs (chatgpt.com, gemini.google.com, claude.ai, copilot.microsoft.com, and grok.com), Akto enforces guardrails on the request path in real time, stopping risky prompts before they ever reach the model. Response-side enforcement for these domains is coming soon.

OWASP Agentic Risk Tags

Guardrail policies include OWASP-aligned risk tags. You can click each tag in the UI to understand the associated risk category and its security impact on agent behaviour.

Save the Guardrail Policy

After completing the required and optional configurations:

  • Click on Create Policy to save the policy and applies enforcement to the selected scope.

Test guardrail behaviour in the playground

The playground allows your team to validate guardrail behaviour before updating the policy.

  • Enter a prompt in the Test your guardrail policy field to simulate a request against the configured guardrail policy.

  • The playground evaluates the prompt using the selected guardrail configuration and displays the enforcement result.

  • You can also use the Quick Test Prompts provided in the playground to test common scenarios such as sensitive data exposure, prompt injection attempts, or abusive language.

Playground probing helps your security team verify that guardrail conditions correctly detect violations and return the expected blocked response message before the policy is finalized.

Preview Change Impact Before Saving

When you edit an existing guardrail policy, the right-hand panel shows Change impact analysis instead of the plain Impact analysis view shown while creating a new policy. It replays recent activity through both the currently saved policy and your unsaved changes, so you can see how an edit would affect detections before you save it.

  • Switch between the Violations and Traffic tabs before running:

    • Violations replays the policy's last few recorded violations.

    • Traffic replays the latest agent traces.

  • Select Run to open the Change impact analysis window and replay that activity against both versions of the policy.

  • The Saved policy and Your changes counts show the total detections under each version, with a badge (e.g. -1) highlighting the net change.

  • The results table lists each replayed Prompt alongside two columns, Saved and Draft, each marked Detected or Missed so you can see exactly which prompts change outcome under your edits.

  • Use Page Size and the pagination controls to page through results, or select Export CSV to download the full comparison.

  • Select Close to return to the policy editor.

Change impact analysis is only available when editing an existing guardrail policy. While creating a new policy, this panel appears as Impact analysis and previews how the new policy would perform against recent agent traffic once you select Run.

What’s Next

You can modify, disable, or delete existing guardrail policies after creation.

To continue, learn how to manage guardrail policies from the Manage Guardrail Policies.

Last updated