AI Guardrails Should Be Rules You Own
Credentials look the same everywhere. Personal data does not. Why guardrails work best as profiles you define, with a baseline teams cannot weaken.
In this post
Why teams switch guardrails off
The most common guardrail failure is not a missed detection.
It is a team turning the guardrail off because it blocked something it should not have.
Picture this
A developer pastes a sample customer record into a prompt while debugging. A vendor-tuned filter flags it, the test run fails, and a team lead is asked why the assistant is broken. Two sprints later the filter is disabled for that app, temporarily. Nobody turns it back on.
Filters you cannot see into, tune, or explain train people to route around them. Guardrails work when the people they affect understand the rule, can change it, and can see what it did.
A guardrail that gets switched off protects nobody.
Two kinds of sensitive data
Not everything sensitive is the same kind of problem.
Some of it looks identical everywhere, and some of it only exists in your business.
Universal: credentials
- API keys, tokens, private keys, password assignments
- The same format in every company
- Cloptima detects them out of the box
Local: personal and business data
- National ID formats, phone numbers, internal codenames, customer identifiers
- Formats and meaning depend on your region and business
- You define the rules, starting from templates
- Cloud and developer credentials: AWS access keys, GitHub, GitLab, Slack, Stripe, npm, Hugging Face, and SendGrid tokens
- AI provider keys: Anthropic, OpenAI, and Google
- Private keys, bearer tokens, and password or secret assignments
- Cloptima's own tokens and virtual keys, so a leaked key in a prompt is caught too
Treating the two differently is what keeps false positives low. You are not asked to trust one global detector with your national ID formats.
Three actions, on two sides
Every guardrail profile has an input side, which covers prompts, and an output side, which covers responses.
Each side has one action, and the action decides what a match does.
| Action | What happens | Good for |
|---|---|---|
| Observe | Records the finding. Nothing changes. | The first week of any rollout |
| Redact | Replaces the match with a marker before the model, or the user, sees it | Credentials and personal data you want to keep flowing |
| Block | Stops the request or response with a clear message | Content that must never leave, such as a customer ID bound for a third-party model |
The sides are independent. A common setup redacts credentials in prompts so the work continues, and blocks them in responses so nothing leaks back out.
Before: Debug this config: api_key = abc123def456ghi789
After: Debug this config: [REDACTED_SECRET]Rules you can read
Custom rules are deliberately plain.
A rule is either a list of terms or one regular expression, with an id and an optional action.
resource "cloptima_llm_guardrail_profile" "customer_data" {
name = "Customer data protection"
definition = jsonencode({
input = {
action = "redact"
detectors = { secret = {} }
custom_rules = [
{ id = "us_ssn", match = { regex = "\\b\\d{3}-\\d{2}-\\d{4}\\b" } },
{ id = "codename", match = { terms = ["Project Falcon"], whole_word = true }, action = "block" },
]
}
output = {
action = "redact"
detectors = { secret = {} }
}
})
}Rules run on a linear-time matcher. In plain words, that is a guarantee: no pattern, however it is written, can make a request slower or stall other traffic. To keep that promise, rules have sensible bounds: up to 10 terms per rule, up to 64 characters per term or pattern, and no look-ahead or back-references.
One baseline, many teams
Security wants one floor. Teams want to add their own rules. A guardrail baseline gives you both.
| Baseline (whole organization) | Team profile (research) | What happens |
|---|---|---|
| Redact credentials | Block the codename Project Falcon | Credentials are redacted everywhere. The codename is blocked for research. |
| Redact credentials | Observe credentials | Credentials are still redacted. The baseline's stricter action wins. |
Each profile is evaluated as written and the strictest outcome wins. A team profile can add protection, and it cannot weaken the baseline.
Roll out in four weeks
A guardrail you trust is one you have watched work.
This schedule moves from evidence to enforcement without surprising anyone.
- 1
Week 1: observe everything
Create the baseline from the Monitor only template. Review what it finds in the audit log.
- 2
Week 2: redact credentials
Switch the credential rules to Redact. Work continues, and secrets stop reaching models.
- 3
Week 3: add personal-data rules in Observe
Start from the US or European template. Tune rules that match too much.
- 4
Week 4: block what must never leave
Move the rules that matter most to Block, with a clear message for developers.
Where meaning-aware checks fit
Local rules are fast and predictable, but they cannot understand meaning.
For prompt injection, jailbreak attempts, harmful content, and sensitive information in free text, connect a safety service you already trust.
| Provider | Good for | You bring |
|---|---|---|
| Azure AI Content Safety | Prompt Shields and harm categories: hate, self-harm, sexual, and violence | A Content Safety resource |
| AWS Bedrock Guardrails | Your own guardrail: content filters, denied topics, word filters, sensitive information | A guardrail id, version, and region |
| Google Model Armor | Prompt injection, responsible-AI filters, and sensitive data in a template you control | A template and a service account |
| Your own webhook | Any custom check | An HTTPS endpoint that answers allowed or not |
Local rules run first, so a request they block never pays for a provider scan. You can cap what a provider scan may cost per request, skip it when a request is already slow, and choose what happens if it cannot run. Cloptima never adds a check you did not configure.
What shows up in the audit trail
Compliance conversations go better with evidence.
Guardrails leave a record you can point to.
- Profile changes are versioned, with who changed what and when
- Credential and custom-rule findings that pass through in Observe appear in the audit log
- A blocked request names the rule that matched, never the content that matched it
- The response a developer sees names the rule, so the fix is obvious