AI Guardrails Should Be Rules You Own

Credentials look the same everywhere. Personal data does not. Why guardrails work best as profiles you define, with a baseline teams cannot weaken.

Cloptima TeamOctober 5, 2026 8 min read
In this post
  1. 01Why teams switch guardrails off
  2. 02Two kinds of sensitive data
  3. 03Three actions, on two sides
  4. 04Rules you can read
  5. 05One baseline, many teams
  6. 06Roll out in four weeks
  7. 07Where meaning-aware checks fit
  8. 08What shows up in the audit trail

Why teams switch guardrails off

The most common guardrail failure is not a missed detection.

It is a team turning the guardrail off because it blocked something it should not have.

Picture this

A developer pastes a sample customer record into a prompt while debugging. A vendor-tuned filter flags it, the test run fails, and a team lead is asked why the assistant is broken. Two sprints later the filter is disabled for that app, temporarily. Nobody turns it back on.

Filters you cannot see into, tune, or explain train people to route around them. Guardrails work when the people they affect understand the rule, can change it, and can see what it did.

A guardrail that gets switched off protects nobody.

Two kinds of sensitive data

Not everything sensitive is the same kind of problem.

Some of it looks identical everywhere, and some of it only exists in your business.

Universal: credentials

  • API keys, tokens, private keys, password assignments
  • The same format in every company
  • Cloptima detects them out of the box

Local: personal and business data

  • National ID formats, phone numbers, internal codenames, customer identifiers
  • Formats and meaning depend on your region and business
  • You define the rules, starting from templates
  • Cloud and developer credentials: AWS access keys, GitHub, GitLab, Slack, Stripe, npm, Hugging Face, and SendGrid tokens
  • AI provider keys: Anthropic, OpenAI, and Google
  • Private keys, bearer tokens, and password or secret assignments
  • Cloptima's own tokens and virtual keys, so a leaked key in a prompt is caught too

Treating the two differently is what keeps false positives low. You are not asked to trust one global detector with your national ID formats.

Three actions, on two sides

Every guardrail profile has an input side, which covers prompts, and an output side, which covers responses.

Each side has one action, and the action decides what a match does.

ActionWhat happensGood for
ObserveRecords the finding. Nothing changes.The first week of any rollout
RedactReplaces the match with a marker before the model, or the user, sees itCredentials and personal data you want to keep flowing
BlockStops the request or response with a clear messageContent that must never leave, such as a customer ID bound for a third-party model

The sides are independent. A common setup redacts credentials in prompts so the work continues, and blocks them in responses so nothing leaks back out.

What a redaction does
Before:  Debug this config: api_key = abc123def456ghi789
After:   Debug this config: [REDACTED_SECRET]

Rules you can read

Custom rules are deliberately plain.

A rule is either a list of terms or one regular expression, with an id and an optional action.

main.tf
resource "cloptima_llm_guardrail_profile" "customer_data" {
  name = "Customer data protection"
  definition = jsonencode({
    input = {
      action    = "redact"
      detectors = { secret = {} }
      custom_rules = [
        { id = "us_ssn", match = { regex = "\\b\\d{3}-\\d{2}-\\d{4}\\b" } },
        { id = "codename", match = { terms = ["Project Falcon"], whole_word = true }, action = "block" },
      ]
    }
    output = {
      action    = "redact"
      detectors = { secret = {} }
    }
  })
}

Rules run on a linear-time matcher. In plain words, that is a guarantee: no pattern, however it is written, can make a request slower or stall other traffic. To keep that promise, rules have sensible bounds: up to 10 terms per rule, up to 64 characters per term or pattern, and no look-ahead or back-references.

One baseline, many teams

Security wants one floor. Teams want to add their own rules. A guardrail baseline gives you both.

Baseline (whole organization)Team profile (research)What happens
Redact credentialsBlock the codename Project FalconCredentials are redacted everywhere. The codename is blocked for research.
Redact credentialsObserve credentialsCredentials are still redacted. The baseline's stricter action wins.

Each profile is evaluated as written and the strictest outcome wins. A team profile can add protection, and it cannot weaken the baseline.

Roll out in four weeks

A guardrail you trust is one you have watched work.

This schedule moves from evidence to enforcement without surprising anyone.

  1. 1

    Week 1: observe everything

    Create the baseline from the Monitor only template. Review what it finds in the audit log.

  2. 2

    Week 2: redact credentials

    Switch the credential rules to Redact. Work continues, and secrets stop reaching models.

  3. 3

    Week 3: add personal-data rules in Observe

    Start from the US or European template. Tune rules that match too much.

  4. 4

    Week 4: block what must never leave

    Move the rules that matter most to Block, with a clear message for developers.

Where meaning-aware checks fit

Local rules are fast and predictable, but they cannot understand meaning.

For prompt injection, jailbreak attempts, harmful content, and sensitive information in free text, connect a safety service you already trust.

ProviderGood forYou bring
Azure AI Content SafetyPrompt Shields and harm categories: hate, self-harm, sexual, and violenceA Content Safety resource
AWS Bedrock GuardrailsYour own guardrail: content filters, denied topics, word filters, sensitive informationA guardrail id, version, and region
Google Model ArmorPrompt injection, responsible-AI filters, and sensitive data in a template you controlA template and a service account
Your own webhookAny custom checkAn HTTPS endpoint that answers allowed or not

Local rules run first, so a request they block never pays for a provider scan. You can cap what a provider scan may cost per request, skip it when a request is already slow, and choose what happens if it cannot run. Cloptima never adds a check you did not configure.

What shows up in the audit trail

Compliance conversations go better with evidence.

Guardrails leave a record you can point to.

  • Profile changes are versioned, with who changed what and when
  • Credential and custom-rule findings that pass through in Observe appear in the audit log
  • A blocked request names the rule that matched, never the content that matched it
  • The response a developer sees names the rule, so the fix is obvious
Put it into practiceCreate your first guardrail profileStart from a template, pick an action for prompts and responses, and set an organization baseline.

Keep reading

Ready to Try Cloptima?

Bring LLM FinOps, governed model access, and cloud cost optimization into one operating model.

No credit card required
5-minute setup
Free trial