All guides

Investigate an Agent Incident from Alert to Cause

Follow a blocked, expensive, or misbehaving agent run from the first signal to the change or step that caused it.

9 min read Updated October 2026LLM FinOps
On this page
  1. 01What you'll do
  2. 02The path
  3. 03Step 1: narrow it in the Explorer
  4. 04Step 2: ask what changed
  5. 05Step 3: read what was blocked
  6. 06Step 4: open the session
  7. 07Step 5: fix and record
  8. 08A worked example
  9. 09A second example: a burst of 429s
  10. 10Write it down
  11. 11Prepare before the next one

01

What you'll do

Incidents rarely announce their cause. In about nine minutes you will walk a standard path through the Explorer, Audit, and Sessions to find what happened, who or what changed, and what to do next.

  • Start from a signal: a spike, a block, or a complaint
  • Narrow it to an app, a key, and a time
  • Find the change or step behind it
  • Fix it and record what you learned

02

The path

Work from the widest view to the narrowest.

From signal to cause
  1. 1Signal

    Spend spike, block, or complaint

  2. 2Explorer

    Which app, key, and hour

  3. 3Audit

    What changed, and what was blocked

  4. 4Sessions

    What the agent actually did

  5. 5Fix

    Change a limit, a tool rule, or a key

03

Step 1: narrow it in the Explorer

Find where and when.

  1. 1

    Choose a window

    Use 24h, or a custom range around the signal.

  2. 2

    Group by App

    Find the app that moved.

  3. 3

    Group by Virtual Key

    Find the integration inside that app.

  4. 4

    Note the hour

    Cost Trend shows when the change began.

04

Step 2: ask what changed

Most incidents follow a change.

  1. 1

    Open the Control Plane Audit Log

    Filter to the hour before the signal.

  2. 2

    Look for policy, binding, credential, and key actions

    A loosened budget or a new tool server is a common cause.

  3. 3

    Check the Approval queue

    A change that is still pending has not taken effect.

05

Step 3: read what was blocked

Blocks tell you which rule is acting.

  1. 1

    Open Policy Violations

    Find the group for your app.

  2. 2

    Read the reason

    A budget block, a rate limit, a tool rule, or a guardrail.

  3. 3

    Drill into recent requests

    See what was asked and which limit applied.

Reason familyWhat it usually means
BudgetSpend reached a limit, often because of a loop
Rate limitA burst, often a retry storm
ToolAn agent tried a tool it may not use
GuardrailA prompt or answer matched a rule

06

Step 4: open the session

Sessions show what the agent actually did.

  1. 1

    Open AI → Sessions

    Filter by the app and the time.

  2. 2

    Sort by spend or requests

    The runaway run usually stands out.

  3. 3

    Open the session

    Read the steps in order.

  4. 4

    Look for the loop

    Repeated tool calls with the same arguments, or a call that never gets a result.

With Tool calls retention you can read the arguments and results. With Zero retention you still see the order, models, and cost.

07

Step 5: fix and record

Close the loop so it does not repeat.

CauseFix
An agent loopsSet Max tool calls and Max loop iterations on the policy
A tool does harmDeny it, or disable its server
Spend ran past planLower the daily budget and add a rate limit
A key leakedRevoke it and create a replacement
A change was a mistakeRevert the policy and review who may approve changes

Write one line in your incident log: the signal, the cause, the change. The audit log keeps the proof.

08

A worked example

A support agent costs three times its normal amount on a Thursday afternoon.

StepWhat you find
ExplorerThe support-assistant app, virtual key support-chatbot-prod, from 15:10
AuditNo policy change, but a new tool server was registered at 15:02
Policy ViolationsNo blocks yet. The agent is under its limits
SessionsOne session repeats the same search tool 60 times
FixSet Max tool calls to 20 and deny the noisy tool until the server is fixed

The numbers are an illustration. The path is the point: wide to narrow, change before behavior.

09

A second example: a burst of 429s

Not every incident is about cost.

StepWhat you find
SignalUsers of the search feature report errors at 09:00
ExplorerRequests from the search app triple at 08:55, and Blocked rises with them
AuditNo change since last week
Policy ViolationsA rate-limit group for the search app
SessionsA new batch job shares the search key and sends requests all at once
FixGive the batch job its own key and a lower limit, and add retry with backoff

The numbers are an illustration. The pattern matters: one noisy caller can use a shared allowance, and one key per app prevents it.

10

Write it down

A short record turns an incident into a lesson.

FieldExample
SignalSpend three times normal on Thursday afternoon
CauseA new tool server made the agent repeat a search
EvidenceAudit row, session id, and Policy Violations group
FixMax tool calls set to 20 and the tool denied
Follow-upAdd a loop cap to every agent policy

11

Prepare before the next one

A few settings make the next incident short.

  • Turn on Tool calls retention for your agents' policies
  • Set a daily budget and a rate limit on every agent policy
  • Set Max tool calls on every agent policy
  • Put the Audit tab on your on-call checklist

Put This Guide Into Practice

Cloptima automates the strategies described in this guide.

No credit card required
5-minute setup
Free trial