On this page
01
What you'll do
Incidents rarely announce their cause. In about nine minutes you will walk a standard path through the Explorer, Audit, and Sessions to find what happened, who or what changed, and what to do next.
- Start from a signal: a spike, a block, or a complaint
- Narrow it to an app, a key, and a time
- Find the change or step behind it
- Fix it and record what you learned
02
The path
Work from the widest view to the narrowest.
1Signal
Spend spike, block, or complaint
2Explorer
Which app, key, and hour
3Audit
What changed, and what was blocked
4Sessions
What the agent actually did
5Fix
Change a limit, a tool rule, or a key
03
Step 1: narrow it in the Explorer
Find where and when.
- 1
Choose a window
Use 24h, or a custom range around the signal.
- 2
Group by App
Find the app that moved.
- 3
Group by Virtual Key
Find the integration inside that app.
- 4
Note the hour
Cost Trend shows when the change began.
04
Step 2: ask what changed
Most incidents follow a change.
- 1
Open the Control Plane Audit Log
Filter to the hour before the signal.
- 2
Look for policy, binding, credential, and key actions
A loosened budget or a new tool server is a common cause.
- 3
Check the Approval queue
A change that is still pending has not taken effect.
05
Step 3: read what was blocked
Blocks tell you which rule is acting.
- 1
Open Policy Violations
Find the group for your app.
- 2
Read the reason
A budget block, a rate limit, a tool rule, or a guardrail.
- 3
Drill into recent requests
See what was asked and which limit applied.
| Reason family | What it usually means |
|---|---|
| Budget | Spend reached a limit, often because of a loop |
| Rate limit | A burst, often a retry storm |
| Tool | An agent tried a tool it may not use |
| Guardrail | A prompt or answer matched a rule |
06
Step 4: open the session
Sessions show what the agent actually did.
- 1
Open AI → Sessions
Filter by the app and the time.
- 2
Sort by spend or requests
The runaway run usually stands out.
- 3
Open the session
Read the steps in order.
- 4
Look for the loop
Repeated tool calls with the same arguments, or a call that never gets a result.
With Tool calls retention you can read the arguments and results. With Zero retention you still see the order, models, and cost.
07
Step 5: fix and record
Close the loop so it does not repeat.
| Cause | Fix |
|---|---|
| An agent loops | Set Max tool calls and Max loop iterations on the policy |
| A tool does harm | Deny it, or disable its server |
| Spend ran past plan | Lower the daily budget and add a rate limit |
| A key leaked | Revoke it and create a replacement |
| A change was a mistake | Revert the policy and review who may approve changes |
Write one line in your incident log: the signal, the cause, the change. The audit log keeps the proof.
08
A worked example
A support agent costs three times its normal amount on a Thursday afternoon.
| Step | What you find |
|---|---|
| Explorer | The support-assistant app, virtual key support-chatbot-prod, from 15:10 |
| Audit | No policy change, but a new tool server was registered at 15:02 |
| Policy Violations | No blocks yet. The agent is under its limits |
| Sessions | One session repeats the same search tool 60 times |
| Fix | Set Max tool calls to 20 and deny the noisy tool until the server is fixed |
The numbers are an illustration. The path is the point: wide to narrow, change before behavior.
09
A second example: a burst of 429s
Not every incident is about cost.
| Step | What you find |
|---|---|
| Signal | Users of the search feature report errors at 09:00 |
| Explorer | Requests from the search app triple at 08:55, and Blocked rises with them |
| Audit | No change since last week |
| Policy Violations | A rate-limit group for the search app |
| Sessions | A new batch job shares the search key and sends requests all at once |
| Fix | Give the batch job its own key and a lower limit, and add retry with backoff |
The numbers are an illustration. The pattern matters: one noisy caller can use a shared allowance, and one key per app prevents it.
10
Write it down
A short record turns an incident into a lesson.
| Field | Example |
|---|---|
| Signal | Spend three times normal on Thursday afternoon |
| Cause | A new tool server made the agent repeat a search |
| Evidence | Audit row, session id, and Policy Violations group |
| Fix | Max tool calls set to 20 and the tool denied |
| Follow-up | Add a loop cap to every agent policy |
11
Prepare before the next one
A few settings make the next incident short.
- Turn on Tool calls retention for your agents' policies
- Set a daily budget and a rate limit on every agent policy
- Set Max tool calls on every agent policy
- Put the Audit tab on your on-call checklist