Alerting on Agent Anomalies: Drift, Hallucinations, and Stalls
A practical alerting strategy for production agents — what to detect, what thresholds to set, and how to wire LiteLLM alerting into Slack, email, and webhooks.
Published on • October 8, 2026
AI Assistant

Agents fail quietly. A model upgrade shifts output style, a tool starts returning empty arrays, a retry loop hangs — and the dashboard looks busy the whole time. The fix is alerting that watches for anomalies in behavior, not just errors in status codes.
This post covers three failure classes every agent deployment needs alerts for — drift, hallucinations, and stalls — and how to actually wire them up using an LLM gateway’s alerting layer, using LiteLLM’s proxy alerting as the concrete example.
Why standard monitoring misses agents
Traditional service monitoring assumes request → response, with latency and error rate as the two signals. An agent run breaks that shape:
- A run is many model calls, tool calls, and retries
- The output can be wrong while every HTTP status is
200 - Runs can hang rather than fail — a tool waiting on a lock, a loop that never terminates
- Cost is a first-class failure mode: an agent can burn $40 in 90 seconds and still “succeed”
So you need alerts on the shape of runs, not just their exit codes.
The three anomaly classes
1. Stalls
A stall is a run that stops making progress: no new tokens, no tool completion, no state transition within a window. Stalls are the easiest anomaly to detect and the most expensive to ignore — the user stares at a spinner.
Signals that indicate a stall:
- Time since last token exceeds N seconds
- Time since last completed tool call exceeds N seconds
- Run duration exceeds a hard deadline
- A tool call has been in-flight past its own timeout
LiteLLM’s proxy makes this a first-class config knob. Its alerting docs describe alerting_threshold as the setting that “sends alerts if requests hang for 5min+ and responses take 5min+”:
general_settings:
alerting: ["slack"]
alerting_threshold: 300 # seconds
major_outage_alert_threshold: 10
max_outage_alert_list_size: 10
log_to_console: false
That default of 300 seconds is generous for chat but aggressive for a long research agent. Set it per workload: a support agent should alert at 30–60s of no progress; an overnight research run might legitimately take an hour.
2. Hallucinations
You can’t alert on “the model was wrong” directly — nobody has that metric. You alert on proxies that correlate with wrongness:
- Grounding failures — an answer cites a document ID that doesn’t exist in your retrieval set
- Schema violations — structured output that fails JSON schema validation, repeatedly
- Tool misuse — calling
get_price(product_id)with a product id that returns “not found,” or calling the same tool three times with identical arguments - Refusal rate spikes — sudden increase in
I don't have that information - Self-reported confidence collapse — if your prompt asks for confidence, watch the distribution
- Judge disagreement — an LLM-as-judge score dropping below a rolling baseline
The pattern is the same each time: compute a per-run metric, compare against a rolling window, alert on deviation.
# Pseudocode: per-run hallucination proxy
def hallucination_score(run) -> float:
score = 0.0
if run.schema_violations > 0:
score += 0.4 * run.schema_violations
if run.dangling_citations > 0:
score += 0.5 * run.dangling_citations
if run.repeat_tool_calls > 2:
score += 0.2 * (run.repeat_tool_calls - 2)
return score
baseline = rolling_p95(last_7_days, per_user_segment=True)
if hallucination_score(run) > baseline * 1.5:
alert("hallucination_spike", run_id=run.id, score=..., baseline=...)
Segment the baseline by task type. A code-writing agent and a summarization agent have completely different score distributions; blending them produces noise.
3. Drift
Drift is the slow version: output quality degrades over days or weeks without any single run looking broken. Typical triggers:
- Model provider silently updates a model behind the same alias
- Your prompt or few-shot examples change
- The upstream data distribution shifts (new product names, new document types)
- A retrieval index rebuild changes which chunks surface
Drift alerts are statistical, not threshold-based:
- Distribution shift on output length, tool-call count, or refusal rate (KS test or PSI against a reference window)
- Rollback-triggering change in an eval score — run your eval suite nightly against a golden set and alert on regression past a delta
- Provider-side: alert when the model version string attached to responses changes
That last one is easy to skip and expensive to miss. Pin your model version where you can, and alert on the version field in every response.
Wiring it up: LiteLLM alerting in practice
LiteLLM’s proxy has a built-in alerting layer, which is a good fit if you’re already routing agent traffic through a gateway.
Enable Slack alerts:
export SLACK_WEBHOOK_URL="https://hooks.slack.com/services/<>/<>/<>"
general_settings:
alerting: ["slack"]
alerting_threshold: 300
spend_report_frequency: "1d"
If SLACK_WEBHOOK_URL is unset, LiteLLM falls back to ALERTING_WEBHOOK_URL — a provider-neutral variable you can point at Rocket.Chat, Mattermost, or any Slack-compatible incoming webhook.
Route by alert type. LiteLLM supports mapping each alert category to its own channel via alert_to_webhook_url:
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
alerting: ["slack"]
alert_to_webhook_url:
llm_exceptions: "https://hooks.slack.com/.../llm-exceptions"
llm_too_slow: "https://hooks.slack.com/.../llm-slow"
llm_requests_hanging: "https://hooks.slack.com/.../llm-hangs"
budget_alerts: "https://hooks.slack.com/.../spend"
db_exceptions: "https://hooks.slack.com/.../db"
outage_alerts: "https://hooks.slack.com/.../outage"
cooldown_deployment: "https://hooks.slack.com/.../cooldown"
daily_reports: "https://hooks.slack.com/.../reports"
spend_reports: "https://hooks.slack.com/.../spend-reports"
new_model_added: "https://hooks.slack.com/.../models"
Stalls go to #agent-halls, spend anomalies to #ai-spend, exceptions to #llm-exceptions. The point of routing is that a page-worthy stall doesn’t get buried under daily spend reports.
Add email for budgets. LiteLLM’s email notifications support soft-budget and max-budget alerts (default 80% of max), deduplicated so each threshold fires at most once per EMAIL_BUDGET_ALERT_TTL (default 24 hours). Enable with:
litellm_settings:
callbacks: ["smtp_email"]
general_settings:
alerting: ["slack", "email"]
Keep sensitive content out of alerts. By default alerts include the messages passed to the LLM. Redact them:
general_settings:
alerting: ["slack"]
alert_types: ["spend_reports"]
litellm_settings:
redact_messages_in_exceptions: True
Verify the channel works. LiteLLM exposes a health check:
curl -X GET 'http://0.0.0.0:4000/health/services?service=slack' \
-H 'Authorization: Bearer sk-1234'
Run this in CI or on deploy. An alerting pipeline nobody has tested is worse than none, because it creates false confidence.
Threshold design
Bad alerting is worse than no alerting — teams learn to ignore it. Some rules that hold up:
| Signal | Suggested start | Tune by |
|---|---|---|
| Stall (no progress) | 30–60s chat, 5–15min research | p95 run duration by task |
| Run deadline | p99 historical × 1.5 | weekly review |
| Schema violation rate | >2% of runs in 15min window | eval suite pass rate |
| Cost per run | >3× median for that task | budget per feature |
| Eval regression | >5% drop vs. golden set | nightly suite |
| Model version change | any change | — (zero tolerance) |
Additional practices:
- Alert on windows, not single events. One schema violation is a bug; 5% over 15 minutes is an incident.
- Separate pages from tickets. Stalls and outage-class events page. Drift and eval regressions file tickets for the morning.
- Attach context. Every alert should carry run id, model, task type, and the metric vs. baseline — enough to start debugging without opening a dashboard.
- Deduplicate. Suppress repeat alerts for the same run or the same root cause within a cooldown window.
- Review silence monthly. Any alert that fired 20+ times without action gets a new threshold or gets deleted.
The alert triage loop
Alerting only works if something happens next:
- Detect — gateway/webhook alert or your own anomaly detector
- Enrich — attach run trace, input/output sample, model version, recent deploys
- Classify — stall, hallucination, drift, cost, outage
- Triage — auto-remediate what you can (timeout a hung tool, fail over to a backup model, roll back a prompt version)
- Retrospect — weekly review of alert volume and false-positive rate
Drift and hallucination alerts almost always trace back to a change: a prompt edit, a model alias bump, a retrieval config change. Correlate your alert stream with your deploy stream and the root cause becomes obvious.
Wrapping up
Agents don’t fail the way APIs do, so don’t monitor them the way you monitor APIs. Alert on stalls (no progress), hallucination proxies (schema violations, dangling citations, tool misuse), and drift (distribution shift, eval regression, model version change) — wire those into a gateway with per-type routing, redaction, and tested delivery paths, and you’ll find problems in minutes instead of in support tickets.