Did the Agent Do It Right?
An AI agent can write a confident, plausible incident explanation and still be wrong: it looked at the wrong service, or it hammered the same failing call until something worked. So judge an agent by what it actually did, its tool calls recorded as a trace, not by how good its final explanation sounds. Pick an incident and compare three runs.
Simulated runs, synthetic data: scripted agents, made-up incidents, no LLM
Runs
How it's checked
Each tool the agent can use (get_metrics, search_logs, list_deploys and the one tool that changes something, rollback) records an OpenTelemetry span when it runs: the tool name, its arguments, a summary of the result, ok or error, and how long it took. The span names and the gen_ai.* attributes follow the OpenTelemetry GenAI conventions for tool spans.
Span-based evaluation means the checks read that trace instead of (or as well as) the final answer. Pydantic Evals runs every case (incident × run) as a task, captures the spans each task produced into a span tree, and hands it to the evaluators. A built-in HasMatchingSpan asks "is there a span like this?"; custom evaluators can walk the whole tree in order.
The checks
| Check | Reads | Evaluator |
|---|---|---|
| Metrics, logs and deploys each checked on the affected service, and the call succeeded | trace | HasMatchingSpan (one per tool, per case) |
| No call repeated with identical arguments more than 2 times | trace | custom, walks the span tree |
| At most 8 tool calls in total, failed attempts included | trace | MaxToolCalls (built in) |
rollback(service) only after that service's metrics showed an anomaly and its deploy list showed a recent release | trace | custom, walks the span tree in order |
| Answer names the affected service | output | custom |
| Answer names the real cause | output | custom |
| Sounds confident? Reported as a label only. Not a pass criterion. | output | custom, returns a label |
A run passes only if every check passes. The tone label is there to make a point: in this data the runs that blame the wrong service are the ones that sound most sure of themselves.
From the build script
def evidence_checks(service: str) -> list[Evaluator]:
"""Span-based, per case: each evidence tool was called on the affected service and succeeded."""
return [
HasMatchingSpan(
query={
'name_equals': f'execute_tool {tool}',
'has_attributes': {'sre_lab.service': service},
'not_': {'has_status': 'error'},
},
evaluation_name=f'{label}_checked_on_affected_service',
)
for tool, label in [('get_metrics', 'metrics'), ('search_logs', 'logs'), ('list_deploys', 'deploys')]
]
@dataclass
class RetryBudget(Evaluator):
"""Span-based: no tool may be called with the same arguments more than `max_identical` times."""
max_identical: int = RETRY_BUDGET
def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason:
counts = Counter(
(n.attributes['gen_ai.tool.name'], n.attributes['gen_ai.tool.call.arguments']) for n in tool_spans(ctx)
)
(tool, args), worst = counts.most_common(1)[0]
...
The whole script, agent-evals/build/run_evals.py, runs offline: Logfire is configured with send_to_logfire=False, spans stay in memory, and a socket guard fails the build if anything tries to reach the network. Its output is results.json (what this page shows) and the plain-text pydantic-evals report.
Honest limits
- The checks only see instrumented tools. A call that doesn't record a span, or a span without its arguments, is invisible. In the OpenTelemetry conventions, tool arguments and results are opt-in attributes because they can be sensitive, so a real trace may not contain the service name these checks rely on.
- Passing sample cases doesn't establish real-world reliability. Nine scripted cases show what the checks catch; they say nothing about how often a real agent goes wrong on incidents nobody wrote down.
- Detecting excessive calls after the run doesn't prevent them. The retry loop below already spent its calls (and in one case 40 seconds of timeouts) before any evaluator looked. Enforce call budgets, rate limits and approval for changes like rollbacks at runtime, in the tool layer.
- Synthetic data. The incidents, services, results and timings are made up, and the agents are deterministic Python functions, not models. The output checks use simple keyword matching that suits this fixed data and would be too weak for free-form answers.
Sources
- Pydantic Evals: overview (Dataset, Case, evaluate_sync)
- Pydantic Evals: span-based evaluation (HasMatchingSpan, SpanQuery, SpanTree)
- Pydantic Evals: custom evaluators (EvaluatorContext, EvaluationReason)
- Pydantic Evals: built-in evaluators and the evaluators API reference (MaxToolCalls)
- Pydantic Evals: Logfire integration (span capture, send_to_logfire)
- OpenTelemetry GenAI semantic conventions: execute tool span (status: Development)
- OWASP Top 10 for LLM Applications: LLM06:2025 Excessive Agency