Evaluating AI agents¶
SignalDB answers one question for teams shipping AI agents: does the new version still behave as specified? You replay a fixed set of test cases (an eval set) on the new version, score every case with your own evaluators — code checks, LLM judges, classifiers — and send the scores to SignalDB. The Evaluate pages then compare the new version with the last one case by case, and link every score to the span it judged.
SignalDB doesn't run your agent or your judges. It stores the results next to the traces of the replay and does the comparing. Offline evaluation (before release) comes first; results sent from production traffic show up under the Production source on the same pages.
What to send¶
Your agent is traced as usual with the OpenTelemetry GenAI conventions: an
invoke_agent {gen_ai.agent.name} span with chat and
execute_tool {gen_ai.tool.name} children.
Each evaluator result is one OTLP log record with
event_name = gen_ai.evaluation.result (the OTel GenAI evaluation event),
carrying the trace id and span id of the span it scores:
| Attribute | Required | Meaning |
|---|---|---|
gen_ai.evaluation.name |
yes | The evaluator, e.g. Correctness. This name is the evaluator. |
gen_ai.evaluation.score.value |
one of both | Numeric score, usually 0–1. |
gen_ai.evaluation.score.label |
one of both | Verdict label, e.g. pass/fail, safe/unsafe. |
gen_ai.evaluation.explanation |
no | The judge's reasoning, shown on the case page. |
error.type |
on failure | Set when the evaluator itself failed (timeout, bad judge output). |
gen_ai.agent.name |
recommended | The agent. Falls back to the record's service.name. |
gen_ai.agent.version |
recommended | The agent version. Falls back to the record's service.version. |
signaldb.eval.run_id |
offline | Groups results into one run. Absent means a production result. |
signaldb.eval.set |
offline | The eval set the run replayed. |
signaldb.eval.case_id |
offline | The case within the set — how two runs are matched. |
signaldb.eval.evaluator |
recommended | Evaluator implementation and version, e.g. trajectory-match@2.1.0. |
signaldb.eval.trial |
no | Trial index when a case runs several times (averaged for now). |
The signaldb.eval.* attributes are SignalDB's own: the OTel conventions
score single responses and have no notion of runs or eval sets yet.
Why a log record? Judges usually run after the agent span has ended, when events can no longer be added to it; a log record with explicit trace context can be sent any time, and several evaluators scoring one span stay separate results. A harness that scores while the span is still open can send span events instead; see Results as span events.
With the OpenTelemetry Python SDK (a release whose LogRecord takes event_name):
from opentelemetry._logs import LogRecord, get_logger
logger = get_logger("eval-harness")
def send_result(scored_span_ctx, case_id, name, score=None, label=None, explanation=None):
logger.emit(LogRecord(
event_name="gen_ai.evaluation.result",
trace_id=scored_span_ctx.trace_id,
span_id=scored_span_ctx.span_id,
trace_flags=scored_span_ctx.trace_flags,
attributes={
"gen_ai.evaluation.name": name,
**({"gen_ai.evaluation.score.value": score} if score is not None else {}),
**({"gen_ai.evaluation.score.label": label} if label else {}),
**({"gen_ai.evaluation.explanation": explanation} if explanation else {}),
"gen_ai.agent.name": "support-triage",
"gen_ai.agent.version": AGENT_VERSION,
"signaldb.eval.run_id": RUN_ID,
"signaldb.eval.set": "triage-golden-200",
"signaldb.eval.case_id": case_id,
"signaldb.eval.evaluator": "trajectory-match@2.1.0",
},
))
Export logs to SignalDB's OTLP endpoint as described in
Sending OTLP. Score the invoke_agent span for
whole-run judgements (correctness, trajectory) and individual
execute_tool or chat spans for per-step checks (valid tool arguments,
toxicity) — the case page shows each result on the span it scored.
Results as span events¶
The GenAI conventions also allow a result as a span event on the span it
scores. SignalDB accepts that form: when a trace export carries a span
event named gen_ai.evaluation.result, the acceptor also stores it as one
log record per event, with the event's time and attributes, the span's
trace id and span id, and the span's resource and scope. The span itself is
stored unchanged. These results appear on the same pages and in the same
logs queries as ones sent as log records.
from opentelemetry import trace
span = trace.get_current_span() # the invoke_agent span, still open
span.add_event(
"gen_ai.evaluation.result",
attributes={
"gen_ai.evaluation.name": "Correctness",
"gen_ai.evaluation.score.value": 0.92,
"signaldb.eval.run_id": RUN_ID,
"signaldb.eval.set": "triage-golden-200",
"signaldb.eval.case_id": case_id,
},
)
If the event has no gen_ai.agent.name or gen_ai.agent.version, the
record takes it from the span's attributes. A value on the event wins.
The record goes through the tenant's log processors, just like a log record you send. It is derived after the tenant's trace processors ran, so any attribute they redacted or rewrote on the span, resource or scope arrives that way on the record too. If the log write fails, the trace export still succeeds: the span is kept and the acceptor logs a warning, but that result is not stored. A resent trace export does not store its results twice.
How results are read¶
- Pass rule. A result with
error.typeis an evaluator error: left out of pass rates and means, never a failure. Otherwise a recognised label decides —pass,passed,true,yes,correct,safepass;fail,failed,false,no,incorrect,unsafefail (case-insensitive). Without one, a score of 0.5 or more passes. Any other result counts toward the mean but has no verdict. - Runs. A run is every result sharing a
signaldb.eval.run_id. It is receiving results until it has been quiet for 10 minutes, then complete, or partial when some results errored or carried no trace context (they still count toward the run's scores). - Comparing runs. Cases are matched on
signaldb.eval.case_id. Per evaluator, a case got worse when it went from pass to fail or its mean dropped by 0.05 or more, and better the other way round. A case is a regression when any evaluator got worse (even if another improved), an improvement when any got better, otherwise unchanged. A case only the candidate ran is a regression marked no baseline when an evaluator fails it, otherwise unchanged.
The pages¶
- Agents & scores (
/evals) — per agent: how many of its runs are scored, the pass rate over all evaluators against the previous window, the evaluator that dropped most, evaluator errors, a mean-score line per evaluator with version markers, and the evaluators table sorted by the biggest drop. An evaluator that sends only labels shows its most common label and that label's share in the Mean column. The source toggle switches between offline runs (default), production results, or both. The page opens on the last 7 days. - Eval sets (
/evals/sets,/evals/sets/{name}) — the stored eval sets with their cases, what they were built from, and how each case scored in the set's newest run. See Eval sets in the UI. - Runs (
/evals/runs) — offline runs with their eval set, version, cases, pass rate and status, filterable by agent and eval set (?set=). Compare › opens the run against the previous run of the same eval set. Upload results… opens the Upload results dialog. - Compare (
/evals/compare?baseline=…&candidate=…) — per evaluator means, deltas, pass rates and how many cases moved, plus latency p95 and tokens per run from theinvoke_agentspans; then the cases, filtered to regressions, improvements or unchanged, with the candidate's tool calls marked against the baseline's (skipped, reordered or repeated, new). Save N regressed cases as eval set turns the regressions into a new set (see Eval sets). - Case (
/evals/compare/case?…&case=…) — the agent trajectory as a timeline with each result on its span, tools the baseline called but the candidate skipped as expected, not called rows, the user input and answer (fromgen_ai.input.messages/gen_ai.output.messageswhen recorded), and one card per evaluator with its explanation. Switch between candidate, baseline and side by side. The breadcrumb reads "Evaluate / Compare / case id", its Compare crumb leading back. - Evaluators (
/evals/evaluators) — every evaluator that sent a result, with the implementation versions seen, the span type it scores, its output kind, and whether it ran offline, in production or both.
Upload a results file¶
A harness that doesn't export OpenTelemetry can upload its results as a file
instead. Each row is one evaluator result; SignalDB turns it into the same
gen_ai.evaluation.result log record described above, so the run shows up on
the Evaluate pages and in the Query IR exactly as if it had been sent over
OTLP.
The file is JSONL (one JSON object per line) or CSV with a header row:
| Column | Required | Becomes |
|---|---|---|
case_id |
yes | signaldb.eval.case_id (at most 128 bytes) |
name |
yes | gen_ai.evaluation.name: the evaluator |
score |
one of | gen_ai.evaluation.score.value (a number) |
label |
one of | gen_ai.evaluation.score.label |
error |
one of | error.type: the evaluator itself failed |
explanation |
no | gen_ai.evaluation.explanation |
trace_id |
no | the record's trace id (32 hex characters); omit for run-level results |
span_id |
no | the record's span id (16 hex characters); needs a trace_id |
evaluator |
no | signaldb.eval.evaluator, e.g. trajectory-match@2.1.0 |
trial |
no | signaldb.eval.trial (a non-negative integer) |
Every row needs a score, a label or an error. An empty CSV cell or a
JSON null counts as absent; other columns or keys are ignored. CSV headers
are matched case-insensitively.
The run comes from the request: the agent (gen_ai.agent.name, and the
records' service.name), its version (gen_ai.agent.version and
service.version), the eval set name (it needn't exist as a stored
eval set), and an optional run id — a UUID is generated when
you leave it out. Every record is stamped with the upload time, plus one
nanosecond per row, so rows keep their file order within the run.
The whole file is checked before anything is written. If any row is invalid,
the upload is rejected with a 400 and nothing is stored; details lists
every problem (the first 100), each with the file line it starts on (a CSV
header is line 1), the column, and why. A missing required CSV column is one
problem naming the column. Files are capped at 100,000 rows and 32 MiB.
curl -sS -X POST \
"$SIGNALDB_URL/api/v1/evals/results?agent=support-triage&version=v1.9.0&set=triage-golden-200" \
-H "Authorization: Bearer $SIGNALDB_API_KEY" -H "X-Tenant-ID: acme" \
-H "Content-Type: text/csv" --data-binary @results.csv
The format comes from the format query parameter (csv or jsonl) when
given, otherwise from the Content-Type: text/csv, or
application/x-ndjson / application/jsonl. The upload needs an API key
with the evals:write scope and answers 201 with the run id and a summary
per evaluator:
{
"run_id": "1f0b7c1e-7c55-4c43-9b8e-3f1a2d6c9e10",
"agent": "support-triage",
"version": "v1.9.0",
"set": "triage-golden-200",
"rows": 600,
"cases": 200,
"span_linked": 596,
"run_level": 4,
"evaluators": [
{
"name": "Correctness",
"results": 200,
"errors": 3,
"mean": 0.87,
"pass_rate": 0.91
}
],
"_links": {
"query": { "href": "/api/v1/query", "method": "POST" },
"runs": { "href": "/evals/runs" }
}
}
mean leaves out evaluator errors; pass_rate is passes / (passes + fails)
under the pass rule, and null when no result has a
verdict. The results appear in queries as soon as the writer commits them
(within seconds).
Durability and retries¶
A 201 means the results are durable: the router forwards them to a writer,
which acks only after they are in its write-ahead log. The router keeps no
log of its own, so an upload that fails or times out (a 5xx, a 504, a
dropped connection) may or may not have been written.
Retrying it is safe as long as you send the same file with the same run id.
The upload's ingest id is a fingerprint of the tenant, dataset, agent,
version, eval set, run id, format and file bytes, and the writer drops a
batch whose ingest id it has already made durable within its dedup window
([writer].ingest_dedup_window, 1 hour by default), so a retry of an upload
that did land is acknowledged without being written twice. Without a run id,
the server generates a new one per request and a retry is a second run; the
CLI and the MCP tool therefore choose the run id themselves when you leave it
out and name it when an upload fails. Uploading a different file under an
existing run id adds its results to that run.
From the CLI, and as a CI gate¶
export SIGNALDB_URL=https://signaldb.example.com SIGNALDB_API_KEY=sk-... SIGNALDB_TENANT_ID=acme
signaldb-cli evals upload results.jsonl \
--agent support-triage --version "$GIT_SHA" --set triage-golden-200 \
--compare-to latest:v1.8.0 \
--fail-if "Correctness.pass_rate < 0.9" \
--fail-if "ToolTrajectory.mean < 0.85"
Without --run-id, the command generates one and prints it before
uploading; if the upload fails without an answer, rerun it with
--run-id <that id> (see Durability and retries).
The command prints the run id, the per-evaluator summary and a link: to the
Compare page (/evals/compare?baseline=…&candidate=…) with --compare-to,
or to the Runs page otherwise. --compare-to takes a run id, or
latest:<version> for the newest other run of the same agent and eval set at
that version in the last 30 days (found with a Query IR read, so the key also
needs logs:read). The format comes from the file extension (.csv,
.jsonl, .ndjson) unless --format csv|jsonl says otherwise.
Each --fail-if is <evaluator>.mean|pass_rate <op> <number> with <,
<=, >, >=, == or !=, checked against the upload's summary. The
command exits non-zero, listing each condition that held, when any does; a
condition naming an evaluator the run doesn't have, or a metric it has no
value for, fails too. Conditions are checked for syntax before anything is
uploaded. The run is uploaded either way, so a blocked release still has its
results to compare.
In GitHub Actions:
- name: Gate on eval results
env:
SIGNALDB_URL: ${{ vars.SIGNALDB_URL }}
SIGNALDB_API_KEY: ${{ secrets.SIGNALDB_EVALS_KEY }}
SIGNALDB_TENANT_ID: acme
run: |
signaldb-cli evals upload eval-results.jsonl \
--agent support-triage --version "${{ github.sha }}" --set triage-golden-200 \
--run-id "gh-${{ github.run_id }}" \
--compare-to latest:${{ vars.RELEASED_VERSION }} \
--fail-if "Correctness.pass_rate < 0.9"
The MCP server offers the same upload as the upload_eval_results tool
(see MCP).
From the Explore UI¶
Upload results… on the Runs page (and in its empty state) opens a dialog with three tabs:
- Upload file — drop or choose a
.jsonlor.csvfile (up to 32 MiB). The browser reads it first and shows the cases, evaluators, span-linked and run-level rows it found, which of the columns above the file has (case_idandgen_ai.evaluation.nameare required), and a warning for rows without atrace_id. The run needs an agent, a version and an eval set name. The version is prefilled from theservice.versionof theinvoke_agentspans the rows link to; when the eval set is a stored one, the dialog says how many of the file's case ids it holds. The set name is required by the endpoint, so an ad-hoc run still needs one; a name that isn't a stored set is accepted and noted. The dialog picks the run id once per chosen file, so uploading the same file again after an error doesn't create a second run. A rejected file lists the server's problems line by line; a successful upload shows the run id, the per-evaluator summary, and links to the run in Runs and to Compare. - CLI / CI — the
signaldb-cli evals uploadcommand above, with a Copy button. - Send over OTLP — the log-record form of one result (see What to send), with a Copy button.
Runs and comparisons outside the UI¶
The Runs and Compare pages have a CLI and an MCP counterpart that read the same Query IR documents and apply the same rules, so a CI job or an AI agent gets the figures a person sees.
signaldb-cli evals runs --agent support-triage # last 7 days, newest first
signaldb-cli evals runs --set triage-golden-200 --from now-30d --json
signaldb-cli evals compare latest:v1.8.0 "$RUN_ID" # baseline, candidate
signaldb-cli evals compare run-0921-0930 run-0927-1004 --tools --limit 20
evals runs lists each run's eval set, agent, version, start time, status,
cases, evaluator errors and pass rate, plus the previous run of the same eval
set (the natural baseline). --agent, --version and --set filter;
--limit (default 50) caps the list and a note says when more runs matched.
--from/--to take RFC3339, relative (now-7d) or epoch-nanosecond times.
evals compare <baseline> <candidate> takes run ids, or latest:<version> for
the newest run of that version of the same agent on the same eval set (the
agent and set come from the other side's run, or from --agent/--set when
both sides are latest:). It reads both runs from the last 30 days by
default and prints per-evaluator means, pass rates and deltas, how many cases
regressed, improved or stayed unchanged, and the regressed cases, largest
drop first (at most --limit, default 50), with the evaluators that got worse
and both trace ids. --tools adds each listed case's tool calls marked
against the baseline's (skipped, reordered, repeated, new); --json prints
the whole comparison. Both commands need logs:read, and --tools also
traces:read.
The MCP server exposes the same reads as list_eval_runs and
compare_eval_runs (see MCP), so an agent such as Claude can answer
"did version B get worse than A, where and why" and follow a regressed case
into its traces.
Querying results yourself¶
Every figure is a Query IR read over logs, so the CLI
and the API can reproduce it. Pass rate and mean per evaluator for one run:
{
"irVersion": 4,
"from": "logs",
"range": { "from": "now-30d", "to": "now" },
"result": "table",
"pipeline": [
{
"where": {
"field": "event_name",
"op": "eq",
"value": "gen_ai.evaluation.result",
},
},
{
"where": {
"field": "signaldb.eval.run_id",
"op": "eq",
"value": "run-0927-1004",
},
},
{
"aggregate": {
"by": [
"gen_ai.evaluation.name",
"gen_ai.evaluation.score.label",
"error.type",
],
"aggs": [
{ "fn": "count", "as": "n" },
{ "fn": "avg", "of": "gen_ai.evaluation.score.value", "as": "mean" },
],
},
},
],
}
signaldb-cli query --ir --file run-scores.json
Eval sets¶
Stored eval sets hold a harness's cases (inputs, expected tool trajectories, reference answers) per tenant and dataset, behind an HTTP API and the Eval sets pages. See Eval sets.