Eval sets¶
An eval set is a named, ordered list of test cases for one AI agent. Your
eval harness pulls a set, runs the agent over each case's input, and reports
the results back as gen_ai.evaluation.result records (see
Evaluating AI agents). Keeping the set in SignalDB means
every run of your harness, in CI or on a laptop, tests against the same
cases.
A set belongs to one tenant and one dataset: the tenant and dataset of the
credential that created it (X-Tenant-ID, and X-Dataset-ID or the
tenant's default dataset). A caller in another tenant or another dataset gets
404 for it, and deleting the dataset deletes its sets.
The set and its cases¶
{
"name": "refund-edge-cases-40",
"agent": "support-triage",
"description": "Refund requests the agent got wrong in March",
"cases": [
{
"id": "edge-01",
"input": "I was charged twice for order 1182, can I get one back?",
"expected_tools": ["lookup_order", "issue_refund"],
"reference": "Refunds the duplicate charge on order 1182.",
"tags": ["refunds", "duplicate-charge"],
"source": {
"kind": "trace",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736"
}
},
{
"id": "edge-02",
"input": "Cancel my subscription and refund this month."
}
]
}
| Field | Required | Notes |
|---|---|---|
name |
yes | Lowercase letters, digits, -, _ and ., starting with a letter or digit; 1-128 characters. |
agent |
yes | The agent under test (gen_ai.agent.name). Must not be empty. |
description |
no | |
cases |
no | Kept in the order given. At most 10,000 per set. |
Each case:
| Field | Required | Notes |
|---|---|---|
id |
yes | 1-128 characters, unique within the set. |
input |
yes | What the agent receives. |
expected_tools |
no | The tool names the agent should call, in order. Empty means the trajectory isn't checked. |
reference |
no | A reference answer, for evaluators that compare against one. |
tags |
no | Free-form labels. |
source |
no | {"kind": "trace", "trace_id": "<32 hex>"}, {"kind": "upload"} or {"kind": "hand_written"} (default). |
Trace ids are stored lower-case.
HTTP API¶
All paths are under /api/v1 and take the usual tenant credentials
(Authorization: Bearer <key>, X-Tenant-ID, optional X-Dataset-ID, or a
browser session). Errors use the shared envelope
{"status": "error", "errorType": "...", "error": "..."}.
| Method | Path | Scope | Result |
|---|---|---|---|
GET |
/eval-sets |
evals:read |
Sets in the dataset, without cases, ordered by name |
POST |
/eval-sets |
evals:write |
201 with a Location header; 409 if the name is taken |
GET |
/eval-sets/{name} |
evals:read |
The set with its cases in order |
PUT |
/eval-sets/{name} |
evals:write |
Replaces agent, description and every case; never creates |
DELETE |
/eval-sets/{name} |
evals:write |
204 |
POST |
/eval-sets/{name}/cases |
evals:write |
Appends cases, skipping ids the set already holds |
POST |
/eval-sets/{name}/cases/from-traces |
evals:write, traces:read |
Appends one case per matching agent trace (below) |
Other statuses: 400 for a body that isn't JSON, 422 for a body that
fails validation (bad name, empty agent, duplicate case ids, a trace id that
isn't 32 hex characters, a PUT whose body name differs from the path),
403 without the scope, 404 for a set that doesn't exist in your dataset,
and 413 for a body over 32 MiB. Responses carry _links to the set and,
when you may write, to its replace, delete, append_cases and
append_cases_from_traces actions.
Create a set:
curl -sS -X POST "$SIGNALDB/api/v1/eval-sets" \
-H "Authorization: Bearer $KEY" -H "X-Tenant-ID: acme" \
-H "Content-Type: application/json" \
-d @refund-edge-cases-40.json
List the sets, then pull one into a harness. Each listed set carries
case_count and sources, how many of its cases came from each source kind
({"trace": 162, "upload": 0, "hand_written": 38}):
curl -sS "$SIGNALDB/api/v1/eval-sets" \
-H "Authorization: Bearer $KEY" -H "X-Tenant-ID: acme"
curl -sS "$SIGNALDB/api/v1/eval-sets/refund-edge-cases-40" \
-H "Authorization: Bearer $KEY" -H "X-Tenant-ID: acme" \
| jq -c '.cases[]'
To replace a set, PUT the same body to /api/v1/eval-sets/{name}.
Append cases. Ids already in the set are reported and left as they are:
curl -sS -X POST "$SIGNALDB/api/v1/eval-sets/refund-edge-cases-40/cases" \
-H "Authorization: Bearer $KEY" -H "X-Tenant-ID: acme" \
-H "Content-Type: application/json" \
-d '{"cases": [
{"id": "edge-01", "input": "..."},
{"id": "edge-41", "input": "Refund the shipping fee only."}
]}'
{
"added": 1,
"already_present": 1,
"added_ids": ["edge-41"],
"already_present_ids": ["edge-01"]
}
Delete a set:
curl -sS -X DELETE "$SIGNALDB/api/v1/eval-sets/refund-edge-cases-40" \
-H "Authorization: Bearer $KEY" -H "X-Tenant-ID: acme"
The Rust SDK (signaldb-sdk) exposes the same operations as
list_eval_sets, create_eval_set, get_eval_set, replace_eval_set,
delete_eval_set, append_eval_cases and append_eval_cases_from_traces.
Build cases from traces¶
POST /api/v1/eval-sets/{name}/cases/from-traces turns production traces
into cases: one case per matching agent trace the set doesn't hold yet.
Sample the runs of support-triage that failed Correctness in the last
7 days:
curl -sS -X POST "$SIGNALDB/api/v1/eval-sets/refund-edge-cases-40/cases/from-traces" \
-H "Authorization: Bearer $KEY" -H "X-Tenant-ID: acme" \
-H "Content-Type: application/json" \
-d '{
"range": {"from": "now-7d", "to": "now"},
"agent": "support-triage",
"failing_evaluator": "Correctness",
"sample": 50,
"expected_tools": true,
"tags": ["correctness-failures"]
}'
{
"matches": 214,
"already_present": 12,
"added": 50,
"added_ids": ["trace-4bf92f3577b34da6", "..."]
}
matches counts the distinct matching traces, already_present those the
set already has a case for, and added_ids lists the new cases in set
order.
| Field | Default | Notes |
|---|---|---|
range |
required | The window, as in a Query IR document: RFC3339, now-7d, or epoch nanoseconds as a string. |
agent |
the set's agent | gen_ai.agent.name of the agent span. A span without one matches on service.name. |
operation |
invoke_agent |
gen_ai.operation.name of the agent span. |
filters |
none | Extra Query IR predicates on the agent span, e.g. {"field": "deployment.environment", "op": "eq", "value": "prod"}. |
failing_evaluator |
none | Keep only traces with at least one failing gen_ai.evaluation.result of this evaluator (gen_ai.evaluation.name) in the window. Needs logs:read. |
sample |
50 | How many cases to add, 1-1000. |
expected_tools |
false |
Set the case's expected tools to the trace's execute_tool spans' gen_ai.tool.names, in call order. |
reference_from_answer |
false |
Set the case's reference to the agent's answer: the last text of the agent span's gen_ai.output.messages. |
tags |
none | Tags put on every new case. |
A result fails by the same rule the Evaluate pages use (see
Evaluating AI agents): a result with error.type is an
evaluator error and never a failure; otherwise a pass/fail-style label
decides; otherwise a score below 0.5 fails.
Each new case:
idistrace-plus the first 16 hex digits of the trace id, so running the same query again never duplicates a case;inputis the last user text of the agent span'sgen_ai.input.messages(a JSON array of{role, parts}), or the attribute as-is when it isn't that shape; empty when the span has none;sourceis{"kind": "trace", "trace_id": "..."}.
Which traces are taken is deterministic: newest agent span first (ties by
trace id), traces already a case source in the set skipped, then the first
sample. Unknown body fields are rejected (422), as are a sample outside
1-1000, an empty agent/operation/failing_evaluator, an unparseable or
empty range, and a filter the query engine refuses. The reads are three
Query IR queries run on the server, bounded to 10,000 matching traces,
50,000 evaluator results and 100,000 spans; matches never exceeds 10,000.
In the Explore UI¶
The Eval sets page (/evals/sets, in the Evaluate group) lists the
dataset's sets: name and description, agent, case count, what the cases
were built from (the list's sources counts: "traces 162 · hand-written
38", "JSONL upload", or "saved from Compare" for a set whose description is
the one Compare writes, "Cases that regressed in …"),
the newest run of the set in the last 30 days with its version, date and
pass rate (a Query IR read; "never run" otherwise), and
when the set last changed. New eval set… asks for a name and agent and
starts the set from:
- Upload JSONL — a file of cases in the case format.
The dialog previews the first rows, counts cases without a reference
answer, and lists every line it can't read; it won't create the set until
the file reads cleanly. Cases without a
sourceare stored as{"kind": "upload"}. - Real traces — creates the set, then builds cases from traces with the agent, window (24 hours, 7 days or 30 days), optional failing evaluator, sample size and the two checkboxes (expected tools from the tools called, on by default; the agent's answer as the reference, off).
- Empty — just the name and agent.
A set's page (/evals/sets/{name}) shows its cases (id, input, expected
tools, reference, source — a trace link for cases built from traces — and
the case's score in the set's newest run: pass when every evaluator
passed, N failing when any failed, P of E when some evaluators reached
no verdict), 50 at a time. Export JSONL downloads the cases in the same
format signaldb-cli eval-sets export writes. Add traces… opens a panel
that builds cases from traces; the endpoint has
no dry run, so the panel shows the matches, already-present and added
counts it returned. The strip under the header links the two newest runs of
the set in Compare, and the Runs tab opens Runs filtered to the set.
Settings shows the set's fields and deletes it. Buttons that write are
disabled when the credential may not (the set's _links carry no write
actions).
On Compare, Save N regressed cases as eval set creates a set (named
regressions-MMDD after the candidate run's day, editable) holding the
regressed cases' inputs, expected tools and references, copied by case id
from the eval set the runs replayed; each case's source is the candidate's
trace when it has one, and every case is tagged saved-from-compare. The
button is disabled, with the reason, when the runs name no eval set or that
set isn't in the dataset's list; the set's cases are read when the dialog
opens, and the dialog says so if that read fails.
CLI¶
Reads use signaldb-cli eval-sets, writes signaldb-cli admin eval-sets.
Both authenticate with a tenant API key (--api-key, --tenant-id,
optional --dataset-id, or the SIGNALDB_* environment), not the instance
admin key.
signaldb-cli eval-sets list # NAME AGENT CASES (per source) UPDATED DESCRIPTION
signaldb-cli eval-sets get refund-edge-cases-40 # the set, then one row per case
signaldb-cli eval-sets export refund-edge-cases-40 > cases.jsonl
signaldb-cli admin eval-sets create --file refund-edge-cases-40.json
signaldb-cli admin eval-sets replace refund-edge-cases-40 --file refund-edge-cases-40.json
signaldb-cli admin eval-sets append refund-edge-cases-40 --file new-cases.jsonl
signaldb-cli admin eval-sets add-traces refund-edge-cases-40 --range 7d \
--failing-evaluator Correctness --sample 50 --expected-tools --tag correctness-failures
signaldb-cli admin eval-sets delete refund-edge-cases-40
list and get print tables; add --json for the raw response. export
writes the set's cases to stdout as JSONL, one case object per line, in
order: the same shape append reads.
create and replace take the set as a YAML or JSON file (.json is read
as JSON, anything else as YAML) in the shape shown
above. For replace, a file without a name
takes the one on the command line, and a file naming a different set is
refused. append takes a JSONL file of cases or a JSON {"cases": [...]}
object, and prints how many cases were added and how many were already
present.
add-traces builds cases from traces. --range
is how far back to look, ending now (7d, 24h; default 7d). The other
flags mirror the body fields: --agent, --operation,
--failing-evaluator, --sample, --expected-tools,
--reference-from-answer, --tag (repeatable) and --filter (a Query IR
predicate as JSON, repeatable).
MCP tools¶
list_eval_sets and get_eval_set (evals:read), and create_eval_set,
replace_eval_set, delete_eval_set, append_eval_cases and
append_eval_cases_from_traces (evals:write; the last also needs
traces:read, plus logs:read with failing_evaluator). The from-traces
tool takes the body fields above, with the
window as from (default now-7d) and to (default now). Each takes the tenant argument and a required dataset:
one MCP session may span several datasets, so there is no implicit default.
get_eval_set returns one page of cases (offset, default 0; limit,
default 200, at most 1000) with the set header plus total_cases, offset,
returned and has_more; page through a large set by raising offset. See
the MCP tool catalogue.
Scopes¶
| Scope | Grants |
|---|---|
evals:read |
Listing and reading eval sets |
evals:write |
Creating, replacing, deleting and appending to eval sets |
Building cases from traces also reads signal data: it needs traces:read,
and logs:read with failing_evaluator. A harness that only pulls sets
needs a key with evals:read. Keys that
predate scopes are unrestricted. In a browser session, every tenant role can
read sets and only tenant admins (and instance admins) can change them.
evals:read is part of the default OAuth grant; evals:write is never
grantable through OAuth. See Authentication.