MCP server (AI agent access)¶
SignalDB ships a Model Context Protocol server, signaldb mcp (a subcommand of the signaldb binary),
that lets AI agents (Claude Code, Claude.ai, IDE assistants) query your traces
directly. It is a thin, credential-forwarding client: it validates the bearer
token you present and forwards that same token to the router's HTTP API. It
holds no credential of its own, so a request can only ever see what your key is
already allowed to see — tenant isolation stays enforced by the router.
What it exposes¶
Neither family below is hidden from tools/list — a call your credential
does not authorize comes back as a clean access-denied tool error, not a
missing tool.
Query and discovery¶
Available to every authenticated tenant session — there is no role gating on these:
| Tool | Purpose |
|---|---|
server_info |
Confirm connectivity and which tenant your credential resolves to. |
connection_info |
This deployment's public OTLP gRPC/HTTP and Prometheus remote-write endpoints, the query API base, headers with your tenant/dataset filled in, the API-key scopes ingest and query need, and ready-to-paste OTEL_EXPORTER_OTLP_* env vars. Takes optional tenant and dataset; a multi-tenant credential must pass tenant. |
discover_datasets |
The tenant and datasets your credential can access, as a nested Markdown list, marking the session's current default dataset. Filtered to your credential's dataset restriction, if any — a dataset outside it never appears, even by name. Call before passing an explicit dataset or tenant argument elsewhere. |
search_traces |
TraceQL search over your tenant's traces. |
get_trace |
Fetch a single trace by ID, with each span's kind (renders as a waterfall — see below). Searches the last 30 days unless start/end are given. |
search_trace_groups |
The grouped RED-metrics view (count, error count, p50/p95 duration, last-seen per group) the UI's traces tab "group by" table shows, along group_by dimensions. No free-text query filter yet — use query_ir for scoped filtering or a custom aggregate. |
get_service_map |
The service dependency graph — nodes per service/external dependency, edges for the calls between them — optionally scoped to one service's neighbourhood (service/depth). Wraps the Query IR graph envelope; returns the graph plus a summary naming the busiest and highest-error edges, and a web UI link (renders as an interactive map — see below). |
get_profile |
Fetch a single profile's flamegraph by ID (renders as an interactive flamegraph — see below). |
get_source_context |
Fetch a source-code snippet around a stack-frame location (path/line) through the tenant's linked GitHub App installation(s). repository may be omitted to probe by path alone; ref may be omitted for the default branch. status: "unavailable" with a reason is a normal answer, not an error. |
discover_attributes |
List queryable attribute/label names, or the values for one. Signal-aware: traces (default, Tempo tags), logs (Loki labels), metrics (Prometheus labels), profiles (Pyroscope labels). |
discover_metrics |
List the distinct metric names visible to your tenant. |
discover_fields |
The queryable fields of a signal source, as logical dotted OTel names with type, origin, coverage and approximate cardinality. Answered from the schema registry and maintained statistics — reads no signal data. |
discover_field_values |
Value suggestions for one field. Exact and free for a declared value set; otherwise it names the query that would answer it, and only sample: true runs that query. |
discover_sources |
The signal sources available to your tenant, with whether each is queryable. |
discover_profile_types |
List the Pyroscope profile types with data for your tenant (e.g. CPU, heap). |
query_metrics |
PromQL query over your tenant's metrics (native Prometheus result); instant by default, or a range query when start/end (and optionally step) are given. |
search_logs |
LogQL query over your tenant's logs (native Loki result); instant or range, same as query_metrics. |
search_profiles |
Search profiles with a Pyroscope selector and a time range; returns the aggregated flame graph (flamebearer encoding). |
compare_profiles |
Compare profiles between two time ranges with a shared Pyroscope selector; returns the differential flame graph. |
profiles_for_trace |
List the profiles correlated with a trace id. |
query_ir |
Native Query IR document (the structured, versioned query surface). |
list_skills |
List the longer-form guidance documents ("skills") this server exposes beyond the tool descriptions — same catalog as the skill://index.json resource. |
get_skill |
Read one guidance document by name (e.g. query-ir) — same content as the corresponding skill://<name>/SKILL.md resource, for clients that don't read MCP resources. |
list_schema_registries |
List the schema registries visible to your tenant in precedence order (custom first, then the bundled signaldb and otel semconv), with definition counts. |
get_schema_registry |
Fetch one registry's summary and full document by namespace/version. |
resolve_attribute |
What an attribute key means: every definition across the visible registries, precedence-ordered (primary first), with brief, type, examples, deprecation. |
resolve_entity |
What an entity type (k8s.pod, service, ...) means: identifying/descriptive attributes, what it extends, associated metrics. |
resolve_metric |
What a metric means: instrument, unit, brief, recorded attributes, associated entities. |
search_schema |
Prefix search over attributes, entities, or metrics (kind, prefix, limit) to find the right vocabulary before querying, or keys to batch-resolve an exact attribute/metric name set instead. |
create_schema_registry |
Upload a custom Weaver-model registry document (JSON object) for your tenant (requires schema:write). |
replace_schema_registry |
Replace a custom registry's document by namespace/version (requires schema:write; bundled registries refuse). |
validate_schema_registry |
Validate a registry document without storing it; errors carry document paths (requires schema:write). |
delete_schema_registry |
Delete a custom registry by namespace/version (requires schema:write; bundled registries refuse). |
list_processors |
List the tenant's OTTL telemetry processors (requires processors:read). See Processors. |
get_processor |
Fetch one processor by name (requires processors:read). |
validate_processor |
Compile-check a processor specification without storing it; errors carry statement index, column, and message (requires processors:read). |
test_processor |
Dry-run a processor set against an OTLP JSON payload — transformed payload plus per-statement match/error counts; never writes to the WAL or forwards data (requires processors:read). |
create_processor |
Create a processor (requires processors:write; tenant-admin only, not OAuth-grantable). |
replace_processor |
Replace a processor's full document by name (requires processors:write; tenant-admin only, not OAuth-grantable). |
delete_processor |
Delete a processor by name (requires processors:write; tenant-admin only, not OAuth-grantable). |
list_eval_sets |
List the agent eval sets in a dataset, without their cases (requires evals:read). See Eval sets. |
get_eval_set |
Fetch one eval set by name with a page of its cases in order: offset (default 0) and limit (default 200, at most 1000); the reply adds total_cases, offset, returned and has_more (requires evals:read). |
create_eval_set |
Create an eval set: name, agent, description, cases (requires evals:write; not OAuth-grantable). |
replace_eval_set |
Replace an eval set's agent, description and every case by name; never creates (requires evals:write; not OAuth-grantable). |
delete_eval_set |
Delete an eval set and its cases by name (requires evals:write; not OAuth-grantable). |
append_eval_cases |
Append cases to an eval set; ids it already holds are skipped and reported (requires evals:write; not OAuth-grantable). |
append_eval_cases_from_traces |
Append one case per matching agent trace the set doesn't hold yet, newest first, up to sample (default 50); failing_evaluator keeps traces with a failing result of that evaluator. Reports matches, already_present, added (requires evals:write and traces:read; not OAuth-grantable). |
upload_eval_results |
Upload a JSONL/CSV results file (content, format) as one offline run of agent/version/set (run_id optional; generated and named in errors, so a retry with it is safe). Any invalid row rejects the file and every problem is listed; returns the run id and per-evaluator summary (requires evals:write; not OAuth-grantable). |
list_eval_runs |
Offline eval runs in a window (default now-7d..now), newest first, filtered by agent, version, set: run id, eval set, agent, version, started_at/last_result_at, status (running/complete/partial with reasons), results, cases, errors, unlinked results, pass rate, per-evaluator mean and pass rate, and previous_run_id (the natural baseline). At most limit runs (default 50, max 500); total_runs and truncated say when more matched (requires logs:read). |
compare_eval_runs |
Compare a candidate run with a baseline run (each a run id or latest:<version>; agent/set default to the other side's run) over a window (default now-30d..now): both runs' summaries, per-evaluator means, pass rates, deltas and cases that moved, counts of regressions/improvements/unchanged, and the regressed cases (largest drop first, at most limit, default 50, max 500) with the evaluators that got worse and both trace ids. include_tools: true adds each listed regression's tool-call diff (requires logs:read, plus traces:read with include_tools). |
Every processor tool takes the tenant parameter validated the same way as
the query tools above. Changes apply at ingest within
[processors].reload_interval (default 30s) — see Processors.
The eval set tools, upload_eval_results, list_eval_runs and
compare_eval_runs take the same tenant parameter and, like the query
tools, a required dataset: a set or a run belongs to one dataset, and one
MCP session may span several, so there is no implicit default. See
Upload a results file and
Runs and comparisons outside the UI.
list_eval_runs and compare_eval_runs answer "did version B of my agent get
worse than A, where and why" with the same figures as the Runs and Compare
pages: they are Query IR reads over the gen_ai.evaluation.result log records
(and, for tool diffs, the execute_tool spans), so they need the read scopes
of those signals rather than evals:read. A typical session lists the runs,
compares the newest with its previous_run_id (or latest:<version>), then
opens a regressed case's traces with get_trace.
Each query tool requires a dataset argument, targeting the dataset your
tenant may access (the router validates access and rejects the rest). Large
results are capped and returned with a truncated: true flag telling the
agent to narrow the query.
Most of these tools also require a tenant argument. For a credential
authorized for exactly one tenant (an API key, or an OAuth connector granted a
single tenant — see below), it's a confirmation check, not a way to switch
tenants: it must equal the tenant this specific call's credential resolves
to (server_info, discover_datasets), and a mismatch fails the call with an
error naming both tenants, before any request reaches the router. For an OAuth
connector granted more than one tenant, tenant is a real selector: pass
the tenant this call should run against, and the server validates it against
the connector's own granted set before forwarding the call — naming a tenant
outside that set fails the same way a mismatch does for a single-tenant
credential. Both arguments are required rather than optional because one MCP
session (one mcp-session-id) can hold credentials for several tenants and
datasets across its calls — there is no single implicit session-wide default
left to fall back to. To reach a second tenant within one session using
separate API-key credentials, present a different credential (Authorization
bearer token plus X-Tenant-ID) on a later call rather than opening a second
connection; the router independently authenticates each call, up to a bounded
number of distinct identities per session. A multi-tenant OAuth connector
reaches its other tenants by simply naming them in tenant on later calls —
no second credential needed, since they all belong to the same grant.
discover_datasets/server_info reflect this too: for a single-tenant
credential they report that one tenant, as before; for a multi-tenant OAuth
connector they enumerate every granted tenant (and, for discover_datasets,
each one's own datasets) so an agent can see the whole reachable set before
picking a tenant for a later call.
Operational control¶
Admin-authenticated (the administrative API key, not a tenant key):
| Tool | Purpose |
|---|---|
compact_run |
Trigger a compaction pass now. |
compact_status |
Active compaction leases and metrics. |
compact_dry_run |
Plan compaction candidates without executing. |
Platform administration¶
Unprefixed, admin-authenticated (the administrative API key can manage
any tenant — this is the same credential the admin CLI group and the
admin HTTP API use):
| Tool | Purpose |
|---|---|
list_tenants / get_tenant |
List every tenant, or fetch one by ID. |
create_tenant / update_tenant |
Create a tenant, or update its name/default dataset. |
delete_tenant |
Delete a tenant and everything under it. Destructive: requires confirm equal to tenant_id. |
create_user |
Create a human user and grant an initial tenant membership. |
list_datasets / create_dataset |
List or create a tenant's datasets. |
delete_dataset |
Delete a dataset by ID. Destructive: requires confirm equal to dataset_id. |
list_api_keys |
List a tenant's API keys with their scopes and dataset set (or that they're unrestricted). Raw secrets are never returned. |
create_api_key |
Create an API key carrying explicit scopes (required; e.g. traces:write, schema:read), an optional dataset_ids set, and an optional allowed_origins set restricting browser (CORS) ingest — see Origin restriction. The raw secret is returned exactly once, in this response. |
update_api_key_scopes |
Change a live key's scopes, dataset set, and/or origin set without rotating its secret; dataset_ids/allowed_origins replace their respective restrictions, clear_dataset_restriction/clear_allowed_origins: true (with no corresponding set field) remove them, and omitting a pair leaves it unchanged — sending a set and its own clear flag together is rejected before any request is made. Revoked keys are rejected. |
revoke_api_key |
Revoke an API key by ID. Destructive: requires confirm equal to key_id. |
Tenant self-management¶
tenant_-prefixed; act as the caller's own identity within its own tenant.
Two sub-groups, by which endpoint they wrap:
Tenant view, tables, and schemas (tenant self-service API) — work with a plain tenant API key, exactly like the query tools above:
| Tool | Purpose |
|---|---|
tenant_info |
The caller's own tenant: id, enabled flag, schema configuration (mirrors signaldb-cli tenant show). |
tenant_list_tables |
List the tenant's provisioned signal tables. Filtered to your credential's dataset restriction, if any, the same way discover_datasets is. |
tenant_create_tables |
Provision (create) the tenant's enabled signal tables — the manual trigger from table provisioning. |
tenant_list_table_schemas |
List the tenant's configured table schema types (distinct from tenant_list_tables, which lists what is actually provisioned). |
list_available_table_schemas |
List every table schema type SignalDB knows how to provision, regardless of tenant configuration. |
Datasets, API keys, memberships, and schema (management API) — need a
management credential: either a human session (a browser session cookie, or
an OAuth access token — see Claude.ai and ChatGPT
below) holding the tenant-admin role or the instance-admin flag, or an API
key that carries the tenant:manage scope for that tenant. Ingest-only
keys and legacy unscoped keys are denied — a deliberate privilege boundary:
an ingest key can write signal data and provision tables, but minting or
revoking other API keys, deleting datasets, or changing memberships is
opt-in, granted only to a tenant admin or a key explicitly scoped for it.
Calling one of these with a key that lacks the scope returns a clean
access-denied error naming the required scope rather than succeeding or
404ing. tenant:manage is never granted through OAuth consent; see
API-key scopes.
| Tool | Purpose |
|---|---|
tenant_list_datasets / tenant_create_dataset |
List or create the caller's own tenant's datasets. |
tenant_delete_dataset |
Delete a dataset by name. Destructive: requires confirm equal to dataset_name. |
tenant_list_api_keys |
List the caller's own tenant's API keys. Raw secrets are never returned. |
tenant_create_api_key |
Create an API key for the caller's own tenant. The raw secret is returned exactly once. |
tenant_update_api_key |
Update the scopes, dataset set, and/or origin set of one of the caller's own tenant's API keys — same dataset_ids/allowed_origins/clear-flag semantics as update_api_key_scopes above. |
tenant_revoke_api_key |
Revoke one of the caller's own tenant's API keys. Destructive: requires confirm equal to key_id. |
tenant_list_memberships / tenant_upsert_membership |
List the caller's own tenant's memberships, or create/update a member's role. |
tenant_remove_membership |
Remove a member from the caller's own tenant. Destructive: requires confirm equal to user_id. |
tenant_get_schema |
The registered logical (client-visible) and physical (storage) schema for every signal source. Takes tenant_id. |
tenant_start_github_link |
Start linking a GitHub App installation. Returns an install_url to open in a browser signed in to SignalDB as an admin of this tenant, and its expires_at. |
tenant_attach_github_installation |
Attach a GitHub App installation that already exists (e.g. linked to another tenant on the same GitHub account) directly, without the OAuth install flow — GitHub skips the consent screen and any redirect back once an account's App installation already exists, so tenant_start_github_link cannot complete a second tenant's link there. Requires instance-admin, unlike every other tool in this section: a tenant:manage grant alone is not enough, since this path has no GitHub code to verify the caller actually controls the installation being linked. |
tenant_list_github_installations |
List the caller's own tenant's linked GitHub App installations, with their repositories. |
tenant_remove_github_installation |
Remove a linked GitHub App installation. Destructive: requires confirm equal to installation_id. |
Destructive tools carry the MCP destructiveHint annotation; read-only tools
carry readOnlyHint — a client that inspects tools/list annotations can
tell which is which without trying the call.
When the router throttles a tool's downstream request (429, the tenant's
query rate limit), the server first lets the SDK's shared retry policy absorb
it — waiting the server-stated Retry-After within the policy's bounds (see
client retry) — so a brief burst is invisible to the
agent. Only once retries are exhausted (or the server asks for more than the
policy's 10 s per-attempt ceiling) does the tool fail, with a throttled
error distinct from an internal failure: the message starts with
throttled: and names the wait (throttled: search_logs was rate limited; the
server asked to retry in 30s), and the error data carries retryAfterMs
(milliseconds; null when no wait was stated) plus http_status: 429. An
agent should wait that long or narrow the query.
When the router rejects a query-backed tool's request outright (query_ir,
search_trace_groups, get_service_map, get_profile, discover_fields,
discover_field_values), the tool error carries the router's own message —
e.g. invalid IR document: unknown field 'all', expected one of … — instead
of a bare status code, so an agent can correct its request rather than
guessing what was wrong.
Prompts¶
prompts/list offers ready-made investigation templates a client can surface
directly (e.g. as a slash command), separate from the tools above:
| Prompt | Arguments | Purpose |
|---|---|---|
investigate_trace |
trace_id (required) |
Seeds a get_trace call and a critical-path/error/self-time analysis. |
find_recent_errors |
service (required), minutes (optional, default 15) |
Seeds a search_traces/search_logs sweep for a service's errors. |
build_promql_query |
metric (required), intent (optional) |
Seeds discover_metrics → query_metrics for a metric. |
investigate_failing_dependency |
service (optional) |
Seeds get_service_map → search_traces on the worst edge → search_logs to trace a failure to its root cause. |
Each prompt renders into a single text message — pure argument substitution, no router call, so prompts work even before your credential has been validated for the session.
Two arguments offer live autocompletion via completion/complete, for
clients that ask for suggestions as you type: find_recent_errors's
service (backed by Tempo service.name tag-value discovery) and
build_promql_query's metric (backed by Prometheus __name__ label
discovery), both scoped to your tenant and filtered by the prefix you've
typed so far. Every other reference/argument returns no suggestions rather
than an error — completions are advisory, so a lookup failure never breaks
the request you're filling in.
Interactive views (MCP Apps)¶
Clients that support the MCP Apps extension render three tools as interactive views instead of raw JSON:
get_traceas a waterfall: span timings, self time versus time spent in child spans, per-span kind, attributes, and span events, with a ruler you can read span offsets against.get_profileas a flamegraph: per-frame self/total sample values across the call stack, colored by function name, with a tooltip on hover.get_service_mapas a graph view: services and external dependencies as nodes, calls between them as edges, with per-node and per-edge rate/error rate/p95.
Nothing needs configuring. A client that negotiates the extension gets all three; every other client keeps receiving the same JSON text result it always did.
How it works, if you are curious or writing a client:
- The client declares
io.modelcontextprotocol/uiin itsinitializecapabilities, namingtext/html;profile=mcp-appin that capability'smimeTypes. A client that declares the extension without naming the type cannot render either view, so it keeps the plain-text tool surface. - The server then marks
get_trace/get_profile/get_service_mapwith_meta.ui.resourceUripointing atui://signaldb/trace,ui://signaldb/profile, andui://signaldb/service-maprespectively, and attaches the trace/flamegraph/graph to the result asstructuredContentalongside the usual text block. - The client fetches the app's URI with
resources/read— served astext/html;profile=mcp-app— and renders it in a sandboxed iframe, handing it the tool result.
Each view is a single self-contained HTML document compiled into the binary.
It makes no network requests of its own and cannot reach the router: its only
data is the tool result the client hands it, which keeps it inside the
strictest sandbox hosts apply (default-src 'none').
Skill resources¶
Alongside the ui:// apps above, resources/list/resources/read also
serve skill:// documents: longer-form guidance a client fetches on demand,
kept out of the always-sent initialize instructions so those stay short.
Each skill follows the common skill://<name>/SKILL.md convention; a
skill://index.json resource lists all of them for discovery. The
list_skills/get_skill tools above mirror the same catalog for clients
that don't read MCP resources on their own. Currently one skill:
| Resource | Covers |
|---|---|
skill://query-ir/SKILL.md |
When query_ir covers more than search_traces/search_logs/query_metrics (a pipeline stage they can't express, or you're already building a document from discover_sources/discover_fields/discover_field_values), plus the full IR document reference — the same content as the Query IR reference, reused rather than duplicated. |
Both ui:// and skill:// resources are static and compiled into the
binary — identical for every client, so resources/list/resources/read
answer with a long cache TTL and public scope.
Running it¶
The server is off by default. Enable it in signaldb.toml:
[mcp]
enabled = true
bind_address = "127.0.0.1:8228" # serves MCP at /mcp; loopback by default
router_url = "http://localhost:3000" # the router HTTP API to forward to
router_timeout = 30 # seconds per forwarded request (default 30)
max_concurrent_tool_calls = 8 # tool calls in flight per session (default 8)
ui_base_url (also --ui-base-url / SIGNALDB__MCP__UI_BASE_URL) points at
the SignalDB UI. When set, search_traces, get_trace, and search_logs
results carry a _links.ui field with a deep link into the matching UI view
(trace search, a single trace, or log search), scoped to the call's tenant
and dataset. Left unset (the default), results carry no _links field at
all.
Each forwarded request is bounded by router_timeout (plus a fixed 5s connect
timeout), so a hung router fails the tool call cleanly instead of hanging the
agent indefinitely. Raise it if your agents run slow analytical queries.
That per-request bound alone does not bound one tools/call: the SDK's
shared retry policy can spend up to 4 attempts plus 30s of retry sleeps
underneath it, so with the default router_timeout a single call could
otherwise run for close to 150s. A total deadline of router_timeout +
30s (60s with defaults) wraps the whole call instead: a call still running
past it fails with a distinct tool call exceeded the Ns deadline error
(outcome=error, error.type=deadline — see
Audit and observability). Raising
router_timeout raises this deadline with it.
max_concurrent_tool_calls bounds how many tool calls one MCP session may
have in flight at once (also --max-concurrent-tool-calls /
SIGNALDB__MCP__MAX_CONCURRENT_TOOL_CALLS). A call that arrives while the
session is at the bound waits up to 2 seconds for an in-flight call to finish
and then fails with the distinct error too many concurrent tool calls (limit
N); wait for in-flight calls to finish (JSON-RPC -32600, data.limit = N)
— the other calls are unaffected, and nothing queues indefinitely. The bound
is per session, so one runaway agent cannot starve another session's tenant.
See Audit and observability for how such a call
is logged.
The server forwards live bearer credentials, so it binds loopback by
default. Exposing it off-host means changing bind_address to a routable
address and putting it behind TLS (direct HTTPS or a trusted terminator).
Then run the standalone binary (or use ./scripts/run-dev.sh services, which
starts it automatically on :8228):
cargo run --bin signaldb -- mcp
# stdio transport for local development (unauthenticated — dev only):
cargo run --bin signaldb -- mcp --stdio
The same settings are available as environment variables (multi-word fields
need the double-underscore form): SIGNALDB__MCP__ENABLED,
SIGNALDB__MCP__BIND_ADDRESS, SIGNALDB__MCP__ROUTER_URL,
SIGNALDB__MCP__ROUTER_TIMEOUT (seconds, also --router-timeout),
SIGNALDB__MCP__MAX_CONCURRENT_TOOL_CALLS, and SIGNALDB__MCP__ALLOWED_HOSTS
(see below). The sidecar also honours [self_monitoring] (from --config,
signaldb.toml, or SIGNALDB__SELF_MONITORING__*) so its own spans, audit
events, and metrics can be exported — see below. It parses the same shared
Configuration as every other service, so unrelated sections like
[wal].max_instances (see WAL Persistence)
are accepted but unused here — the standalone MCP server holds no WAL of its
own.
The Host allowlist (serving beyond localhost)¶
The Streamable HTTP transport carries a DNS-rebinding guard that validates the
inbound Host header, and by default accepts only loopback hosts
(localhost, 127.0.0.1, ::1). A client that reaches the server by any other
name or IP — a LAN address, a public hostname — is rejected with
403 Forbidden: Host header is not allowed before authentication runs. (Node's
fetch, which Claude Code uses for HTTP MCP, will not let a client override the
Host header, so there is no client-side workaround.)
When you serve the MCP off-localhost, name the reachable authority in the
allowlist. The value is a comma-separated list of host or host:port
authorities, appended to the loopback defaults:
# reached as mcp.example.org (behind TLS) and, for a bare LAN sidecar, by IP:port
signaldb mcp --allowed-hosts mcp.example.org,10.0.0.5:30228
# or via env
SIGNALDB__MCP__ALLOWED_HOSTS="mcp.example.org,10.0.0.5:30228" signaldb mcp
The single value * disables the guard entirely. The server still authenticates
every request (bearer + tenant), so * drops only the rebinding guard, never
authorization — but prefer an explicit list where you can.
Running as a sidecar¶
The MCP server ships as its own image, ghcr.io/cedricziel/signaldb/mcp
(the same signaldb binary as every other image, with signaldb mcp as its
entrypoint), so it runs as a sidecar next to a signaldb router/monolith. The
deployment (not the Dockerfile) makes it reachable: bind
a non-loopback address, point it at the router by service name, and publish
the port (EXPOSE alone does not publish anything).
services:
signaldb: # your router/monolith, serving the router on :3000
image: ghcr.io/cedricziel/signaldb:main
volumes: ["./data:/data"]
working_dir: /data
signaldb-mcp:
image: ghcr.io/cedricziel/signaldb/mcp:main # dedicated MCP image
# SDK-only + forward-only: no config file or catalog needed, just the
# router URL. It validates nothing itself — the router does.
environment:
# 0.0.0.0 so the published port is reachable (loopback is the default)
SIGNALDB__MCP__BIND_ADDRESS: "0.0.0.0:8228"
# the router, by compose service name
SIGNALDB__MCP__ROUTER_URL: "http://signaldb:3000"
# the authority clients reach this by — otherwise the Host guard 403s
# them before auth (see "The Host allowlist" above). Use your TLS
# hostname, or the bare host:port for a LAN sidecar.
SIGNALDB__MCP__ALLOWED_HOSTS: "mcp.example.org"
ports: ["8228:8228"] # publish it — required for reachability
depends_on: [signaldb]
restart: unless-stopped
Because it forwards live bearer credentials, a non-loopback bind should sit behind TLS — front it with your reverse proxy rather than publishing the raw port to an untrusted network.
Connecting an agent¶
The server speaks MCP over Streamable HTTP at /mcp. Authenticate with the
same headers as any SignalDB HTTP caller:
Authorization: Bearer <api-key>X-Tenant-ID: <tenant>X-Dataset-ID: <dataset>(optional)
Use the URL that matches your deployment: http://localhost:8228/mcp for a
loopback dev instance, or your HTTPS reverse-proxy URL (e.g.
https://mcp.example.org/mcp) for anything off-host — the server forwards live
bearer credentials, so a remote endpoint must be TLS-terminated.
Claude Code (CLI + IDE)¶
The first-class path — it passes arbitrary headers, which is how the server receives the bearer and tenant:
# local dev instance
claude mcp add --transport http signaldb http://localhost:8228/mcp \
--header "Authorization: Bearer sk-your-key" \
--header "X-Tenant-ID: your-tenant"
# deployed behind TLS
claude mcp add --transport http signaldb https://mcp.example.org/mcp \
--header "Authorization: Bearer sk-your-key" \
--header "X-Tenant-ID: your-tenant"
A request that carries no bearer token or no X-Tenant-ID is rejected with
401 at the MCP server before it reaches the transport. The MCP server does
not validate the credential itself — it forwards it, and the router decides
whether it is valid; an invalid or revoked key is rejected downstream and comes
back as a clean MCP tool error.
Claude.ai and ChatGPT (OAuth connector)¶
Claude.ai and OpenAI/ChatGPT register a remote MCP server through OAuth 2.1 with
Dynamic Client Registration — no headers, no pre-registration. Add the /mcp
URL under Settings → Connectors → Add custom connector; the client
discovers SignalDB's authorization server, registers itself, and sends you
through a sign-in + consent screen. The sign-in step is an ordinary browser
session, so when the operator has configured SSO you authenticate through your
identity provider there (or the connector reuses an existing SignalDB session);
password sign-in works unless the operator disabled it. On the consent screen —
which proceeds identically either way — you check every tenant you want
this connector to reach (a multi-select checklist of the tenants you belong
to) and approve the read scopes it requested; the token it receives is bound
to that whole set. One connector is enough for every tenant you need: there is
no more "add it a second time" workaround, and nothing stops you from checking
just one tenant if that's all you want.
For each tenant you check, the consent screen also offers a dataset choice:
all datasets in that tenant (the default — identical to every connector
granted before this choice existed) or only these datasets, which reveals
a checklist of that tenant's datasets and requires at least one checked box to
approve — set independently per tenant, so restricting one tenant's grant
never affects another's. Picking specific datasets binds that tenant's part of
the grant to exactly that set — queries against any other dataset in that
tenant are refused, and a query naming no dataset at all is rejected rather
than silently falling back to the tenant default when the set has more than
one dataset (a single-dataset restriction resolves to that dataset the same
way an unrestricted token resolves to the tenant default). A refresh preserves
whichever restriction each tenant's original grant had. Restricting a grant to
specific datasets is refused, naming the dataset_restriction_rollout_complete
config key, until an operator has set
[auth] dataset_restriction_rollout_complete = true on every router node —
see Multi-dataset rollout;
choosing "all datasets" is unaffected by this and always available.
The endpoint must be HTTPS with a valid certificate — these clients will not connect to a raw LAN port, so the TLS reverse proxy is required here.
Operator setup. The authorization server is served by the router, off by default. Enable it and point it at the externally-reachable URLs clients use:
# router config (signaldb.toml)
[mcp.oauth]
enabled = true
issuer_url = "https://signaldb.example.org" # this AS, as clients reach it
resource_url = "https://signaldb.example.org/mcp" # the MCP resource tokens bind to
# access_token_ttl = "1h"; refresh_token_ttl = "30d"; authorization_code_ttl = "60s"
The signaldb mcp sidecar advertises the same resource so an
unauthenticated request is challenged toward discovery — pass the matching URLs:
signaldb mcp \
--oauth-resource-url https://signaldb.example.org/mcp \
--oauth-issuer-url https://signaldb.example.org
Tokens are opaque, catalog-backed, and audience-bound to resource_url;
revoking one is a row delete. POST /oauth/introspect (RFC 7662) reports
whether a bearer token is active and, if so, its full granted-tenant set,
scopes, audience, and expiry — this is how the signaldb mcp sidecar learns
a multi-tenant connector's whole reachable set before any one tenant has been
selected for a call; it isn't something you call directly as an operator or
agent. The read scopes a token may hold —
traces:read, logs:read, metrics:read, profiles:read, schema:read,
processors:read, evals:read —
gate the corresponding query surface (see the
multi-tenancy model); a request with no scope
is granted all of them, and schema:write/processors:write/evals:write are never
grantable through OAuth (a request naming only one of them is rejected with
invalid_scope). The
existing Bearer <api-key> + X-Tenant-ID path is unchanged; OAuth is an
added credential type, not a replacement.
Audit and observability¶
Every tool call is audited: after it completes, the server emits exactly one
structured log event (target signaldb_mcp::audit) with bounded fields —
never the arguments, the query expression, or the result:
| Field | Meaning |
|---|---|
tool |
The tool name (search_traces, get_trace, …). |
tenant_id |
The tenant the router resolved the caller to. |
dataset |
The dataset the call named (its dataset argument, else X-Dataset-ID); absent when neither is set. |
session_id |
The Mcp-Session-Id (stdio on the stdio transport). |
outcome |
ok, truncated (result cut at the size cap), denied, throttled, or error. |
duration_ms |
Wall time of the call, including any wait for a concurrency permit. |
error.type |
Only for outcome=error: concurrency_limit, deadline, tool_error, the router's HTTP status (500), or the JSON-RPC code. |
Levels: ok, truncated, and throttled log at info; denied (the router
rejected the credential or the tenant/dataset access — a 401/403) at
warn, so probing is visible; error at error. A call refused at the
concurrency bound is outcome=error, error.type=concurrency_limit; a call
still running past the total per-call deadline (see Running it)
is outcome=error, error.type=deadline.
The same call is one tools/call {tool} span (INTERNAL, gen_ai.tool.name,
mcp.session.id, signaldb.tenant.id, signaldb.dataset.id; status Error
only for outcome=error), the HTTP request that carried it is a POST /mcp
server span parented to the client's traceparent, and two metrics count the
calls: signaldb.mcp.tool_calls by tool and outcome and
signaldb.mcp.tool_call.duration by tool (Prometheus:
signaldb_mcp_tool_calls_total{gen_ai_tool_name,signaldb_mcp_outcome} and
signaldb_mcp_tool_call_duration_seconds{gen_ai_tool_name}). All of it is
exported when [self_monitoring] is enabled for the sidecar (service name
signaldb-mcp); see docs/operations/self-monitoring-traces.md.
Example flow¶
Configuring a new application to send data here? Call connection_info first —
it returns the deployment's public OTLP endpoints, the headers to send, and
ready-to-paste OTEL_EXPORTER_OTLP_* env vars — then mint an ingest key with
tenant_create_api_key and substitute it for the placeholder.
server_info— confirm you are connected as the expected tenant.discover_attributes— list tag names, then values forservice.name.search_traceswith{ .service.name = "checkout" && status = error }and a time range to find failing requests.get_tracewith an ID from the search results to inspect the full trace.
To explore logs or metrics instead: discover_attributes with signal:
"logs" lists Loki labels (add tag for a label's values); signal:
"metrics" does the same for Prometheus labels. discover_metrics lists
metric names directly, for building a query_metrics PromQL expression.
Before filtering or grouping by a name you are unsure of, ask the schema
registry what it means: resolve_attribute with key: "k8s.pod.uid" (or
resolve_entity / resolve_metric, or search_schema with kind:
"attribute", prefix: "k8s.pod.") returns namespace-tagged, precedence-ordered
definitions — a tenant's own conventions (uploaded with
create_schema_registry) come first, the bundled OpenTelemetry definition is
kept as an alternative. discover_* tells you which names have data;
resolve_* tells you what they mean. resolve_entity with name:
"gen_ai.agent" tells an agent which attribute identifies an AI agent
(gen_ai.agent.id).
From the CLI¶
The same discovery is available outside an agent session, via
signaldb-sdk like every other CLI capability:
# Native surface (Query IR, logical dotted names, no scan)
signaldb-cli discover sources
signaldb-cli discover fields --source logs
signaldb-cli discover values --source traces --field span.kind
signaldb-cli discover values --source traces --field http.route --sample
# Compatibility-dialect view (Tempo tags, Loki/Prometheus labels)
signaldb-cli discover attributes --signal traces --tag service.name
signaldb-cli discover attributes --signal logs
signaldb-cli discover attributes --signal metrics --tag job
signaldb-cli discover metrics
discover fields/values/sources are the native surface: they speak the same
logical names as a Query IR document and are answered from metadata rather than
by scanning. discover values reads data only when you pass --sample, and the
response says so — without it you are told what would answer the question
instead. discover attributes remains the dialect-shaped view, for parity with
what Grafana sees. See the Query IR reference.
Schema-registry lookup and custom-registry management mirror the schema tools
(reads need a key with schema:read, mutations schema:write):
signaldb-cli schema registry list
signaldb-cli schema registry get otel 1.43.0
signaldb-cli schema attribute get k8s.pod.uid
signaldb-cli schema entity get k8s.pod
signaldb-cli schema metric search k8s.pod. --limit 20
signaldb-cli admin schema validate --file conventions.yaml
signaldb-cli admin schema create --file conventions.yaml # YAML or JSON
signaldb-cli admin schema replace acme 1.0.0 --file conventions.yaml
signaldb-cli admin schema delete acme 1.0.0
server_info mirrors signaldb-cli whoami; query_metrics/search_logs's
range mode mirrors signaldb query --promql|--logql ... --start ... --end
...; get_trace mirrors signaldb query --trace-id <id>:
signaldb-cli whoami
signaldb-cli query --promql 'up' --start 0 --end 3600 --step 15s
signaldb-cli query --trace-id 4bf92f3577b34da6a3ce929d0e0e4736
get_source_context mirrors signaldb-cli tenant source-context — a
read tool, like get_trace/get_profile, so any valid key of the tenant
works; it does not need tenant:manage despite living under the CLI's
tenant verb group:
signaldb-cli tenant source-context --path src/main.rs --line 42 --api-key sk-your-key --tenant-id your-tenant
The tenant_* tools mirror signaldb-cli tenant: tenant_info is tenant
show, the table tools are tenant table ... (any valid key of the tenant),
and the management tools are
tenant dataset|api-key|membership|schema|github ... (a key carrying
tenant:manage; destructive verbs prompt on a TTY unless --yes):
signaldb-cli tenant show --api-key sk-your-key --tenant-id your-tenant
signaldb-cli tenant table list --api-key sk-your-key --tenant-id your-tenant
signaldb-cli tenant table provision --api-key sk-your-key --tenant-id your-tenant
signaldb-cli tenant table schemas --api-key sk-your-key --tenant-id your-tenant
signaldb-cli tenant table available-schemas --api-key sk-your-key
signaldb-cli tenant dataset create staging --api-key sk-manage-key --tenant-id your-tenant
signaldb-cli tenant api-key create --name ci --scope traces:write --api-key sk-manage-key --tenant-id your-tenant
signaldb-cli tenant membership set alice@example.com --role member --api-key sk-manage-key --tenant-id your-tenant
Platform administration (list_tenants, create_dataset, revoke_api_key,
...) mirrors the admin command group, authenticated with the administrative
key instead of a tenant key — see signaldb-cli admin --help.