Skip to content

MCP server (AI agent access)

SignalDB ships a Model Context Protocol server, signaldb mcp (a subcommand of the signaldb binary), that lets AI agents (Claude Code, Claude.ai, IDE assistants) query your traces directly. It is a thin, credential-forwarding client: it validates the bearer token you present and forwards that same token to the router's HTTP API. It holds no credential of its own, so a request can only ever see what your key is already allowed to see — tenant isolation stays enforced by the router.

What it exposes

Neither family below is hidden from tools/list — a call your credential does not authorize comes back as a clean access-denied tool error, not a missing tool.

Query and discovery

Available to every authenticated tenant session — there is no role gating on these:

Tool Purpose
server_info Confirm connectivity and which tenant your credential resolves to.
connection_info This deployment's public OTLP gRPC/HTTP and Prometheus remote-write endpoints, the query API base, headers with your tenant/dataset filled in, the API-key scopes ingest and query need, and ready-to-paste OTEL_EXPORTER_OTLP_* env vars. Takes optional tenant and dataset; a multi-tenant credential must pass tenant.
discover_datasets The tenant and datasets your credential can access, as a nested Markdown list, marking the session's current default dataset. Filtered to your credential's dataset restriction, if any — a dataset outside it never appears, even by name. Call before passing an explicit dataset or tenant argument elsewhere.
search_traces TraceQL search over your tenant's traces.
get_trace Fetch a single trace by ID, with each span's kind (renders as a waterfall — see below). Searches the last 30 days unless start/end are given.
search_trace_groups The grouped RED-metrics view (count, error count, p50/p95 duration, last-seen per group) the UI's traces tab "group by" table shows, along group_by dimensions. No free-text query filter yet — use query_ir for scoped filtering or a custom aggregate.
get_service_map The service dependency graph — nodes per service/external dependency, edges for the calls between them — optionally scoped to one service's neighbourhood (service/depth). Wraps the Query IR graph envelope; returns the graph plus a summary naming the busiest and highest-error edges, and a web UI link (renders as an interactive map — see below).
get_profile Fetch a single profile's flamegraph by ID (renders as an interactive flamegraph — see below).
get_source_context Fetch a source-code snippet around a stack-frame location (path/line) through the tenant's linked GitHub App installation(s). repository may be omitted to probe by path alone; ref may be omitted for the default branch. status: "unavailable" with a reason is a normal answer, not an error.
discover_attributes List queryable attribute/label names, or the values for one. Signal-aware: traces (default, Tempo tags), logs (Loki labels), metrics (Prometheus labels), profiles (Pyroscope labels).
discover_metrics List the distinct metric names visible to your tenant.
discover_fields The queryable fields of a signal source, as logical dotted OTel names with type, origin, coverage and approximate cardinality. Answered from the schema registry and maintained statistics — reads no signal data.
discover_field_values Value suggestions for one field. Exact and free for a declared value set; otherwise it names the query that would answer it, and only sample: true runs that query.
discover_sources The signal sources available to your tenant, with whether each is queryable.
discover_profile_types List the Pyroscope profile types with data for your tenant (e.g. CPU, heap).
query_metrics PromQL query over your tenant's metrics (native Prometheus result); instant by default, or a range query when start/end (and optionally step) are given.
search_logs LogQL query over your tenant's logs (native Loki result); instant or range, same as query_metrics.
search_profiles Search profiles with a Pyroscope selector and a time range; returns the aggregated flame graph (flamebearer encoding).
compare_profiles Compare profiles between two time ranges with a shared Pyroscope selector; returns the differential flame graph.
profiles_for_trace List the profiles correlated with a trace id.
query_ir Native Query IR document (the structured, versioned query surface).
list_skills List the longer-form guidance documents ("skills") this server exposes beyond the tool descriptions — same catalog as the skill://index.json resource.
get_skill Read one guidance document by name (e.g. query-ir) — same content as the corresponding skill://<name>/SKILL.md resource, for clients that don't read MCP resources.
list_schema_registries List the schema registries visible to your tenant in precedence order (custom first, then the bundled signaldb and otel semconv), with definition counts.
get_schema_registry Fetch one registry's summary and full document by namespace/version.
resolve_attribute What an attribute key means: every definition across the visible registries, precedence-ordered (primary first), with brief, type, examples, deprecation.
resolve_entity What an entity type (k8s.pod, service, ...) means: identifying/descriptive attributes, what it extends, associated metrics.
resolve_metric What a metric means: instrument, unit, brief, recorded attributes, associated entities.
search_schema Prefix search over attributes, entities, or metrics (kind, prefix, limit) to find the right vocabulary before querying, or keys to batch-resolve an exact attribute/metric name set instead.
create_schema_registry Upload a custom Weaver-model registry document (JSON object) for your tenant (requires schema:write).
replace_schema_registry Replace a custom registry's document by namespace/version (requires schema:write; bundled registries refuse).
validate_schema_registry Validate a registry document without storing it; errors carry document paths (requires schema:write).
delete_schema_registry Delete a custom registry by namespace/version (requires schema:write; bundled registries refuse).
list_processors List the tenant's OTTL telemetry processors (requires processors:read). See Processors.
get_processor Fetch one processor by name (requires processors:read).
validate_processor Compile-check a processor specification without storing it; errors carry statement index, column, and message (requires processors:read).
test_processor Dry-run a processor set against an OTLP JSON payload — transformed payload plus per-statement match/error counts; never writes to the WAL or forwards data (requires processors:read).
create_processor Create a processor (requires processors:write; tenant-admin only, not OAuth-grantable).
replace_processor Replace a processor's full document by name (requires processors:write; tenant-admin only, not OAuth-grantable).
delete_processor Delete a processor by name (requires processors:write; tenant-admin only, not OAuth-grantable).
list_eval_sets List the agent eval sets in a dataset, without their cases (requires evals:read). See Eval sets.
get_eval_set Fetch one eval set by name with a page of its cases in order: offset (default 0) and limit (default 200, at most 1000); the reply adds total_cases, offset, returned and has_more (requires evals:read).
create_eval_set Create an eval set: name, agent, description, cases (requires evals:write; not OAuth-grantable).
replace_eval_set Replace an eval set's agent, description and every case by name; never creates (requires evals:write; not OAuth-grantable).
delete_eval_set Delete an eval set and its cases by name (requires evals:write; not OAuth-grantable).
append_eval_cases Append cases to an eval set; ids it already holds are skipped and reported (requires evals:write; not OAuth-grantable).
append_eval_cases_from_traces Append one case per matching agent trace the set doesn't hold yet, newest first, up to sample (default 50); failing_evaluator keeps traces with a failing result of that evaluator. Reports matches, already_present, added (requires evals:write and traces:read; not OAuth-grantable).
upload_eval_results Upload a JSONL/CSV results file (content, format) as one offline run of agent/version/set (run_id optional; generated and named in errors, so a retry with it is safe). Any invalid row rejects the file and every problem is listed; returns the run id and per-evaluator summary (requires evals:write; not OAuth-grantable).
list_eval_runs Offline eval runs in a window (default now-7d..now), newest first, filtered by agent, version, set: run id, eval set, agent, version, started_at/last_result_at, status (running/complete/partial with reasons), results, cases, errors, unlinked results, pass rate, per-evaluator mean and pass rate, and previous_run_id (the natural baseline). At most limit runs (default 50, max 500); total_runs and truncated say when more matched (requires logs:read).
compare_eval_runs Compare a candidate run with a baseline run (each a run id or latest:<version>; agent/set default to the other side's run) over a window (default now-30d..now): both runs' summaries, per-evaluator means, pass rates, deltas and cases that moved, counts of regressions/improvements/unchanged, and the regressed cases (largest drop first, at most limit, default 50, max 500) with the evaluators that got worse and both trace ids. include_tools: true adds each listed regression's tool-call diff (requires logs:read, plus traces:read with include_tools).

Every processor tool takes the tenant parameter validated the same way as the query tools above. Changes apply at ingest within [processors].reload_interval (default 30s) — see Processors.

The eval set tools, upload_eval_results, list_eval_runs and compare_eval_runs take the same tenant parameter and, like the query tools, a required dataset: a set or a run belongs to one dataset, and one MCP session may span several, so there is no implicit default. See Upload a results file and Runs and comparisons outside the UI.

list_eval_runs and compare_eval_runs answer "did version B of my agent get worse than A, where and why" with the same figures as the Runs and Compare pages: they are Query IR reads over the gen_ai.evaluation.result log records (and, for tool diffs, the execute_tool spans), so they need the read scopes of those signals rather than evals:read. A typical session lists the runs, compares the newest with its previous_run_id (or latest:<version>), then opens a regressed case's traces with get_trace.

Each query tool requires a dataset argument, targeting the dataset your tenant may access (the router validates access and rejects the rest). Large results are capped and returned with a truncated: true flag telling the agent to narrow the query.

Most of these tools also require a tenant argument. For a credential authorized for exactly one tenant (an API key, or an OAuth connector granted a single tenant — see below), it's a confirmation check, not a way to switch tenants: it must equal the tenant this specific call's credential resolves to (server_info, discover_datasets), and a mismatch fails the call with an error naming both tenants, before any request reaches the router. For an OAuth connector granted more than one tenant, tenant is a real selector: pass the tenant this call should run against, and the server validates it against the connector's own granted set before forwarding the call — naming a tenant outside that set fails the same way a mismatch does for a single-tenant credential. Both arguments are required rather than optional because one MCP session (one mcp-session-id) can hold credentials for several tenants and datasets across its calls — there is no single implicit session-wide default left to fall back to. To reach a second tenant within one session using separate API-key credentials, present a different credential (Authorization bearer token plus X-Tenant-ID) on a later call rather than opening a second connection; the router independently authenticates each call, up to a bounded number of distinct identities per session. A multi-tenant OAuth connector reaches its other tenants by simply naming them in tenant on later calls — no second credential needed, since they all belong to the same grant.

discover_datasets/server_info reflect this too: for a single-tenant credential they report that one tenant, as before; for a multi-tenant OAuth connector they enumerate every granted tenant (and, for discover_datasets, each one's own datasets) so an agent can see the whole reachable set before picking a tenant for a later call.

Operational control

Admin-authenticated (the administrative API key, not a tenant key):

Tool Purpose
compact_run Trigger a compaction pass now.
compact_status Active compaction leases and metrics.
compact_dry_run Plan compaction candidates without executing.

Platform administration

Unprefixed, admin-authenticated (the administrative API key can manage any tenant — this is the same credential the admin CLI group and the admin HTTP API use):

Tool Purpose
list_tenants / get_tenant List every tenant, or fetch one by ID.
create_tenant / update_tenant Create a tenant, or update its name/default dataset.
delete_tenant Delete a tenant and everything under it. Destructive: requires confirm equal to tenant_id.
create_user Create a human user and grant an initial tenant membership.
list_datasets / create_dataset List or create a tenant's datasets.
delete_dataset Delete a dataset by ID. Destructive: requires confirm equal to dataset_id.
list_api_keys List a tenant's API keys with their scopes and dataset set (or that they're unrestricted). Raw secrets are never returned.
create_api_key Create an API key carrying explicit scopes (required; e.g. traces:write, schema:read), an optional dataset_ids set, and an optional allowed_origins set restricting browser (CORS) ingest — see Origin restriction. The raw secret is returned exactly once, in this response.
update_api_key_scopes Change a live key's scopes, dataset set, and/or origin set without rotating its secret; dataset_ids/allowed_origins replace their respective restrictions, clear_dataset_restriction/clear_allowed_origins: true (with no corresponding set field) remove them, and omitting a pair leaves it unchanged — sending a set and its own clear flag together is rejected before any request is made. Revoked keys are rejected.
revoke_api_key Revoke an API key by ID. Destructive: requires confirm equal to key_id.

Tenant self-management

tenant_-prefixed; act as the caller's own identity within its own tenant. Two sub-groups, by which endpoint they wrap:

Tenant view, tables, and schemas (tenant self-service API) — work with a plain tenant API key, exactly like the query tools above:

Tool Purpose
tenant_info The caller's own tenant: id, enabled flag, schema configuration (mirrors signaldb-cli tenant show).
tenant_list_tables List the tenant's provisioned signal tables. Filtered to your credential's dataset restriction, if any, the same way discover_datasets is.
tenant_create_tables Provision (create) the tenant's enabled signal tables — the manual trigger from table provisioning.
tenant_list_table_schemas List the tenant's configured table schema types (distinct from tenant_list_tables, which lists what is actually provisioned).
list_available_table_schemas List every table schema type SignalDB knows how to provision, regardless of tenant configuration.

Datasets, API keys, memberships, and schema (management API) — need a management credential: either a human session (a browser session cookie, or an OAuth access token — see Claude.ai and ChatGPT below) holding the tenant-admin role or the instance-admin flag, or an API key that carries the tenant:manage scope for that tenant. Ingest-only keys and legacy unscoped keys are denied — a deliberate privilege boundary: an ingest key can write signal data and provision tables, but minting or revoking other API keys, deleting datasets, or changing memberships is opt-in, granted only to a tenant admin or a key explicitly scoped for it. Calling one of these with a key that lacks the scope returns a clean access-denied error naming the required scope rather than succeeding or 404ing. tenant:manage is never granted through OAuth consent; see API-key scopes.

Tool Purpose
tenant_list_datasets / tenant_create_dataset List or create the caller's own tenant's datasets.
tenant_delete_dataset Delete a dataset by name. Destructive: requires confirm equal to dataset_name.
tenant_list_api_keys List the caller's own tenant's API keys. Raw secrets are never returned.
tenant_create_api_key Create an API key for the caller's own tenant. The raw secret is returned exactly once.
tenant_update_api_key Update the scopes, dataset set, and/or origin set of one of the caller's own tenant's API keys — same dataset_ids/allowed_origins/clear-flag semantics as update_api_key_scopes above.
tenant_revoke_api_key Revoke one of the caller's own tenant's API keys. Destructive: requires confirm equal to key_id.
tenant_list_memberships / tenant_upsert_membership List the caller's own tenant's memberships, or create/update a member's role.
tenant_remove_membership Remove a member from the caller's own tenant. Destructive: requires confirm equal to user_id.
tenant_get_schema The registered logical (client-visible) and physical (storage) schema for every signal source. Takes tenant_id.
tenant_start_github_link Start linking a GitHub App installation. Returns an install_url to open in a browser signed in to SignalDB as an admin of this tenant, and its expires_at.
tenant_attach_github_installation Attach a GitHub App installation that already exists (e.g. linked to another tenant on the same GitHub account) directly, without the OAuth install flow — GitHub skips the consent screen and any redirect back once an account's App installation already exists, so tenant_start_github_link cannot complete a second tenant's link there. Requires instance-admin, unlike every other tool in this section: a tenant:manage grant alone is not enough, since this path has no GitHub code to verify the caller actually controls the installation being linked.
tenant_list_github_installations List the caller's own tenant's linked GitHub App installations, with their repositories.
tenant_remove_github_installation Remove a linked GitHub App installation. Destructive: requires confirm equal to installation_id.

Destructive tools carry the MCP destructiveHint annotation; read-only tools carry readOnlyHint — a client that inspects tools/list annotations can tell which is which without trying the call.

When the router throttles a tool's downstream request (429, the tenant's query rate limit), the server first lets the SDK's shared retry policy absorb it — waiting the server-stated Retry-After within the policy's bounds (see client retry) — so a brief burst is invisible to the agent. Only once retries are exhausted (or the server asks for more than the policy's 10 s per-attempt ceiling) does the tool fail, with a throttled error distinct from an internal failure: the message starts with throttled: and names the wait (throttled: search_logs was rate limited; the server asked to retry in 30s), and the error data carries retryAfterMs (milliseconds; null when no wait was stated) plus http_status: 429. An agent should wait that long or narrow the query.

When the router rejects a query-backed tool's request outright (query_ir, search_trace_groups, get_service_map, get_profile, discover_fields, discover_field_values), the tool error carries the router's own message — e.g. invalid IR document: unknown field 'all', expected one of … — instead of a bare status code, so an agent can correct its request rather than guessing what was wrong.

Prompts

prompts/list offers ready-made investigation templates a client can surface directly (e.g. as a slash command), separate from the tools above:

Prompt Arguments Purpose
investigate_trace trace_id (required) Seeds a get_trace call and a critical-path/error/self-time analysis.
find_recent_errors service (required), minutes (optional, default 15) Seeds a search_traces/search_logs sweep for a service's errors.
build_promql_query metric (required), intent (optional) Seeds discover_metrics → query_metrics for a metric.
investigate_failing_dependency service (optional) Seeds get_service_map → search_traces on the worst edge → search_logs to trace a failure to its root cause.

Each prompt renders into a single text message — pure argument substitution, no router call, so prompts work even before your credential has been validated for the session.

Two arguments offer live autocompletion via completion/complete, for clients that ask for suggestions as you type: find_recent_errors's service (backed by Tempo service.name tag-value discovery) and build_promql_query's metric (backed by Prometheus __name__ label discovery), both scoped to your tenant and filtered by the prefix you've typed so far. Every other reference/argument returns no suggestions rather than an error — completions are advisory, so a lookup failure never breaks the request you're filling in.

Interactive views (MCP Apps)

Clients that support the MCP Apps extension render three tools as interactive views instead of raw JSON:

  • get_trace as a waterfall: span timings, self time versus time spent in child spans, per-span kind, attributes, and span events, with a ruler you can read span offsets against.
  • get_profile as a flamegraph: per-frame self/total sample values across the call stack, colored by function name, with a tooltip on hover.
  • get_service_map as a graph view: services and external dependencies as nodes, calls between them as edges, with per-node and per-edge rate/error rate/p95.

Nothing needs configuring. A client that negotiates the extension gets all three; every other client keeps receiving the same JSON text result it always did.

How it works, if you are curious or writing a client:

  • The client declares io.modelcontextprotocol/ui in its initialize capabilities, naming text/html;profile=mcp-app in that capability's mimeTypes. A client that declares the extension without naming the type cannot render either view, so it keeps the plain-text tool surface.
  • The server then marks get_trace/get_profile/get_service_map with _meta.ui.resourceUri pointing at ui://signaldb/trace, ui://signaldb/profile, and ui://signaldb/service-map respectively, and attaches the trace/flamegraph/graph to the result as structuredContent alongside the usual text block.
  • The client fetches the app's URI with resources/read — served as text/html;profile=mcp-app — and renders it in a sandboxed iframe, handing it the tool result.

Each view is a single self-contained HTML document compiled into the binary. It makes no network requests of its own and cannot reach the router: its only data is the tool result the client hands it, which keeps it inside the strictest sandbox hosts apply (default-src 'none').

Skill resources

Alongside the ui:// apps above, resources/list/resources/read also serve skill:// documents: longer-form guidance a client fetches on demand, kept out of the always-sent initialize instructions so those stay short. Each skill follows the common skill://<name>/SKILL.md convention; a skill://index.json resource lists all of them for discovery. The list_skills/get_skill tools above mirror the same catalog for clients that don't read MCP resources on their own. Currently one skill:

Resource Covers
skill://query-ir/SKILL.md When query_ir covers more than search_traces/search_logs/query_metrics (a pipeline stage they can't express, or you're already building a document from discover_sources/discover_fields/discover_field_values), plus the full IR document reference — the same content as the Query IR reference, reused rather than duplicated.

Both ui:// and skill:// resources are static and compiled into the binary — identical for every client, so resources/list/resources/read answer with a long cache TTL and public scope.

Running it

The server is off by default. Enable it in signaldb.toml:

[mcp]
enabled = true
bind_address = "127.0.0.1:8228"      # serves MCP at /mcp; loopback by default
router_url = "http://localhost:3000" # the router HTTP API to forward to
router_timeout = 30                  # seconds per forwarded request (default 30)
max_concurrent_tool_calls = 8        # tool calls in flight per session (default 8)

ui_base_url (also --ui-base-url / SIGNALDB__MCP__UI_BASE_URL) points at the SignalDB UI. When set, search_traces, get_trace, and search_logs results carry a _links.ui field with a deep link into the matching UI view (trace search, a single trace, or log search), scoped to the call's tenant and dataset. Left unset (the default), results carry no _links field at all.

Each forwarded request is bounded by router_timeout (plus a fixed 5s connect timeout), so a hung router fails the tool call cleanly instead of hanging the agent indefinitely. Raise it if your agents run slow analytical queries.

That per-request bound alone does not bound one tools/call: the SDK's shared retry policy can spend up to 4 attempts plus 30s of retry sleeps underneath it, so with the default router_timeout a single call could otherwise run for close to 150s. A total deadline of router_timeout + 30s (60s with defaults) wraps the whole call instead: a call still running past it fails with a distinct tool call exceeded the Ns deadline error (outcome=error, error.type=deadline — see Audit and observability). Raising router_timeout raises this deadline with it.

max_concurrent_tool_calls bounds how many tool calls one MCP session may have in flight at once (also --max-concurrent-tool-calls / SIGNALDB__MCP__MAX_CONCURRENT_TOOL_CALLS). A call that arrives while the session is at the bound waits up to 2 seconds for an in-flight call to finish and then fails with the distinct error too many concurrent tool calls (limit N); wait for in-flight calls to finish (JSON-RPC -32600, data.limit = N) — the other calls are unaffected, and nothing queues indefinitely. The bound is per session, so one runaway agent cannot starve another session's tenant. See Audit and observability for how such a call is logged.

The server forwards live bearer credentials, so it binds loopback by default. Exposing it off-host means changing bind_address to a routable address and putting it behind TLS (direct HTTPS or a trusted terminator).

Then run the standalone binary (or use ./scripts/run-dev.sh services, which starts it automatically on :8228):

cargo run --bin signaldb -- mcp
# stdio transport for local development (unauthenticated — dev only):
cargo run --bin signaldb -- mcp --stdio

The same settings are available as environment variables (multi-word fields need the double-underscore form): SIGNALDB__MCP__ENABLED, SIGNALDB__MCP__BIND_ADDRESS, SIGNALDB__MCP__ROUTER_URL, SIGNALDB__MCP__ROUTER_TIMEOUT (seconds, also --router-timeout), SIGNALDB__MCP__MAX_CONCURRENT_TOOL_CALLS, and SIGNALDB__MCP__ALLOWED_HOSTS (see below). The sidecar also honours [self_monitoring] (from --config, signaldb.toml, or SIGNALDB__SELF_MONITORING__*) so its own spans, audit events, and metrics can be exported — see below. It parses the same shared Configuration as every other service, so unrelated sections like [wal].max_instances (see WAL Persistence) are accepted but unused here — the standalone MCP server holds no WAL of its own.

The Host allowlist (serving beyond localhost)

The Streamable HTTP transport carries a DNS-rebinding guard that validates the inbound Host header, and by default accepts only loopback hosts (localhost, 127.0.0.1, ::1). A client that reaches the server by any other name or IP — a LAN address, a public hostname — is rejected with 403 Forbidden: Host header is not allowed before authentication runs. (Node's fetch, which Claude Code uses for HTTP MCP, will not let a client override the Host header, so there is no client-side workaround.)

When you serve the MCP off-localhost, name the reachable authority in the allowlist. The value is a comma-separated list of host or host:port authorities, appended to the loopback defaults:

# reached as mcp.example.org (behind TLS) and, for a bare LAN sidecar, by IP:port
signaldb mcp --allowed-hosts mcp.example.org,10.0.0.5:30228
# or via env
SIGNALDB__MCP__ALLOWED_HOSTS="mcp.example.org,10.0.0.5:30228" signaldb mcp

The single value * disables the guard entirely. The server still authenticates every request (bearer + tenant), so * drops only the rebinding guard, never authorization — but prefer an explicit list where you can.

Running as a sidecar

The MCP server ships as its own image, ghcr.io/cedricziel/signaldb/mcp (the same signaldb binary as every other image, with signaldb mcp as its entrypoint), so it runs as a sidecar next to a signaldb router/monolith. The deployment (not the Dockerfile) makes it reachable: bind a non-loopback address, point it at the router by service name, and publish the port (EXPOSE alone does not publish anything).

services:
  signaldb: # your router/monolith, serving the router on :3000
    image: ghcr.io/cedricziel/signaldb:main
    volumes: ["./data:/data"]
    working_dir: /data

  signaldb-mcp:
    image: ghcr.io/cedricziel/signaldb/mcp:main # dedicated MCP image
    # SDK-only + forward-only: no config file or catalog needed, just the
    # router URL. It validates nothing itself — the router does.
    environment:
      # 0.0.0.0 so the published port is reachable (loopback is the default)
      SIGNALDB__MCP__BIND_ADDRESS: "0.0.0.0:8228"
      # the router, by compose service name
      SIGNALDB__MCP__ROUTER_URL: "http://signaldb:3000"
      # the authority clients reach this by — otherwise the Host guard 403s
      # them before auth (see "The Host allowlist" above). Use your TLS
      # hostname, or the bare host:port for a LAN sidecar.
      SIGNALDB__MCP__ALLOWED_HOSTS: "mcp.example.org"
    ports: ["8228:8228"] # publish it — required for reachability
    depends_on: [signaldb]
    restart: unless-stopped

Because it forwards live bearer credentials, a non-loopback bind should sit behind TLS — front it with your reverse proxy rather than publishing the raw port to an untrusted network.

Connecting an agent

The server speaks MCP over Streamable HTTP at /mcp. Authenticate with the same headers as any SignalDB HTTP caller:

  • Authorization: Bearer <api-key>
  • X-Tenant-ID: <tenant>
  • X-Dataset-ID: <dataset> (optional)

Use the URL that matches your deployment: http://localhost:8228/mcp for a loopback dev instance, or your HTTPS reverse-proxy URL (e.g. https://mcp.example.org/mcp) for anything off-host — the server forwards live bearer credentials, so a remote endpoint must be TLS-terminated.

Claude Code (CLI + IDE)

The first-class path — it passes arbitrary headers, which is how the server receives the bearer and tenant:

# local dev instance
claude mcp add --transport http signaldb http://localhost:8228/mcp \
  --header "Authorization: Bearer sk-your-key" \
  --header "X-Tenant-ID: your-tenant"

# deployed behind TLS
claude mcp add --transport http signaldb https://mcp.example.org/mcp \
  --header "Authorization: Bearer sk-your-key" \
  --header "X-Tenant-ID: your-tenant"

A request that carries no bearer token or no X-Tenant-ID is rejected with 401 at the MCP server before it reaches the transport. The MCP server does not validate the credential itself — it forwards it, and the router decides whether it is valid; an invalid or revoked key is rejected downstream and comes back as a clean MCP tool error.

Claude.ai and ChatGPT (OAuth connector)

Claude.ai and OpenAI/ChatGPT register a remote MCP server through OAuth 2.1 with Dynamic Client Registration — no headers, no pre-registration. Add the /mcp URL under Settings → Connectors → Add custom connector; the client discovers SignalDB's authorization server, registers itself, and sends you through a sign-in + consent screen. The sign-in step is an ordinary browser session, so when the operator has configured SSO you authenticate through your identity provider there (or the connector reuses an existing SignalDB session); password sign-in works unless the operator disabled it. On the consent screen — which proceeds identically either way — you check every tenant you want this connector to reach (a multi-select checklist of the tenants you belong to) and approve the read scopes it requested; the token it receives is bound to that whole set. One connector is enough for every tenant you need: there is no more "add it a second time" workaround, and nothing stops you from checking just one tenant if that's all you want.

For each tenant you check, the consent screen also offers a dataset choice: all datasets in that tenant (the default — identical to every connector granted before this choice existed) or only these datasets, which reveals a checklist of that tenant's datasets and requires at least one checked box to approve — set independently per tenant, so restricting one tenant's grant never affects another's. Picking specific datasets binds that tenant's part of the grant to exactly that set — queries against any other dataset in that tenant are refused, and a query naming no dataset at all is rejected rather than silently falling back to the tenant default when the set has more than one dataset (a single-dataset restriction resolves to that dataset the same way an unrestricted token resolves to the tenant default). A refresh preserves whichever restriction each tenant's original grant had. Restricting a grant to specific datasets is refused, naming the dataset_restriction_rollout_complete config key, until an operator has set [auth] dataset_restriction_rollout_complete = true on every router node — see Multi-dataset rollout; choosing "all datasets" is unaffected by this and always available.

The endpoint must be HTTPS with a valid certificate — these clients will not connect to a raw LAN port, so the TLS reverse proxy is required here.

Operator setup. The authorization server is served by the router, off by default. Enable it and point it at the externally-reachable URLs clients use:

# router config (signaldb.toml)
[mcp.oauth]
enabled = true
issuer_url = "https://signaldb.example.org"        # this AS, as clients reach it
resource_url = "https://signaldb.example.org/mcp"  # the MCP resource tokens bind to
# access_token_ttl = "1h"; refresh_token_ttl = "30d"; authorization_code_ttl = "60s"

The signaldb mcp sidecar advertises the same resource so an unauthenticated request is challenged toward discovery — pass the matching URLs:

signaldb mcp \
  --oauth-resource-url https://signaldb.example.org/mcp \
  --oauth-issuer-url   https://signaldb.example.org

Tokens are opaque, catalog-backed, and audience-bound to resource_url; revoking one is a row delete. POST /oauth/introspect (RFC 7662) reports whether a bearer token is active and, if so, its full granted-tenant set, scopes, audience, and expiry — this is how the signaldb mcp sidecar learns a multi-tenant connector's whole reachable set before any one tenant has been selected for a call; it isn't something you call directly as an operator or agent. The read scopes a token may hold — traces:read, logs:read, metrics:read, profiles:read, schema:read, processors:read, evals:read — gate the corresponding query surface (see the multi-tenancy model); a request with no scope is granted all of them, and schema:write/processors:write/evals:write are never grantable through OAuth (a request naming only one of them is rejected with invalid_scope). The existing Bearer <api-key> + X-Tenant-ID path is unchanged; OAuth is an added credential type, not a replacement.

Audit and observability

Every tool call is audited: after it completes, the server emits exactly one structured log event (target signaldb_mcp::audit) with bounded fields — never the arguments, the query expression, or the result:

Field Meaning
tool The tool name (search_traces, get_trace, …).
tenant_id The tenant the router resolved the caller to.
dataset The dataset the call named (its dataset argument, else X-Dataset-ID); absent when neither is set.
session_id The Mcp-Session-Id (stdio on the stdio transport).
outcome ok, truncated (result cut at the size cap), denied, throttled, or error.
duration_ms Wall time of the call, including any wait for a concurrency permit.
error.type Only for outcome=error: concurrency_limit, deadline, tool_error, the router's HTTP status (500), or the JSON-RPC code.

Levels: ok, truncated, and throttled log at info; denied (the router rejected the credential or the tenant/dataset access — a 401/403) at warn, so probing is visible; error at error. A call refused at the concurrency bound is outcome=error, error.type=concurrency_limit; a call still running past the total per-call deadline (see Running it) is outcome=error, error.type=deadline.

The same call is one tools/call {tool} span (INTERNAL, gen_ai.tool.name, mcp.session.id, signaldb.tenant.id, signaldb.dataset.id; status Error only for outcome=error), the HTTP request that carried it is a POST /mcp server span parented to the client's traceparent, and two metrics count the calls: signaldb.mcp.tool_calls by tool and outcome and signaldb.mcp.tool_call.duration by tool (Prometheus: signaldb_mcp_tool_calls_total{gen_ai_tool_name,signaldb_mcp_outcome} and signaldb_mcp_tool_call_duration_seconds{gen_ai_tool_name}). All of it is exported when [self_monitoring] is enabled for the sidecar (service name signaldb-mcp); see docs/operations/self-monitoring-traces.md.

Example flow

Configuring a new application to send data here? Call connection_info first — it returns the deployment's public OTLP endpoints, the headers to send, and ready-to-paste OTEL_EXPORTER_OTLP_* env vars — then mint an ingest key with tenant_create_api_key and substitute it for the placeholder.

  1. server_info — confirm you are connected as the expected tenant.
  2. discover_attributes — list tag names, then values for service.name.
  3. search_traces with { .service.name = "checkout" && status = error } and a time range to find failing requests.
  4. get_trace with an ID from the search results to inspect the full trace.

To explore logs or metrics instead: discover_attributes with signal: "logs" lists Loki labels (add tag for a label's values); signal: "metrics" does the same for Prometheus labels. discover_metrics lists metric names directly, for building a query_metrics PromQL expression.

Before filtering or grouping by a name you are unsure of, ask the schema registry what it means: resolve_attribute with key: "k8s.pod.uid" (or resolve_entity / resolve_metric, or search_schema with kind: "attribute", prefix: "k8s.pod.") returns namespace-tagged, precedence-ordered definitions — a tenant's own conventions (uploaded with create_schema_registry) come first, the bundled OpenTelemetry definition is kept as an alternative. discover_* tells you which names have data; resolve_* tells you what they mean. resolve_entity with name: "gen_ai.agent" tells an agent which attribute identifies an AI agent (gen_ai.agent.id).

From the CLI

The same discovery is available outside an agent session, via signaldb-sdk like every other CLI capability:

# Native surface (Query IR, logical dotted names, no scan)
signaldb-cli discover sources
signaldb-cli discover fields --source logs
signaldb-cli discover values --source traces --field span.kind
signaldb-cli discover values --source traces --field http.route --sample

# Compatibility-dialect view (Tempo tags, Loki/Prometheus labels)
signaldb-cli discover attributes --signal traces --tag service.name
signaldb-cli discover attributes --signal logs
signaldb-cli discover attributes --signal metrics --tag job
signaldb-cli discover metrics

discover fields/values/sources are the native surface: they speak the same logical names as a Query IR document and are answered from metadata rather than by scanning. discover values reads data only when you pass --sample, and the response says so — without it you are told what would answer the question instead. discover attributes remains the dialect-shaped view, for parity with what Grafana sees. See the Query IR reference.

Schema-registry lookup and custom-registry management mirror the schema tools (reads need a key with schema:read, mutations schema:write):

signaldb-cli schema registry list
signaldb-cli schema registry get otel 1.43.0
signaldb-cli schema attribute get k8s.pod.uid
signaldb-cli schema entity get k8s.pod
signaldb-cli schema metric search k8s.pod. --limit 20
signaldb-cli admin schema validate --file conventions.yaml
signaldb-cli admin schema create --file conventions.yaml     # YAML or JSON
signaldb-cli admin schema replace acme 1.0.0 --file conventions.yaml
signaldb-cli admin schema delete acme 1.0.0

server_info mirrors signaldb-cli whoami; query_metrics/search_logs's range mode mirrors signaldb query --promql|--logql ... --start ... --end ...; get_trace mirrors signaldb query --trace-id <id>:

signaldb-cli whoami
signaldb-cli query --promql 'up' --start 0 --end 3600 --step 15s
signaldb-cli query --trace-id 4bf92f3577b34da6a3ce929d0e0e4736

get_source_context mirrors signaldb-cli tenant source-context — a read tool, like get_trace/get_profile, so any valid key of the tenant works; it does not need tenant:manage despite living under the CLI's tenant verb group:

signaldb-cli tenant source-context --path src/main.rs --line 42 --api-key sk-your-key --tenant-id your-tenant

The tenant_* tools mirror signaldb-cli tenant: tenant_info is tenant show, the table tools are tenant table ... (any valid key of the tenant), and the management tools are tenant dataset|api-key|membership|schema|github ... (a key carrying tenant:manage; destructive verbs prompt on a TTY unless --yes):

signaldb-cli tenant show --api-key sk-your-key --tenant-id your-tenant
signaldb-cli tenant table list --api-key sk-your-key --tenant-id your-tenant
signaldb-cli tenant table provision --api-key sk-your-key --tenant-id your-tenant
signaldb-cli tenant table schemas --api-key sk-your-key --tenant-id your-tenant
signaldb-cli tenant table available-schemas --api-key sk-your-key
signaldb-cli tenant dataset create staging --api-key sk-manage-key --tenant-id your-tenant
signaldb-cli tenant api-key create --name ci --scope traces:write --api-key sk-manage-key --tenant-id your-tenant
signaldb-cli tenant membership set alice@example.com --role member --api-key sk-manage-key --tenant-id your-tenant

Platform administration (list_tenants, create_dataset, revoke_api_key, ...) mirrors the admin command group, authenticated with the administrative key instead of a tenant key — see signaldb-cli admin --help.