Skip to content

Signal Table Provisioning

Every tenant/dataset SignalDB knows about converges on an Iceberg table for each signal type enabled for it, so a dataset becomes complete without waiting for telemetry to arrive for each signal.

Creating a dataset — through the management API or the MCP tools that proxy them — provisions its tables synchronously, best-effort, before the creation call returns: the new dataset is queryable immediately in the common case. That provisioning attempt never blocks or fails the dataset's creation itself; a partial or total provisioning failure is logged as a warning and leaves the dataset row committed regardless.

For everything that provisioning attempt doesn't catch — the writer was down when the dataset was created, the attempt itself failed, or the dataset predates this behavior — convergence is eventual: the next reconcile pass (within table_reconcile_interval), an on-demand request, or the dataset's first write, whichever comes first. Queries do not wait for any of that — a signal with no table yet reads as an empty result, never an error.

The writer runs the reconciler continuously as the backstop; nothing needs to be run by hand. This page describes the knob, what a pass does, and how to tell whether it is working.

What gets provisioned

Up to five tables per dataset, gated on the signal types enabled for that tenant — a tenant that carries its own [schema] block narrows the set, so one that disabled metrics gets no metrics/metric_exemplars tables:

Signal Tables Gate
Traces traces default_schemas.traces_enabled
Logs logs default_schemas.logs_enabled
Metrics metrics, metric_exemplars default_schemas.metrics_enabled
Profiles profiles default_schemas.profiles_enabled

Custom tables declared in default_schemas.custom_schemas are a config-only concept and are never provisioned.

Provisioned tables are what the ingest path would have created: the same schemas, partition specs, bloom-filter and compression settings, and metadata-pruning properties. A later write lands in them — it does not create a second table and cannot conflict.

Configuration

[writer]
# How often the writer re-runs the reconciler over the tenant registry.
# A pass always runs at startup; this governs only the periodic re-run.
# "0s" keeps the startup pass and disables the periodic one, so datasets
# created after startup wait for a first write, an on-demand request, or a
# writer restart.
table_reconcile_interval = "5m"

Environment override: SIGNALDB__WRITER__TABLE_RECONCILE_INTERVAL=90s.

Set it shorter if you create tenants frequently and want their tables to appear sooner; longer if your catalog is remote and you want less background traffic. A converged deployment issues no catalog calls per pass either way (see Steady-state cost), so the interval mostly bounds how long a newly created dataset waits.

What a pass does

  1. Enumerate the tenant registry — config-defined tenants and tenants created through the management API. Provisioning is driven by the registry and each tenant's enabled signal types alone; it is independent of who (if anyone) holds a membership in the tenant, so a dataset is provisioned whether its users authenticate by password, API key, or SSO.
  2. For each tenant, take its datasets. The registry guarantees a tenant's default_dataset is among them even when no dataset row names it, so a tenant whose default exists only as a column on its tenant row is provisioned like any other.
  3. For each dataset, load-or-create every enabled table.
  4. For each dataset, drop any of the five legacy per-type metric tables (metrics_gauge, metrics_sum, metrics_histogram, metrics_exponential_histogram, metrics_summary) still present — the otel-native-schema layer 7 cutover to the typed metrics/metric_exemplars tables. Dropped data is not migrated; see the metrics table.

Datasets created while the writer was down, datasets predating this behavior, and datasets added at runtime all converge on a later pass — no restart, no configuration-file edit.

Loading an existing table also reconciles its metadata, which is how tables created by older builds catch up. Each step is idempotent, metadata-only, and commits at most once per table: metadata-pruning properties are backfilled, the canonical sort order is declared if the table predates it, and the table's schema is evolved to the current schemas.toml version. A step that loses a commit race is logged at warn and retried on the next load; it never fails the pass.

Every table named by a pass also has its stale WAL idempotency markers retired, using the same [writer].wal_marker_retention rules the WAL commit path applies. This is what retires markers on a table that has gone dormant — one no IcebergTableWriter commits to anymore, and so is never reached by the commit-path sweep — since the reconciler walks the tenant/dataset registry directly rather than a process's cache of active writers.

Steady-state cost

The writer remembers, per process, which (tenant, dataset, table) triples it has confirmed, and skips those datasets entirely on later passes. A fully converged deployment therefore issues no catalog traffic per pass. The memory is rebuilt on restart, so the startup pass is always a real check.

Failures

A provisioning failure is never fatal and never blocks ingest or queries. It is logged at warn with the tenant, dataset, and table it concerns, and retried on the next pass. One table's failure does not abort its siblings, and one dataset's failure does not abort the rest of the pass.

Because the ingest path independently creates any table it needs, a persistently failing reconciler degrades to create-on-first-write — the prior behavior — rather than to data loss.

Provisioning at dataset creation

POST /api/v1/tenants/{tenant_id}/datasets — reachable by a tenant-admin session/tenant:manage-scoped key, or by the break-glass admin key with no tenant — and the MCP create_dataset/tenant_create_dataset tools, which proxy it — provision the new dataset's enabled tables right after the dataset row commits, using the same CatalogManager::ensure_dataset_tables path a reconcile pass uses per dataset. The response still reflects the dataset's creation, not the provisioning outcome:

  • The dataset is created (201) whether or not provisioning succeeds.
  • A provisioning failure — any table, or building the catalog manager itself — is logged at warn with the tenant, dataset, and reason, and never surfaces in the HTTP response.
  • The periodic reconciler still runs over every tenant and dataset regardless, so a provisioning failure here converges on the next pass rather than staying broken.

This makes the reconcile interval matter only for datasets this synchronous step didn't cover — not the common "just created a dataset" case.

Provisioning a tenant on demand

To create a tenant's tables immediately instead of waiting for the next pass, and to list what is already provisioned, POST /api/v1/tenants/{id}/tables/create and GET /api/v1/tenants/{id}/tables are in the OpenAPI contract, so every client surface reaches them through the generated SDK — raw HTTP, the CLI, an MCP tool, and the UI's management area:

# raw HTTP
curl -X POST http://localhost:3000/api/v1/tenants/acme/tables/create \
  -H "Authorization: Bearer $SIGNALDB_API_KEY" \
  -H "X-Tenant-ID: acme"

# CLI (signaldb-cli tenant table)
signaldb-cli tenant table list --api-key "$SIGNALDB_API_KEY" --tenant-id acme
signaldb-cli tenant table provision --api-key "$SIGNALDB_API_KEY" --tenant-id acme

Through an MCP agent session: the tenant_list_tables and tenant_create_tables tools (see the MCP server doc). Through the web UI: the management area's Tables section for a dataset, with a Provision tables action.

The call provisions the tenant's datasets before returning 201, and returns 500 if any table could not be created. Tenant-administrator privileges are required — in practice any valid tenant API key, since a tenant-scoped key is already trusted to shape its own tenant's infrastructure (see TenantContext::can_manage_tenant); this is a lower bar than the management API's dataset/API-key/membership endpoints, which require a tenant-admin session or a key carrying tenant:manage.

GET /api/v1/tenants/{id}/tables (tenant table list / tenant_list_tables) lists what is really in the Iceberg catalog — the same place provisioning writes to — grouped by dataset. Each table carries the dataset it belongs to (dataset on every entry); the response also groups them under datasets (one entry per dataset, each with its own tables), while the flat tables list is kept for existing consumers.

  • datasets always has one entry per dataset the tenant is known to have, even one with nothing provisioned in it yet — that entry's tables is simply empty, it is never omitted. A tenant with no datasets at all (or one this endpoint cannot resolve) gets datasets: [].
  • tables (the flat list) contains only tables that actually exist — never invented, never one per known dataset.
  • A listing right after POST …/tables/create reflects exactly what was just created: the provisioned dataset's group gains those tables, and tables gains their flat entries.
$ signaldb-cli tenant table list --api-key "$SIGNALDB_API_KEY" --tenant-id acme
DATASET     TABLE    TYPE
production  logs     logs
production  traces   traces

Pass --json for the raw response instead of the formatted table.

Verifying

Metrics, exported over OTLP by self-monitoring and attributed by signaldb.tenant.id and signaldb.dataset.id:

Instrument Meaning
signaldb.writer.tables_provisioned Tables the reconciler created
signaldb.writer.table_provisioning_failures Tables it could not create

A rising failure count means the deployment has degraded to create-on-first-write; check the writer's logs for the tenant, dataset, and table, and confirm the Iceberg catalog is reachable.

Logs: a pass that created or failed anything logs at info with datasets_checked, datasets_skipped, tables_created, and tables_failed. A converged pass logs at debug.

Consequences to expect

  • Empty tables multiply catalog objects. Each provisioned dataset gains up to five catalog rows and five metadata.json files. They carry no snapshots and no data files, so retention, orphan cleanup, storage accounting, and the compactor all treat them as no-ops.
  • Materialized labels are fixed at creation time. [schema] materialized_labels is applied when a table is created; ensure_table's own catch-up path (below) only brings a table's schemas.toml-declared columns forward, it does not retrofit a changed label configuration onto an existing table. Provisioning means the initial labels-at-creation behavior now happens for every dataset, not only for ones that ingest, so prefer setting materialized_labels before a dataset is provisioned. The compactor no longer adds label_<key> columns; its attribute-promotion pass adds typed attr_<level>_<key> copies of hot attributes instead — see Attribute Promotion.
  • ensure_table does evolve an existing table's schema — every schemas.toml-sourced signal (traces, logs, both metrics tables, and profiles). Every load, not just creation, brings the table's schema forward to the current schemas.toml version if it's behind. New columns are always nullable and historical rows are never rewritten to backfill them; beyond that, a table whose recorded starting version is trusted (found on the version chain) also gets renames and removals applied hop by hop, not just additions — an untrusted starting version (no recorded version, or one not found on the chain) skips straight to current with additions only. See schema evolution.
  • The typed-attribute-layout cutover recreates a legacy table instead of evolving it, and that recreation drops its data. Every signal's current schemas.toml version is now the typed attribute layout (one typed map per canonical type plus a binary residue per attribute container, replacing the single map<string,string> column -- see Typed attribute layout). Iceberg schema evolution cannot add or remove a map-typed column on a live table, so ensure_table detects a table still in the legacy layout and, instead of evolving it, drops and recreates it fresh at the current typed version -- with the same schema, partitioning, bloom/compression properties, and sort order a brand-new table gets. The dropped table's data is not migrated; this is a deliberate one-shot cutover under the post-1.0 breaking-changes policy, not a bug. It happens the first time any reconcile pass or write touches a still-legacy table after upgrading past the cutover, so plan the upgrade around each table's retention window if the pre-cutover data matters.
  • Not every table property is set at creation. Provisioning applies the bloom-filter, column-statistics, compression and metadata-pruning properties, but deliberately not write.target-file-size-bytes: compaction reconciles that one against [compactor].target_file_size_mb before each rewrite, so it cannot drift when an operator retunes the target. A provisioned table therefore shows no such property until its first compaction — see output file size.