Skip to content

LogQL reference

SignalDB exposes a Grafana Loki-compatible query API for the logs signal, served by the router at http://<router-host>:3000/loki. Point Grafana's Loki datasource at that base URL, or call the endpoints directly.

This page documents the supported LogQL surface and — importantly — the approximations and gaps, so you know what to expect. For sending logs, see Sending OTLP data; for the trace API, the Tempo API reference.

Authentication

Every request is authenticated and tenant-scoped, like the other query APIs (see Authentication):

  • Authorization: Bearer <api-key>
  • X-Tenant-ID: <tenant>
  • X-Dataset-ID: <dataset> (optional; the tenant's default dataset is used otherwise)

Endpoints

Endpoint Status
GET /loki/api/v1/query_range Range query. Returns a streams result for log queries, a matrix for metric queries
GET /loki/api/v1/query Instant query. Returns the most recent lines over a one-hour window ending at time (log queries only)
GET /loki/api/v1/labels Label names available in the window
GET /loki/api/v1/label/{name}/values Distinct values of one label
GET /loki/api/v1/series Series (label sets) matching a selector
GET /loki/api/v1/detected_fields Discover attribute fields in a window: name, inferred type, approximate cardinality (samples the data; no declaration or indexing needed)
GET /loki/api/v1/tail Not implemented — live tail is tracked separately

Labels are also reachable without raw HTTP: signaldb-cli discover attributes --signal logs [--tag NAME], and the MCP discover_attributes (signal: "logs") tool for AI agents — see the MCP server doc.

Common query parameters: query (the LogQL string), start/end (unix nanoseconds, unix seconds, or RFC3339), limit, direction (forward/backward), and step (metric queries; Go duration like 30s or a number of seconds).

Labels and how they map to storage

SignalDB stores logs in columnar form, not as free-form label sets. LogQL labels resolve as follows:

LogQL label Resolves to
service_name, service, job, service.name the service_name column
level, severity, detected_level the severity_text column
trace_id, span_id the matching columns
a materialized label (see below) its dedicated label_<key> column
any other label the log_attributes / resource_attributes maps

Labels backed by a column are exact. On tables created since attributes became typed maps, any other label is also exact: the value is looked up per key, and regex plus ordered comparisons (| status >= 500) work on every attribute. Older tables store attributes as serialized JSON, where a label is matched by its "key":"value" fragment — an approximation that can over-match and supports only =/!=; the querier picks the right form per table automatically.

A label name may contain dots ({k8s.pod.name="checkout-7c9f"}, | http.response.status_code >= 500), so a query can name an attribute by its real OTel key. Apart from the well-known aliases in the table above (service.name reaches the service_name column), a dotted key resolves directly against the attribute maps by exact key — no materialization needed. The underscore spelling of the same attribute (k8s_pod_name) only resolves to that data once the label has been materialized (see below): both spellings sanitize to the identical label_<key> column, so either works once the column exists, but only the dotted form is guaranteed to match beforehand.

Materialized labels

Configured attribute keys are promoted to dedicated columns at ingest via [schema.materialized_labels] (see signaldb.dist.toml and storage layout). A materialized label is matched exactly — no substring over-match — and, unlike JSON attributes, supports regex (=~ / !~) and ordered comparisons (| latency_ms > 500). The value is compared as a string for = / != / =~, and cast to a number for > / >= / < / <=.

Materialization applies to data written after the label was configured and to tables created after that point; a table that predates the column falls back to the JSON substring match for that label. Because the promoted value is also kept in the attribute JSON, label discovery (/labels, /label/{name}/values) is unchanged.

Two distinct label keys can sanitize to the same label_<key> column name (see the dotted-vs-underscore example above); the writer resolves that by suffixing the later key's column (label_<key>_2). As an interim guard (#1533), a query against either colliding key currently falls back to the attribute-map extraction path rather than risk reading the wrong key's column; full per-key resolution of the collision is still open.

Series identity (in /series results and bare range aggregations such as count_over_time(...) with no vector wrapper) is the service_name and level labels. A vector aggregation with no by clause — sum(count_over_time(...)) — collapses every matching series into one, as in Loki; add by (...) to keep a grouping.

Structured metadata

A returned log entry carries per-line fields as Loki structured metadata: the trace context (trace_id, span_id) and the record's log and resource attributes. These are per-entry rather than stream labels because their values vary line to line — promoting them to the label set would split every distinct combination into its own stream.

Structured metadata is a flat string map, so the three OTel attribute scopes (resource, instrumentation scope, log record) are flattened into one map here, and a key present at more than one scope resolves to the log-record value. This is a limitation of the Loki wire format, not of storage. To read attributes with their scopes intact — and to reach instrumentation-scope attributes, trace_flags, severity_number or observed_timestamp at all — use the native Query IR, which keeps each container separate.

Discovering fields

GET /loki/api/v1/detected_fields answers "what attributes exist here?" without any declaration: for a window (and optional query stream selector) it samples the stored attribute documents and returns each key with an inferred type (string/int/float/boolean) and an approximate distinct-value count. Parameters: start, end, query, limit (default 100). Works on both JSON-string and map-typed attribute tables. Cardinality is a lower-bound estimate from the sample — intended for exploration UIs (faceted field pickers), not exact counts.

Log queries

A log query is a stream selector followed by an optional pipeline, and returns log lines.

{service_name="api"}
{service_name="api", level="error"}
{namespace=~"prod-.*"}
{service_name="api"} |= "timeout"
{service_name="api"} |= "error" != "healthcheck" |~ "5\\d\\d"
{service_name="api"} | json | level="error"

Stream selectors

All four matchers are supported: =, !=, =~, !~. Regex matchers on a column (service_name=~"api.*") are pushed down as regexp_like; regex against an attribute label is not supported (attribute labels support only = and !=). As in Loki, a negative matcher (!=, !~) also matches a stream that lacks the label entirely, not only one holding a different value.

Line filters

|=, !=, |~, !~ filter the log body (|~/!~ via regexp_like). The ip("...") filter form is not supported yet.

Pipeline (parser) stages

Parser and formatter stages parse — json, logfmt (with flags), regexp, pattern, unpack, decolorize, line_format, label_format, drop, keep, distinct — but do not themselves filter rows: they are accepted and pass through. A label filter after a stage (| level="error", | status="500") does filter, resolving labels via the mapping above.

Label filters support =, !=, =~, !~ against columns (including materialized labels) and against map-typed attributes, combined with and/or/comma. Ordered comparisons (>, >=, <, <=, e.g. | status >= 500) work on materialized labels and on any attribute of a map-typed table — values are cast to a number. Legacy JSON-attribute tables support only =/!= on non-materialized attributes.

Metric queries

Metric queries aggregate log streams into a numeric matrix and are evaluated through query_range. Supported forms:

count_over_time({service_name="api"}[5m])
rate({service_name="api"} |= "error" [5m])
bytes_over_time({service_name="api"}[5m])
bytes_rate({service_name="api"}[5m])
sum_over_time({service_name="api"} | unwrap duration_ms [5m])
avg_over_time({service_name="api"} | unwrap duration_ms [5m])
sum by (level) (rate({service_name="api"}[5m]))
  • Range functions: count_over_time, rate, bytes_over_time, bytes_rate, and the unwrap-based sum_over_time / avg_over_time / min_over_time / max_over_time / stddev_over_time / stdvar_over_time / first_over_time / last_over_time / quantile_over_time.
  • Vector aggregation: a single outer sum / avg / min / max / count / stddev / stdvar, with by (...) or without (), wrapping one of the above. sum folds into the grouped range aggregate; the others reduce the per-series range aggregation across the series in each group (stddev/stdvar are population statistics, as in Prometheus).
  • Selection: topk(k, ...) / bottomk(k, ...) keep the k highest / lowest series by value in each time bucket, applied after aggregation.
  • Ordering: sort(...) / sort_desc(...) order the result series by value. A range matrix has a value per bucket, so ordering uses each series' value at its latest bucket (sort is strictly defined only for instant queries).
  • Scalar arithmetic: a metric query combined with a scalar literal using + / - / * / / (rate(...) / 60, 100 * rate(...)), scaling every value. Operand order is preserved for - and /.
  • Vector arithmetic: two metric queries combined with + / - / * / / (sum(rate(...)) / sum(rate(...)) for a ratio). Series are matched one-to-one on (bucket, labels); unmatched series are dropped.
  • Vector matching: an on (...) / ignoring (...) clause after the operator restricts which labels series match on (a / on(service) b, a / ignoring(level) b). Matching is one-to-one; a group_left / group_right modifier is accepted but does not enable many-to-one.
  • Vector comparison: two metric queries compared with == / != / > / >= / < / <=. Without bool the left series is kept where the comparison holds (a filter); with bool (a > bool b) each matched pair becomes 1 or 0.
  • Logical / set: and (left series with a right match), or (union of both sides), unless (left series with no right match). Matching is on (bucket, labels); values are carried through, never combined.
  • vector(N): a scalar promoted to a constant, no-label series valued N at every step bucket (as in ... or vector(0) in Prometheus, though the or form itself is not supported yet).
  • label_replace(v, dst, replacement, src, regex): rewrites the dst label from a regex capture of src ($1 / ${name} in replacement). The regex is anchored to the whole source value; a non-match leaves the series unchanged, and an empty result deletes the label. A new dst becomes an additional series label.
  • Grouping by a label backed by a column works as expected: the built-in ones (service_name, level, ...) and any materialized label (sum by (namespace) (...) groups on its label_namespace column; the result series carry the label under its sanitized name). Grouping by any other name is also accepted — it resolves as an attribute, coalesced across containers the same way a where clause on that name would — and if no record in the queried window actually carries that attribute, every row collapses into one series with that label absent, same as real Loki: sum by (foo) (...) for a label no stream carries returns one series with foo absent/empty rather than an error, whether foo is a column we happen to store or an attribute we don't.

Time bucketing — the key approximation

Loki evaluates a range aggregation over a sliding [range] window at each step. SignalDB instead buckets into fixed, step-aligned windows with date_bin(step, timestamp). This is exact when step equals the range (Grafana's default for count_over_time panels) and an approximation otherwise. Set the panel step equal to the range window for exact results.

Not supported yet

These parse but return an "unsupported" error at execution:

  • Many-to-one vector matching (group_left / group_right) and ... or vector(0) fallbacks
  • sum without (labels) with a non-empty label list

Examples

Range log query for one service, newest first:

curl -G 'http://localhost:3000/loki/api/v1/query_range' \
  -H 'Authorization: Bearer sk-...' -H 'X-Tenant-ID: acme' \
  --data-urlencode 'query={service_name="api"} |= "error"' \
  --data-urlencode 'start=1700000000' --data-urlencode 'end=1700003600' \
  --data-urlencode 'limit=100'

Error rate per level over five-minute buckets:

curl -G 'http://localhost:3000/loki/api/v1/query_range' \
  -H 'Authorization: Bearer sk-...' -H 'X-Tenant-ID: acme' \
  --data-urlencode 'query=sum by (level) (rate({service_name="api"}[5m]))' \
  --data-urlencode 'start=1700000000' --data-urlencode 'end=1700003600' \
  --data-urlencode 'step=5m'

Errors

Failures return a JSON body in the Prometheus/Loki error shape:

{
  "status": "error",
  "errorType": "bad_data",
  "error": "expected matcher operator, found end of input at line 1, column 10"
}

errorType is bad_data (400), not_found (404), rate_limited (429), timeout (504), unavailable (503, no querier), not_implemented (501), or internal (500). A query that matches nothing is not an error — it returns 200 with an empty result array.

A 429 (per-tenant query rate limit exceeded) additionally carries retryAfterMs in the body and three response headers computed from the tenant's actual token-bucket state: Retry-After (whole seconds, rounded up, at least 1), X-RateLimit-Limit (the per-second budget), and X-RateLimit-Burst (the burst allowance):

{
  "status": "error",
  "errorType": "rate_limited",
  "error": "tenant 'acme' exceeded its query request rate limit; retry after 1s or raise the tenant's limits",
  "retryAfterMs": 1000
}

Query-demand statistics

Attribute labels used in filters (any label that is not a dedicated column) are counted as query demand and flushed to the catalog's advisory attribute_stats table, where they inform which attributes are worth materializing (see the storage layout docs on materialized labels).