LogQL reference¶
SignalDB exposes a Grafana Loki-compatible query API for the logs signal,
served by the router at http://<router-host>:3000/loki. Point Grafana's
Loki datasource at that base URL, or call the endpoints directly.
This page documents the supported LogQL surface and — importantly — the approximations and gaps, so you know what to expect. For sending logs, see Sending OTLP data; for the trace API, the Tempo API reference.
Authentication¶
Every request is authenticated and tenant-scoped, like the other query APIs (see Authentication):
Authorization: Bearer <api-key>X-Tenant-ID: <tenant>X-Dataset-ID: <dataset>(optional; the tenant's default dataset is used otherwise)
Endpoints¶
| Endpoint | Status |
|---|---|
GET /loki/api/v1/query_range |
Range query. Returns a streams result for log queries, a matrix for metric queries |
GET /loki/api/v1/query |
Instant query. Returns the most recent lines over a one-hour window ending at time (log queries only) |
GET /loki/api/v1/labels |
Label names available in the window |
GET /loki/api/v1/label/{name}/values |
Distinct values of one label |
GET /loki/api/v1/series |
Series (label sets) matching a selector |
GET /loki/api/v1/detected_fields |
Discover attribute fields in a window: name, inferred type, approximate cardinality (samples the data; no declaration or indexing needed) |
GET /loki/api/v1/tail |
Not implemented — live tail is tracked separately |
Labels are also reachable without raw HTTP: signaldb-cli discover
attributes --signal logs [--tag NAME], and the MCP discover_attributes
(signal: "logs") tool for AI agents — see the MCP server doc.
Common query parameters: query (the LogQL string), start/end
(unix nanoseconds, unix seconds, or RFC3339), limit, direction
(forward/backward), and step (metric queries; Go duration like
30s or a number of seconds).
Labels and how they map to storage¶
SignalDB stores logs in columnar form, not as free-form label sets. LogQL labels resolve as follows:
| LogQL label | Resolves to |
|---|---|
service_name, service, job, service.name |
the service_name column |
level, severity, detected_level |
the severity_text column |
trace_id, span_id |
the matching columns |
| a materialized label (see below) | its dedicated label_<key> column |
| any other label | the log_attributes / resource_attributes maps |
Labels backed by a column are exact. On tables created since attributes
became typed maps, any other label is also exact: the value is looked
up per key, and regex plus ordered comparisons (| status >= 500) work on
every attribute. Older tables store attributes as serialized JSON, where a
label is matched by its "key":"value" fragment — an approximation that
can over-match and supports only =/!=; the querier picks the right
form per table automatically.
A label name may contain dots ({k8s.pod.name="checkout-7c9f"},
| http.response.status_code >= 500), so a query can name an attribute by
its real OTel key. Apart from the well-known aliases in the table above
(service.name reaches the service_name column), a dotted key resolves
directly against the attribute maps by exact key — no materialization
needed. The underscore spelling of
the same attribute (k8s_pod_name) only resolves to that data once the
label has been materialized (see below): both spellings sanitize to
the identical label_<key> column, so either works once the column
exists, but only the dotted form is guaranteed to match beforehand.
Materialized labels¶
Configured attribute keys are promoted to dedicated columns at ingest via
[schema.materialized_labels] (see signaldb.dist.toml and
storage layout).
A materialized label is matched exactly — no substring over-match — and,
unlike JSON attributes, supports regex (=~ / !~) and ordered
comparisons (| latency_ms > 500). The value is compared as a string for
= / != / =~, and cast to a number for > / >= / < / <=.
Materialization applies to data written after the label was configured and
to tables created after that point; a table that predates the column falls
back to the JSON substring match for that label. Because the promoted value
is also kept in the attribute JSON, label discovery (/labels,
/label/{name}/values) is unchanged.
Two distinct label keys can sanitize to the same label_<key> column name
(see the dotted-vs-underscore example above); the writer resolves that by
suffixing the later key's column (label_<key>_2). As an interim guard
(#1533), a query against either colliding key currently falls back to the
attribute-map extraction path rather than risk reading the wrong key's
column; full per-key resolution of the collision is still open.
Series identity (in /series results and bare range aggregations such as
count_over_time(...) with no vector wrapper) is the service_name and
level labels. A vector aggregation with no by clause —
sum(count_over_time(...)) — collapses every matching series into one, as in
Loki; add by (...) to keep a grouping.
Structured metadata¶
A returned log entry carries per-line fields as Loki structured metadata:
the trace context (trace_id, span_id) and the record's log and resource
attributes. These are per-entry rather than stream labels because their values
vary line to line — promoting them to the label set would split every distinct
combination into its own stream.
Structured metadata is a flat string map, so the three OTel attribute scopes
(resource, instrumentation scope, log record) are flattened into one map
here, and a key present at more than one scope resolves to the log-record
value. This is a limitation of the Loki wire format, not of storage. To read
attributes with their scopes intact — and to reach instrumentation-scope
attributes, trace_flags, severity_number or observed_timestamp at all —
use the native Query IR, which keeps each container separate.
Discovering fields¶
GET /loki/api/v1/detected_fields answers "what attributes exist here?"
without any declaration: for a window (and optional query stream
selector) it samples the stored attribute documents and returns each key
with an inferred type (string/int/float/boolean) and an
approximate distinct-value count. Parameters: start, end, query,
limit (default 100). Works on both JSON-string and map-typed attribute tables. Cardinality is a lower-bound estimate from the
sample — intended for exploration UIs (faceted field pickers), not exact
counts.
Log queries¶
A log query is a stream selector followed by an optional pipeline, and returns log lines.
{service_name="api"}
{service_name="api", level="error"}
{namespace=~"prod-.*"}
{service_name="api"} |= "timeout"
{service_name="api"} |= "error" != "healthcheck" |~ "5\\d\\d"
{service_name="api"} | json | level="error"
Stream selectors¶
All four matchers are supported: =, !=, =~, !~. Regex matchers on
a column (service_name=~"api.*") are pushed down as regexp_like;
regex against an attribute label is not supported (attribute labels
support only = and !=). As in Loki, a negative matcher (!=, !~)
also matches a stream that lacks the label entirely, not only one holding
a different value.
Line filters¶
|=, !=, |~, !~ filter the log body (|~/!~ via regexp_like).
The ip("...") filter form is not supported yet.
Pipeline (parser) stages¶
Parser and formatter stages parse — json, logfmt (with flags),
regexp, pattern, unpack, decolorize, line_format,
label_format, drop, keep, distinct — but do not themselves
filter rows: they are accepted and pass through. A label filter after
a stage (| level="error", | status="500") does filter, resolving
labels via the mapping above.
Label filters support =, !=, =~, !~ against columns (including
materialized labels) and against map-typed attributes, combined with
and/or/comma. Ordered comparisons (>, >=, <, <=, e.g.
| status >= 500) work on materialized labels and on any attribute of
a map-typed table — values are cast to a number. Legacy JSON-attribute
tables support only =/!= on non-materialized attributes.
Metric queries¶
Metric queries aggregate log streams into a numeric matrix and are
evaluated through query_range. Supported forms:
count_over_time({service_name="api"}[5m])
rate({service_name="api"} |= "error" [5m])
bytes_over_time({service_name="api"}[5m])
bytes_rate({service_name="api"}[5m])
sum_over_time({service_name="api"} | unwrap duration_ms [5m])
avg_over_time({service_name="api"} | unwrap duration_ms [5m])
sum by (level) (rate({service_name="api"}[5m]))
- Range functions:
count_over_time,rate,bytes_over_time,bytes_rate, and the unwrap-basedsum_over_time/avg_over_time/min_over_time/max_over_time/stddev_over_time/stdvar_over_time/first_over_time/last_over_time/quantile_over_time. - Vector aggregation: a single outer
sum/avg/min/max/count/stddev/stdvar, withby (...)orwithout (), wrapping one of the above.sumfolds into the grouped range aggregate; the others reduce the per-series range aggregation across the series in each group (stddev/stdvarare population statistics, as in Prometheus). - Selection:
topk(k, ...)/bottomk(k, ...)keep thekhighest / lowest series by value in each time bucket, applied after aggregation. - Ordering:
sort(...)/sort_desc(...)order the result series by value. A range matrix has a value per bucket, so ordering uses each series' value at its latest bucket (sortis strictly defined only for instant queries). - Scalar arithmetic: a metric query combined with a scalar literal
using
+/-/*//(rate(...) / 60,100 * rate(...)), scaling every value. Operand order is preserved for-and/. - Vector arithmetic: two metric queries combined with
+/-/*//(sum(rate(...)) / sum(rate(...))for a ratio). Series are matched one-to-one on(bucket, labels); unmatched series are dropped. - Vector matching: an
on (...)/ignoring (...)clause after the operator restricts which labels series match on (a / on(service) b,a / ignoring(level) b). Matching is one-to-one; agroup_left/group_rightmodifier is accepted but does not enable many-to-one. - Vector comparison: two metric queries compared with
==/!=/>/>=/</<=. Withoutboolthe left series is kept where the comparison holds (a filter); withbool(a > bool b) each matched pair becomes1or0. - Logical / set:
and(left series with a right match),or(union of both sides),unless(left series with no right match). Matching is on(bucket, labels); values are carried through, never combined. vector(N): a scalar promoted to a constant, no-label series valuedNat every step bucket (as in... or vector(0)in Prometheus, though theorform itself is not supported yet).label_replace(v, dst, replacement, src, regex): rewrites thedstlabel from a regex capture ofsrc($1/${name}inreplacement). The regex is anchored to the whole source value; a non-match leaves the series unchanged, and an empty result deletes the label. A newdstbecomes an additional series label.- Grouping by a label backed by a column works as expected: the built-in
ones (
service_name,level, ...) and any materialized label (sum by (namespace) (...)groups on itslabel_namespacecolumn; the result series carry the label under its sanitized name). Grouping by any other name is also accepted — it resolves as an attribute, coalesced across containers the same way awhereclause on that name would — and if no record in the queried window actually carries that attribute, every row collapses into one series with that label absent, same as real Loki:sum by (foo) (...)for a label no stream carries returns one series withfooabsent/empty rather than an error, whetherfoois a column we happen to store or an attribute we don't.
Time bucketing — the key approximation¶
Loki evaluates a range aggregation over a sliding [range] window at
each step. SignalDB instead buckets into fixed, step-aligned windows
with date_bin(step, timestamp). This is exact when step equals the
range (Grafana's default for count_over_time panels) and an
approximation otherwise. Set the panel step equal to the range window for
exact results.
Not supported yet¶
These parse but return an "unsupported" error at execution:
- Many-to-one vector matching (
group_left/group_right) and... or vector(0)fallbacks sum without (labels)with a non-empty label list
Examples¶
Range log query for one service, newest first:
curl -G 'http://localhost:3000/loki/api/v1/query_range' \
-H 'Authorization: Bearer sk-...' -H 'X-Tenant-ID: acme' \
--data-urlencode 'query={service_name="api"} |= "error"' \
--data-urlencode 'start=1700000000' --data-urlencode 'end=1700003600' \
--data-urlencode 'limit=100'
Error rate per level over five-minute buckets:
curl -G 'http://localhost:3000/loki/api/v1/query_range' \
-H 'Authorization: Bearer sk-...' -H 'X-Tenant-ID: acme' \
--data-urlencode 'query=sum by (level) (rate({service_name="api"}[5m]))' \
--data-urlencode 'start=1700000000' --data-urlencode 'end=1700003600' \
--data-urlencode 'step=5m'
Errors¶
Failures return a JSON body in the Prometheus/Loki error shape:
{
"status": "error",
"errorType": "bad_data",
"error": "expected matcher operator, found end of input at line 1, column 10"
}
errorType is bad_data (400), not_found (404), rate_limited (429),
timeout (504), unavailable (503, no querier), not_implemented (501),
or internal (500). A query that matches nothing is not an error — it
returns 200 with an empty result array.
A 429 (per-tenant query rate limit exceeded) additionally carries
retryAfterMs in the body and three response headers computed from the
tenant's actual token-bucket state: Retry-After (whole seconds, rounded
up, at least 1), X-RateLimit-Limit (the per-second budget), and
X-RateLimit-Burst (the burst allowance):
{
"status": "error",
"errorType": "rate_limited",
"error": "tenant 'acme' exceeded its query request rate limit; retry after 1s or raise the tenant's limits",
"retryAfterMs": 1000
}
Query-demand statistics¶
Attribute labels used in filters (any label that is not a dedicated
column) are counted as query demand and flushed to the catalog's
advisory attribute_stats table, where they inform which attributes are
worth materializing (see the storage layout docs on materialized labels).
Related¶
- Grafana datasource — connecting Grafana
- Sending OTLP data — ingesting logs
- Authentication — API keys and tenant headers