Compactor Operations Guide¶
This guide covers day-to-day operations for SignalDB Compactor retention and lifecycle management (retention enforcement, snapshot expiration, and orphan-file cleanup).
Table of Contents¶
- Overview
- Enabling Retention Enforcement
- Enabling Orphan Cleanup
- Monitoring and Metrics
- Common Operations
- Emergency Procedures
- Performance Tuning
- Attribute Promotion
Overview¶
The compactor provides automatic data lifecycle management through:
- Compaction: Merges many small ingest files into fewer large ones
- Retention Enforcement: Drops expired partitions based on configurable policies
- Snapshot Expiration: Maintains bounded metadata by expiring old snapshots
- Orphan Cleanup: Reclaims storage by deleting unreferenced files
Compaction, retention and snapshot expiration are Iceberg metadata commits and respect its transactional guarantees and snapshot isolation. Orphan cleanup is different in kind: it deletes objects from storage outside any commit. It is snapshot-aware (a file is a candidate only when no retained snapshot references it), bounded by a grace period, and re-validated against a freshly rebuilt live set immediately before each deletion batch — but it is not atomic, and a deletion cannot be rolled back by a snapshot.
Within one compactor process, the three committing actors take turns per table: compaction, retention partition drops, and snapshot expiration each acquire that table's lock before doing their work, so they cannot interleave on the same table. Different tables never wait on each other. This covers both ways compaction is triggered — the background cycle and the compact_now Flight action, which an admin can fire at any moment — so a manually triggered compaction cannot land in the middle of a retention pass. It is an in-process ordering only: across multiple compactor instances, safety still rests on Iceberg catalog CAS and on compaction validating that its own input files are still live at commit time. The practical consequence to know about is that a long rewrite delays that table's next retention pass until it finishes; the pass is deferred, never skipped.
Compaction is partition-scoped. A job operates on exactly one closed
timestamp_hourpartition and commits a delta — its input files are removed and the compacted outputs added in a single snapshot, leaving every other partition referenced as it was. Two consequences matter operationally: cost is proportional to the partition being compacted rather than to the table, and concurrent ingest does not invalidate the commit. A partition becomes eligible once its hour has ended and[compactor] partition_lateness(default10m) has elapsed — the partition still receiving writes is the one whose files would change under a running rewrite, so it is deliberately left alone. Rewrites run their DataFusion operators under a[compactor] memory_limit_mbbudget (default 512 MB), spilling to disk past it. The rewrite streams the partition in two passes rather than collecting it, so what sits outside that budget is bounded by one output file rather than by the partition. The rewrite sorts with a fan-out of[compactor] target_partitions(default1): the sorters share the one budget, so raising the fan-out divides it and can exhaust the pool on concurrency alone. It reads the partition in batches of[compactor] scan_batch_sizerows (default1024, below DataFusion's 8192): the sort reserves roughly twice a batch's bytes before it can spill anything, so on a table with wide rows the default row count would put a single unspillable reservation above the whole pool. The compactor warns at startup when these settings cannot work together — a target file size at or above the pool, a per-sorter share too small for a spilling sort to use, or a spill reservation claiming half of that share — since none is visible from any single value.What the rewrite sorts by: the table's own declared sort order (see Storage Layout), not a key list held by the compactor. Its output files record that order, so a partition that held files written before the declaration existed comes out fully attested — compaction is how such files converge, and there is no backfill job. A table that has no declaration yet is still sorted by the canonical key, but its output is written unattested, since there is no declared order for it to claim. See Ordered queries and sort-order attestation for what that means for query behavior and how to turn it off.
Default behavior: The compactor and retention enforcement are enabled by default with
dry_run = falseand a 30-day retention period for traces, logs, metrics, and profiles. A default deployment deletes data older than 30 days. To keep data indefinitely, set[compactor.retention].enabled = false; to keep it longer, raise the per-signal durations. Orphan cleanup is also enabled by default withdry_run = falseand physically reclaims files no retained snapshot references — data Parquet and unreferenced metadata files (old metadata.json versions, expired snapshots' manifests) alike; set[compactor.orphan_cleanup].enabled = falseto opt out ordry_run = trueto observe first.
Enabling Retention Enforcement¶
Step 1: Plan Your Retention Policies¶
Determine appropriate retention periods for each signal type:
| Signal Type | Typical Retention | Production Example |
|---|---|---|
| Traces | 7-30 days | 7 days (dev), 30 days (prod) |
| Logs | 3-14 days | 3 days (dev), 7 days (prod) |
| Metrics | 30-90 days | 30 days (dev), 90 days (prod) |
Consider:
- Regulatory requirements (GDPR, HIPAA, etc.)
- Storage costs vs. query needs
- Incident investigation timeframes
- Audit requirements
Step 2: Configure Retention Policies¶
Edit signaldb.toml:
[compactor.retention]
enabled = true
dry_run = true # Start with dry-run mode
retention_check_interval = "1h" # Check every hour
# Global defaults
traces = "7d"
logs = "30d"
metrics = "90d"
# Safety settings
grace_period = "1h" # Safety margin
timezone = "UTC" # For logging
snapshots_to_keep = 10 # Keep last 10 snapshots
Step 3: Test with Dry-Run Mode¶
Start the compactor with dry-run enabled:
# Start compactor (logs to stdout; redirect to a file if you want to tail it)
cargo run --bin signaldb -- compactor 2>&1 | tee compactor.log
# Or in monolithic mode via the dev script (logs to .data/logs/monolithic.log)
./scripts/run-dev.sh
# Monitor logs. Note: retention logs use "[DRY RUN]" (space), orphan
# cleanup logs use "[DRY-RUN]" (hyphen) — match both:
tail -f .data/logs/monolithic.log | grep -E "DRY.RUN"
Look for log entries like:
INFO compactor::retention::enforcer: [DRY RUN] Would drop expired partitions signaldb.tenant.id=acme signaldb.dataset.id=prod signaldb.table=traces signaldb.job.dry_run=true signaldb.job.partitions_dropped=48 signaldb.job.bytes_reclaimed=1073741824
DEBUG compactor::retention::enforcer: [DRY RUN] Would drop partition tenant_id=acme dataset_id=prod table_name=traces partition_hour=Some("492245") file_count=12 size_bytes=Some(10485760)
The per-partition breakdown is debug-level; run with
RUST_LOG=info,compactor::retention=debug to see it.
Validate:
- Partitions identified for deletion are expected
- Cutoff timestamps are correct
- No unexpected data would be deleted
Step 4: Enable for Test Environment¶
Once dry-run looks good, enable for a test tenant:
[compactor.retention]
enabled = true
dry_run = false # Enable actual deletion
# Use short retention for testing
traces = "1d" # 1 day for fast testing
[compactor.retention.tenant_overrides.test]
traces = "1d"
logs = "1d"
metrics = "1d"
Restart the compactor and verify. Signal data is stored in Iceberg tables on the object store (not in PostgreSQL), so use the compactor's observability endpoint and log output:
# Record the retention counters before the cycle
curl -s localhost:9091/status | jq .retention
# Wait for retention cycle (check interval + processing time)
# Typically 1-5 minutes
# Verify partitions were dropped
curl -s localhost:9091/status | jq .retention
curl -s localhost:9091/metrics | grep compactor_partitions_dropped_total
# Check the drop logs
grep "Dropped expired partitions" .data/logs/monolithic.log
# Confirm old data is gone by querying through the router
# (Tempo search API; expired time ranges should return no results)
curl -s "http://localhost:3000/api/search?start=<old-unix-ts>&end=<old-unix-ts>" \
-H "Authorization: Bearer <api-key>"
Step 5: Rollout to Production¶
After successful test tenant validation:
[compactor.retention]
enabled = true
dry_run = false
retention_check_interval = "1h"
# Production retention periods
traces = "7d"
logs = "30d"
metrics = "90d"
# Production tenant overrides
[compactor.retention.tenant_overrides.production]
traces = "30d" # Keep production traces longer
logs = "7d"
metrics = "90d"
# Critical dataset overrides
[compactor.retention.tenant_overrides.production.dataset_overrides.critical]
traces = "90d" # Critical data kept 90 days
Rollout Checklist:
- [ ] Dry-run validation completed
- [ ] Test tenant validation successful
- [ ] Retention periods reviewed and approved
- [ ] Monitoring and alerts configured
- [ ] Backup/restore procedures verified
- [ ] Stakeholders notified
Enabling Orphan Cleanup¶
Step 1: Identify Orphan Files (Dry-Run)¶
Enable orphan cleanup in dry-run mode:
[compactor.orphan_cleanup]
enabled = true
dry_run = true # Don't delete, just identify
grace_period_hours = 24
cleanup_interval_hours = 24
batch_size = 1000
Start the compactor and monitor its stdout (or .data/logs/monolithic.log when using ./scripts/run-dev.sh):
tail -f .data/logs/monolithic.log | grep -E "(orphan|cleanup)"
Look for:
INFO compactor::orphan::detector: Starting orphan detection tenant_id=acme dataset_id=prod table_name=traces
INFO compactor::orphan::detector: Identified orphan candidates tenant_id=acme dataset_id=prod table_name=traces orphan_candidates=42
INFO compactor::orphan::cleaner: Starting batch deletion of orphan files signaldb.job.candidates=42 signaldb.job.dry_run=true signaldb.job.batch_size=100
INFO compactor::orphan::cleaner: Batch deletion complete signaldb.job.files_deleted=42 signaldb.job.bytes_reclaimed=2147483648 signaldb.job.deletion_failures=0 signaldb.job.dry_run=true
The per-file lines ([DRY-RUN] Would delete orphan file … for each dry-run
candidate, Deleted orphan file … for each successful deletion, with path,
size_bytes, table) are logged at DEBUG — a backlog run can delete tens
of thousands of files in one pass, which would otherwise flood the log at
startup. Failed deletions still log Failed to delete orphan file at
ERROR. Enable the per-file lines with
RUST_LOG=compactor::orphan::cleaner=debug when auditing individual
deletions; the per-batch and per-run summaries above stay at INFO.
Validate:
- Orphan count seems reasonable (expect 0-5% of total files)
- Ages are all beyond grace period (24+ hours)
- No recently modified files flagged
Step 2: Enable Cleanup¶
After validating orphan identification:
[compactor.orphan_cleanup]
enabled = true
dry_run = false # Enable actual deletion
grace_period_hours = 24
cleanup_interval_hours = 24
batch_size = 1000
Restart and monitor:
# Monitor deletion progress
watch -n 5 'curl -s localhost:9091/metrics | grep compactor_files_deleted_total'
# Check logs for errors (stdout, or monolithic.log with run-dev.sh)
tail -f .data/logs/monolithic.log | grep -E "(ERROR|Failed to delete)"
Step 3: Verify Storage Reclamation¶
After cleanup runs:
# Check metrics
curl -s localhost:9091/metrics | grep -E "compactor_(orphan_candidates_identified|files_deleted|bytes_freed)"
# Example output (counters are process-global, no per-tenant labels):
# compactor_orphan_candidates_identified_total 42
# compactor_files_deleted_total 42
# compactor_bytes_freed_total 2147483648
Validation:
compactor_files_deleted_totalshould equalcompactor_orphan_candidates_identified_totalcompactor_bytes_freed_totalshows actual storage reclaimed- No deletion failures (
compactor_deletion_failures_total= 0)
Monitoring and Metrics¶
Key Metrics to Monitor¶
All lifecycle counters are exported at localhost:9091/metrics (see src/compactor/src/http.rs for the authoritative list). Most are process-global, with three exceptions worth knowing:
- the job counters —
compactor_jobs_started_total,compactor_jobs_succeeded_totalandcompactor_jobs_failed_total— carrysignaldb_tenant_id,signaldb_dataset_idandsignaldb_table, and the failure counter additionally carrieserror_type(see Compaction Retries below). Their names and label sets are declared in the SignalDB convention registry,otel/registry/signaldb.yaml; compactor_orphan_cleanup_skipped_total{reason="live_files_threshold_exceeded"};cycle="compaction"|"lease_expiry"|"retention"|"orphan_cleanup"oncompactor_cycle_panics_total/compactor_cycle_down(see Lifecycle Task Recovery below).
Compaction Retries¶
A failed compaction attempt is classified as conflict (a lost
optimistic-concurrency race), transient (object store blip, network hiccup,
catalog contention), or terminal (validation, schema, malformed input).
Conflicts and transient failures are retried with exponential backoff;
terminal failures fail on the first attempt, because retrying repeats the whole
rewrite to reach the same error.
# Retries cover both conflicts and transient infrastructure failures
increase(compactor_retries_attempted_total[1h])
# How much of that is contention specifically
increase(compactor_conflicts_detected_total[1h])
Retries far in excess of conflicts mean infrastructure flakiness rather than
contention. The error_class field on each job's failure log says which class
a given failure was.
Which table is failing. The job counters carry the tenant, dataset and table they acted on, and failures additionally carry the error type — so the first question after "jobs are failing" is answerable from the metric rather than from the logs:
# Failures by table, worst first
topk(5, sum by (signaldb_tenant_id, signaldb_dataset_id, signaldb_table) (
increase(compactor_jobs_failed_total[6h])
))
# Is a table failing occasionally, or every time?
sum by (signaldb_table) (increase(compactor_jobs_failed_total[6h]))
/ sum by (signaldb_table) (increase(compactor_jobs_started_total[6h]))
# Self-resolving contention, or a job that will never succeed as configured?
sum by (signaldb_table, error_type) (increase(compactor_jobs_failed_total[6h]))
An error_type of conflict means another actor committed first and the work
will be retried. Anything else on a table that fails every attempt means the
partition is heading for cooldown and will stop being
attempted at all — that is the shape to alert on. Label names are the
Prometheus rendering of the attributes declared in the SignalDB convention
registry (otel/registry/signaldb.yaml); the label set is bounded, and tables
past that bound are counted under __overflow__ rather than dropped.
Breaking change: these three counters previously had no labels. A query that read
compactor_jobs_failed_totalas a bare scalar must now aggregate:sum(compactor_jobs_failed_total).
Lease Recovery¶
Stale Leases Expired:
A partition whose compactor instance crashed stays unclaimable until its lease is expired. The lease-expiry task sweeps every 30s on its own task, so this counter keeps advancing even while a long compaction cycle is in flight.
# Leases reclaimed from crashed instances (last 24h)
increase(compactor_stale_leases_expired_total[24h])
The counter says a lease reached its expiry without a successful renewal; it
does not say why. Before touching lease_ttl_seconds, rule out the causes in
order of likelihood: instances dying mid-compaction (restarts, OOM kills),
renewal calls failing against the catalog, catalog or network latency
swallowing the ttl / 3 renewal window, and process pauses (long GC-like
stalls, suspended containers). A TTL that is simply too short for your job
durations is the last of these, not the first.
A single renewal miss against a SQLite catalog is not by itself a sign of
trouble: the renewal retries transient SQLITE_BUSY/SQLITE_LOCKED
contention internally and logs at WARN, only escalating to ERROR once
renewal has been failing continuously past the full lease TTL — which is
the point a concurrent instance could actually steal the lease (#1495). Grep
for that ERROR line, not the WARN, when triaging a lease actually lost to
renewal failure.
Lifecycle Task Recovery¶
Each of the four lifecycle tasks (compaction, lease expiry, retention, orphan
cleanup) guards its own iterations against panics via catch_unwind: a panic
is caught, counted, and the task retries on its normal cadence plus a short
exponential backoff, rather than the task ending permanently. This closes the
same class of failure #1011 fixed for slow cycles — a bug in one cycle no
longer takes that cycle down for good.
catch_unwindrequires an unwinding panic strategy. This workspace's[profile.release]setspanic = "abort", so in a release build a lifecycle-cycle panic still aborts the whole compactor process — the guard is exercised bycargo test(defaultunwindstrategy) but is not yet load-bearing in a production binary. If you rely on this recovery behavior, build withpanic = "unwind"instead.
# Panics recovered per cycle (last 24h) — should normally be flat at 0
increase(compactor_cycle_panics_total[24h])
# Is any cycle currently in its post-panic backoff?
compactor_cycle_down
compactor_cycle_down{cycle="..."} is 1 only while that cycle sits in its
post-panic backoff window; it clears as soon as the backoff ends and the
cycle resumes retrying, so it does not stay latched after a one-off failure.
/health deliberately does not reflect this state — it is a pure
liveness probe (200 "ok" whenever the process is serving requests) so that
a cycle recovering on its own backoff schedule never causes a container
orchestrator to restart the process mid-recovery, which would abort the
in-flight retry and turn a bounded backoff into a crash-restart loop. Watch
compactor_cycle_down and /status for the actual per-cycle state instead:
curl -s localhost:9091/status | jq '.lifecycle'
Alerting: any nonzero increase(compactor_cycle_panics_total[1h]) is
worth paging on — a healthy compactor should show zero. A cycle that keeps
reappearing in compactor_cycle_down indicates a persistent bug rather than
a transient failure; check the logs for the cycle's name and panic message
(tracing::error! logs it on every recovery) before assuming a restart will
fix it.
Retention Enforcement¶
Partitions Dropped:
# Partitions dropped (last 24h)
increase(compactor_partitions_dropped_total[24h])
# Partitions evaluated vs dropped
increase(compactor_partitions_evaluated_total[24h])
Retention Duration:
# Wall-clock milliseconds spent enforcing retention per second
rate(compactor_retention_duration_ms_total[5m])
Bytes Reclaimed by Retention:
increase(compactor_bytes_reclaimed_total[24h])
Unclassifiable Files:
Data files whose timestamp_hour partition value could not be determined
from the manifest entry (or the legacy file-path fallback). Such files are
kept and excluded from retention, so a non-zero value means some data is
never expired — investigate the table's manifests.
# Should be 0; alert if it grows
increase(compactor_unclassifiable_files_total[24h])
Compaction Backoff¶
Partitions Skipped While Cooling Down:
When a compaction job fails, the scheduler suppresses that partition for 15 minutes, doubling on each consecutive failure up to a 6-hour ceiling. A success clears the suppression and resets the escalation. Commit conflicts are not failures for this purpose — they mean another actor committed first and the job should be retried, not backed off.
The counter increments once per partition per cycle that is withheld, so it climbs steadily while a partition is stuck rather than reporting a single event.
# Steady growth means some partition cannot be compacted at all
increase(compactor_cooldown_partitions_skipped_total[1h])
A non-zero value is not itself an error: it is the compactor declining to spend capacity on work that just failed. A value that never returns to zero means a partition is permanently stuck — find it in the logs, which name the tenant, dataset, table, partition, and consecutive failure count at each skip:
journalctl -u signaldb-compactor | grep "cooling down"
Orphan Cleanup¶
Storage Reclaimed:
# Total storage freed (last 24h)
increase(compactor_bytes_freed_total[24h])
# Storage reclaimed rate (bytes/second)
rate(compactor_bytes_freed_total[5m])
Cleanup Success Rate:
# Success rate (should be ~100%)
sum(rate(compactor_files_deleted_total[5m]))
/
sum(rate(compactor_orphan_candidates_identified_total[5m]))
Deletion Failures:
# Should be 0 or very low
increase(compactor_deletion_failures_total[1h])
Skipped Cleanups:
# Cleanup runs skipped because the live-file estimate exceeded
# max_live_files_threshold
increase(compactor_orphan_cleanup_skipped_total[24h])
Recommended Alerts¶
Critical Alerts¶
High Deletion Failure Rate:
alert: CompactorHighDeletionFailureRate
expr: |
rate(compactor_deletion_failures_total[5m]) > 0.01
for: 10m
labels:
severity: critical
annotations:
summary: "Compactor orphan deletion failures"
description: "{{ $value }} deletion failures/sec"
Retention Enforcement Stuck:
# No last-run timestamp metric exists; alert on the cutoff-computation
# counter stalling instead (it increments on every retention cycle).
alert: CompactorRetentionStuck
expr: |
increase(compactor_retention_cutoffs_computed_total[2h]) == 0
for: 15m
labels:
severity: critical
annotations:
summary: "Compactor retention hasn't run in 2 hours"
Warning Alerts¶
High Orphan Rate:
alert: CompactorHighOrphanRate
expr: |
increase(compactor_orphan_candidates_identified_total[24h]) > 10000
for: 1h
labels:
severity: warning
annotations:
summary: "Unusually many orphan file candidates"
description: "{{ $value }} orphan candidates identified in 24h"
Orphan Cleanup Skipped:
alert: CompactorOrphanCleanupSkipped
expr: |
increase(compactor_orphan_cleanup_skipped_total[24h]) > 0
for: 1h
labels:
severity: warning
annotations:
summary: "Orphan cleanup skipped (live-file threshold exceeded)"
Grafana Dashboard¶
Example dashboard queries:
Panel: Storage Reclaimed (Bytes)
# Orphan cleanup
increase(compactor_bytes_freed_total[24h])
# Retention enforcement
increase(compactor_bytes_reclaimed_total[24h])
Panel: Partitions Dropped Over Time
rate(compactor_partitions_dropped_total[5m]) * 300
Panel: Retention Duration
rate(compactor_retention_duration_ms_total[5m])
Common Operations¶
Adjusting Retention Periods¶
To change retention for a tenant:
- Update Configuration:
[compactor.retention.tenant_overrides.production]
traces = "14d" # Changed from 30d to 14d
- Reload Configuration:
# Restart compactor (graceful)
pkill -TERM compactor
cargo run --bin signaldb -- compactor
# Or restart monolithic service
systemctl restart signaldb
- Monitor Next Retention Cycle:
# Wait for next retention check (check interval)
# Monitor logs for new cutoff (stdout, or monolithic.log with run-dev.sh).
# The cutoff line is debug-level: run with RUST_LOG=info,compactor::retention=debug
tail -f .data/logs/monolithic.log | grep "Retention cutoff computed"
# Expected log:
# DEBUG compactor::retention::enforcer: Retention cutoff computed tenant_id=production dataset_id=default table_name=traces cutoff_timestamp=2026-01-26 10:00:00 UTC retention_period=14d source=Tenant
Force Immediate Retention Check¶
To trigger retention enforcement immediately (without waiting for interval):
# Option 1: Restart compactor (runs on startup)
systemctl restart signaldb-compactor
# Option 2: Send SIGUSR1 signal (if implemented)
pkill -USR1 compactor
# Option 3: Temporarily reduce interval
# In signaldb.toml:
# retention_check_interval = "1m"
# Then restart
Force Immediate Orphan Cleanup¶
To trigger orphan cleanup immediately:
# Restart compactor (cleanup runs on startup)
systemctl restart signaldb-compactor
# Monitor progress (journalctl for systemd units; stdout otherwise)
journalctl -u signaldb-compactor -f | grep orphan
Verify Retention Cutoff Computation¶
To check what the current retention cutoff would be:
# Enable debug logging (logs go to stdout)
RUST_LOG=debug,compactor::retention=trace cargo run --bin signaldb -- compactor 2>&1 | \
grep "Retention cutoff computed"
# Example output:
# DEBUG compactor::retention::enforcer: Retention cutoff computed tenant_id=acme dataset_id=prod table_name=traces cutoff_timestamp=2026-01-25 09:00:00 UTC retention_period=7d source=Global
Inspect Orphan Candidates¶
To see what files would be identified as orphans (without deleting):
# Enable dry-run mode
# In signaldb.toml:
[compactor.orphan_cleanup]
enabled = true
dry_run = true
# Restart and check logs (stdout, or monolithic.log with run-dev.sh)
tail -f .data/logs/monolithic.log | grep "DRY-RUN.*Would delete"
# Example output:
# INFO compactor::orphan::cleaner: [DRY-RUN] Would delete orphan file path=acme/prod/traces/data/orphan-001.parquet size_bytes=10485760 last_modified=2026-02-04T10:00:00Z table=acme/prod/traces
Check Storage Utilization¶
Signal data lives on the object store, not in the SQL catalog (the catalog's iceberg_tables table only maps table names to metadata locations). Inspect the object store directly:
# Local filesystem storage
du -sh .data/storage
find .data/storage -name "*.parquet" | wc -l
# Per-tenant/dataset/table breakdown (paths are {tenant}/{dataset}/{table}/...)
du -sh .data/storage/*/*/*
# Check object store directly (S3 example)
aws s3 ls s3://signaldb-data/ --recursive --summarize | grep "Total Size"
# Bytes reclaimed by retention and cleanup so far
curl -s localhost:9091/metrics | grep -E "compactor_bytes_(freed|reclaimed)_total"
Emergency Procedures¶
Emergency: Stop All Retention Operations¶
If retention is deleting unexpected data:
# Option 1: Disable in config and restart
# In signaldb.toml:
[compactor.retention]
enabled = false
systemctl restart signaldb-compactor
# Option 2: Stop compactor immediately
systemctl stop signaldb-compactor
# or
pkill -KILL compactor
Emergency: Stop Orphan Cleanup¶
If orphan cleanup is deleting live files (should never happen with proper grace period):
# Disable orphan cleanup
# In signaldb.toml:
[compactor.orphan_cleanup]
enabled = false
systemctl restart signaldb-compactor
Investigate (against compactor stdout, journalctl, or .data/logs/monolithic.log):
# Check revalidation logs
grep -i "revalidation" .data/logs/monolithic.log
# e.g. "File no longer orphan after revalidation, skipping deletion"
# "Revalidation failed, skipping file for safety"
# Verify grace period is filtering recent files
grep "within grace period" .data/logs/monolithic.log
# e.g. "Skipping recent file (within grace period)"
Emergency: Restore Accidentally Deleted Data¶
Retention enforcement drops partitions as Iceberg commits, so the pre-drop snapshot survives until snapshot expiration removes it and orphan cleanup deletes the underlying files. However, SignalDB currently has no supported restore path: snapshot metadata is not queryable via SQL (there is no iceberg_snapshots table in the catalog database), and time-travel queries (FOR SYSTEM_TIME AS OF) are not supported by the query path.
- Stop the Compactor Immediately so snapshot expiration and orphan cleanup cannot delete the files still referenced by the pre-drop snapshot:
systemctl stop signaldb-compactor
- Locate the Pre-Drop Snapshot in the Iceberg Metadata (object store):
# Local filesystem storage: metadata JSON lives next to the table data
ls .data/storage/<tenant>/<dataset>/traces/metadata/
# Inspect the snapshot list (snapshot-id, timestamp-ms, summary)
jq '.snapshots[] | {snapshot_id: ."snapshot-id", timestamp_ms: ."timestamp-ms", summary}' \
.data/storage/<tenant>/<dataset>/traces/metadata/<latest>.metadata.json
- Restore: rolling the table back to the pre-drop snapshot requires manual Iceberg surgery with external Iceberg tooling; it is not currently possible through SignalDB itself.
Prevention:
- Always test retention in dry-run mode first
- Use test tenants before production rollout
- Keep
snapshots_to_keephigh enough for recovery window - Monitor
compactor_partitions_dropped_totalfor unexpected spikes
Performance Tuning¶
Retention Enforcement Performance¶
Symptoms:
- Retention checks taking too long (> 5 minutes)
- High CPU usage during retention cycles
Tuning Options:
[compactor.retention]
# Increase check interval (less frequent = less overhead)
retention_check_interval = "2h"
# Reduce snapshots to keep (less metadata to process)
snapshots_to_keep = 3
Orphan Cleanup Performance¶
Symptoms:
- Cleanup taking hours to complete
- High memory usage during cleanup
- Object store rate limiting
Tuning Options:
[compactor.orphan_cleanup]
# Reduce batch size (less memory, more checkpoints)
batch_size = 500
# Run less frequently (lower peak load)
cleanup_interval_hours = 48 # Every 2 days
# Keep fewer snapshots retained (less history to scan)
[compactor.retention]
snapshots_to_keep = 5
# Bound memory by skipping tables with too many estimated live files
# (skips are counted in compactor_orphan_cleanup_skipped_total)
max_live_files_threshold = 500000
How detection scales¶
Detection cost per table is bounded by three quantities, none of which grows with the number of snapshots that reference the same file:
| Resource | Cost | Notes |
|---|---|---|
| Manifest-list reads | one per retained snapshot | snapshots_to_keep is the lever |
| Manifest file reads | one per distinct manifest | manifests shared by several retained snapshots are deduplicated by path before any of them is fetched, so a manifest referenced by all 14 retained snapshots is read once |
| Manifest entries in memory | one manifest's worth | entries are streamed; only the live path is inspected, never collected |
| Live set in memory | one 64-bit fingerprint per live file (≈16 bytes per hash-table slot) | a 500k-file table costs single-digit MB, not the ~75 MB the same set of path strings would |
| Object-store listing in memory | zero | the data/ (and metadata/) listing is streamed and each entry is decided as it arrives |
| Candidates in memory | one entry per orphan candidate | this, plus the live set, is the working set of a cleanup pass |
Consequences for tuning:
- Peak memory tracks the live file count and the orphan count, not the total object-store listing length and not the snapshot count. A table with tens of thousands of files across a handful of shared manifests is cheap.
- Lowering
snapshots_to_keepreduces manifest-list reads and can shrink the live set (files referenced only by expired snapshots stop being protected), but it does not change the per-file memory cost. max_live_files_thresholdremains the backstop for pathological tables. It is evaluated from manifest-list metadata before any manifest is fetched, so a table over the cap costs one manifest-list read per retained snapshot and nothing else.
Two deliberate non-goals:
- Fingerprints, not paths. The live set stores a 64-bit hash of each live path. A collision can only make an orphan look live — the file is kept, not deleted — so the failure direction is "reclaim later", never "delete a live file".
- One table-scoped listing, not per-partition listings. Orphans can sit
under partition prefixes that table metadata no longer mentions (that is
precisely what makes them orphans), so cleanup lists the whole
data/prefix. Listing per partition prefix would issue more requests for the same objects and would silently skip orphans in dropped partitions.
Concurrent Operation Tuning¶
Symptoms:
- Queries failing during retention operations
- Snapshot conflicts
Tuning Options:
[compactor.retention]
# Keep more snapshots (longer isolation window)
snapshots_to_keep = 10
# Run retention less frequently
retention_check_interval = "2h"
# Stagger operations (retention at 2 AM, cleanup at 3 AM)
Object Store Optimization¶
For S3-compatible stores:
# Increase connection pool size
export AWS_MAX_ATTEMPTS=10
export AWS_RETRY_MODE=adaptive
# Use faster instance types for cleanup
# (more CPU = faster manifest reading)
# Enable S3 Transfer Acceleration
export AWS_S3_USE_ACCELERATE_ENDPOINT=true
Attribute Promotion¶
With [compactor.attr_promotion] enabled and dry_run = false, the compactor promotes frequently queried attributes to typed columns as part of a normal compaction rewrite. Promotion works per (attribute level, key): the same key at resource, scope, and record level is three candidates, each with its own column attr_<level>_<key> typed as the key's canonical type (String → string, Int64 → long, Float64 → double, Bool → boolean). See docs/architecture/storage-layout.md for the naming rule.
A promoted column is a copy. The key's typed map ({container}_str/_int/_double/_bool) stays its one home and keeps every value; the column only duplicates it so filters can use column statistics. Demotion therefore drops the column in a metadata-only commit, loses nothing, and does not change query results.
A (level, key) is promoted when all of these hold for promote_streak consecutive cycles: it has a canonical type at that level, the table has that level's attribute container, the key is not machine-generated-looking or capped by the analyzer, its presence is at least min_presence, it has at least min_query_hits query hits, and it was queried within demote_after_idle. Candidates are ranked by query_hits × presence; at most max_promotions_per_cycle are promoted per cycle, within max_labels_per_table.
Demotion runs before promotion in each cycle:
- A promoted column whose (level, key) was not queried within
demote_after_idleis demoted. - If the table is still over
max_labels_per_table, the least recently queried promoted columns are demoted until it fits.
A pair demoted in a cycle is not re-promoted in the same cycle. The budget counts promoted columns and label_<key> columns together.
Each acted-on promotion makes two commits per table:
- Schema flip (before the rewrite): a metadata-only
AddSchema+SetCurrentSchemacommit adds the promoted columns with fresh Iceberg field ids. No data files change; readers null-fill the new columns until the rewrite lands. - Rewrite/delta commit (the normal compaction commit): every row in the partition being compacted is rewritten with each promoted column filled from its level's typed home. Because compaction is partition-scoped, backfill reaches older rows as their partitions are compacted. A column whose type no longer matches the key's canonical type (the key was repinned) is left null, and the querier ignores it.
A schema-evolution failure is logged as a warning and the compaction continues under the old schema — promotion never fails a rewrite.
Queries stay correct during the transition window: the querier reads coalesce(promoted column, typed home) per level, so a row the rewrite has not backfilled yet still returns its value from the map. The writer does not fill promoted columns.
Legacy label_<key> columns: the compactor no longer adds new ones. Existing ones and [schema.materialized_labels] pins stay; pins are never demoted, and an unpinned label column with no query hits is dropped at rewrite as before.
What operators see in the logs:
Typed attribute promotion decision(info) — per table:dry_run, the (level, key) pairs to promote, and those stillbuildingtheir streak.Typed attribute demotion decision(info) — per table:dry_runand the (level, key) pairs to demote (idle or over budget).Added typed promoted attribute columns via schema evolution/Removed typed promoted attribute columns via schema evolution(info) — a promotion or demotion schema commit landed; lists the table, schema id, and columns.Failed to evolve schema for attribute promotion; continuing compaction without itandFailed to evolve schema for attribute demotion; continuing compaction without it(warn) — the schema commit failed; the rewrite proceeded under the old schema.- The usual
Rewrote table data into compacted filesline covers the backfilled rewrite — there is no separate backfill log line, and no promotion-specific Prometheus metric yet.
See troubleshooting for the skip warnings.
Additional Resources¶
Note: every compaction rewrite also runs a read-only attribute-statistics pass that logs per-key presence, approximate cardinality, and advisory materialization candidates (
Attribute-stats analyzerlog line), and persists the per-key statistics to the service catalog'sattribute_statstable (joined there with query-demand counters flushed by the querier). The same pass also counts presence per (attribute level, key) — resource/scope/record, from the typed layout's four typed homes only, never the off-type residue — into the catalog'sattribute_level_statstable, which, together with per-level query demand recorded by the querier, drives Attribute Promotion; the flatattribute_statsdoes not. The same pass records a bounded per-key value sketch (the most frequent values with their counts, sized byvalue_sketch_size, default 100) intoattribute_value_stats, which is what lets query discovery suggest values without reading data — a key whose distinct values exceed the analyzer's cardinality cap keeps no sketch, so a runaway key is reported as uncovered rather than partially suggested. Each pass replaces a key's sketch wholesale, so suggestions follow the data rather than accumulating values that have stopped occurring. Apart fromvalue_sketch_sizethis pass requires no configuration and changes no table data; the promotion pass built on it is covered in Attribute Promotion.