Operations¶
Docker Compose Modes¶
Docker build, compose modes, and AWS credential options live in docs/docker.md.
Testing¶
# Unit + integration tests across modules (in-memory)
make test
# Run full test suite against LocalStack (fixtures + catalog storage)
make test-localstack
# Only run unit or integration suites
make unit-test
make integration-test
Builtin Catalog Validator¶
Validate bundled or bespoke builtin catalog protobufs without running the service.
mvn -pl tools/builtin-validator package
Engine mode – load via ServiceLoader (extension JAR must be on the classpath):
java -jar tools/builtin-validator/target/builtin-validator.jar \
--engine example
Directory mode – point at a catalog directory; the validator reads _index.txt automatically:
java -jar tools/builtin-validator/target/builtin-validator.jar \
/path/to/builtins/example
File mode – validate a single merged protobuf file (binary or text format):
java -jar tools/builtin-validator/target/builtin-validator.jar \
/path/to/catalog.pbtxt
Flags:
--engine <kind>– load the registeredEngineSystemCatalogExtensionfor the given engine kind viaServiceLoaderinstead of reading a file.--json– emit machine-readable output (for CI or scripting).--strict– fail the run when warnings are present (warnings are currently reserved for future checks).
Observability & Operations¶
Outbound token endpoint allowlist¶
Floecat validates outbound token endpoint hosts before performing client credentials or token exchange flows on behalf of connectors or internal workers. This is an SSRF guard on the shared auth resolution path, not a connector-specific feature.
The relevant settings are:
FLOECAT_SECURITY_ALLOWED_TOKEN_ENDPOINT_DOMAINS– comma-separated host/domain allowlist for outbound token endpoints. Exact hosts match exactly;*.example.commatches subdomains only.*allows any token endpoint host.FLOECAT_SECURITY_ALLOW_PRIVATE_TOKEN_ENDPOINTS_FOR_ALLOWED_HOSTS=true– permits allowlisted HTTPS hosts that resolve to private or loopback addresses.FLOECAT_SECURITY_ALLOW_LOOPBACK_TOKEN_ENDPOINTS=true– permits loopback-only HTTP token endpoints for local development, on the shared connector auth-resolution path.FLOECAT_SECURITY_ALLOW_LOOPBACK_CATALOG_ENDPOINTS=true– the equivalent for a Unity Catalog Integration. Its provider holds its own token endpoint to the catalog setting rather than the token one, because whentoken_uriis omitted the endpoint is derived fromcatalog_uriand inherits its scheme: gating the two differently would let a cleartext token endpoint outlive the cleartext catalog it came from. A Unity Integration against anhttp://localhostcatalog therefore needs this variable, not the token one.
Important behavior:
- These settings restrict token endpoint hosts, not upstream catalog hosts in general.
FLOECAT_SECURITY_ALLOWED_TOKEN_ENDPOINT_DOMAINS=*only bypasses the host allowlist check. It does not disable the private-address or loopback HTTP guards.- The same shared validation applies to Delta/Unity, Iceberg REST, and any other connector auth flow that performs service-side token acquisition.
- Where the allowlist is set, it is also applied when a Catalog Integration's OAuth
client-credentials
token_uriis created or updated, for every integration type. An integration whosetoken_uriis outside the list cannot be written until the list covers it. This is a write-path check only: records persisted before the check existed are not re-validated, and the endpoint derived fromcatalog_uriwhentoken_uriis omitted is not covered.
Cleartext S3 endpoints¶
A Unity Catalog or Delta Sharing Integration may set s3.endpoint to reach an S3-compatible store.
Floecat requires HTTPS there by default, because a storage vend on either is only published when it
carries an AWS session token: that token travels in the X-Amz-Security-Token header on every signed
request, and anyone who observes it can replay it against the table's storage prefix until it
expires. The value is also republished to reconcile and query workers, so the exposure is not limited
to validation.
FLOECAT_SECURITY_ALLOW_CLEARTEXT_S3_ENDPOINTS=true– permits anhttp://s3.endpoint. Set this only where the network between Floecat, its workers, and the object store is trusted; MinIO and LocalStack deployments are the usual reason. The bundled Docker Compose sets it because its S3 is LocalStack over HTTP.
The address-class guard applies to s3.endpoint under either scheme, and it is a separate control
from the one above. Link-local, wildcard and multicast literals -- including 169.254.169.254 -- are
always refused. A private address literal (10/8, 172.16/12, 192.168/16, IPv6 unique-local) is
also refused unless floecat.security.allow-private-catalog-endpoints=true (or
FLOECAT_SECURITY_ALLOW_PRIVATE_CATALOG_ENDPOINTS=true) -- the same property that governs a Unity
catalog URI, described in Delta connectors. Permitting cleartext does not
permit a private address, so a MinIO or LocalStack endpoint written as an IP needs both settings.
As with the catalog URI, these checks apply only to address literals: a hostname is never resolved
during validation, so an endpoint written as a hostname passes the address guard regardless of what
it resolves to. That is why the bundled Compose stack, which addresses LocalStack as localstack,
needs only the cleartext setting.
Delta Sharing access modes¶
A Delta Sharing table is readable through this integration only where its provider offers directory
access. A table that offers url access alone returns presigned per-file URLs rather than
credentials, which the storage contract cannot represent, and the vend refuses it by name.
A table that states no access modes at all is asked rather than refused. The protocol reads an absent
field as url only, but the reference server implements the temporary-table-credentials endpoint
while never sending the field, so holding to that reading refuses every table on a server that would
have answered. Floecat attempts the vend and reports what the server says: a refusal or a missing
endpoint becomes CIVI_CREDENTIAL_VENDING_UNSUPPORTED on validation, while a rejected token or an
unreachable server is reported as itself.
A table this Integration cannot read is refused when the overlay reconciles, and counted in
objects_skipped, rather than materialized and left to fail at query time. That covers a table
offering url access alone; a table stating no access modes under
delta.sharing.strict-access-modes; a table reporting auxiliaryLocations on either surface, since
the storage contract carries one credential over one scope prefix and vending only the root would
fail at scan time on the files it does not reach; a table whose location is on a cloud this
provider cannot vend for, meaning abfss://, gs://, or an s3:// whose bucket cannot be read;
and a table whose location no surface states, where the credential endpoint then refuses with a 404
or the protocol's own 400.
Two shapes nearby are not skips. Where the credential endpoint answers 200 without a location, or
with a location the stated one is not under, the result is a classified failure rather than a
skipped table -- a response this client cannot read is not a property of one table. And where the
storage probe finds no Delta log object under the location, validation reports
CIVI_STORAGE_ACCESS_FAILED: that means the location is not a table root -- a wrong region, a wrong
prefix, a broadly scoped credential -- which is correctable, unlike a capability the provider does
not have. The reason for refusing the rest is
that a table reconciled without a storage location keeps the Integration's own catalog URI as its
upstream reference, so it exists in the catalog and cannot be opened by anything reading object
storage. A share offering only url access is therefore unsupported here, not partially supported.
A table's access modes are reported by integration validate, which names the table and the reason
when it refuses one. They are not stored on the reconciled table: the overlay reconciler builds a
table's properties from the Integration and Overlay identities and the storage location, and does
not carry the upstream properties a provider reports. That is true of every provider, not only this
one, so reading a share's access modes means validating the Integration rather than describing the
materialized table.
delta.sharing.strict-access-modes=true-- a connection property that restores the protocol reading, refusing a table that states no modes without asking. Set it against a server known to advertise its modes correctly, where a doomed request per table is waste. It carries a cost on the vend, and against a large schema the cost is the wrong way round. Where the metaData action states no modes, strict mode consults the table listing before refusing, because the listing is where the protocol definesaccessModesand the metaData spelling is the newer one -- refusing without looking there would refuse a conforming table. A credential is vended on every storage-authority resolve with a client that lives for that one resolve, so that fallback pages the whole schema listing once per read. It is bounded byfloecat.delta-sharing.max-listing-bytesrather than unbounded, but it is a listing per read. Leave this off for a share whose schemas hold many tables, or set it only against a server that states modes on the metaData action, where the fallback never runs.delta.sharing.reader-features-- a comma-separated list of Delta reader features this deployment can process, empty by default. The capability header tells the server what the client can handle, and claiming a feature the read path cannot process turns a refusal the server would have made into a failure partway through a scan. Enforcement is the server's: the list is sent on the metadata call and is not compared against the features the server reports back, so a server that ignores the header is not caught here.floecat.delta-sharing.max-pages(system property) -- maximum pages fetched by one listing, default 10,000. A repeated page token is refused outright; this bound is what stops a server minting a fresh one forever. Exceeding it is reported asINVALID_RESPONSE.floecat.delta-sharing.max-response-bytes(system property) -- cap on a single response body, default 32 MiB. A larger body is refused rather than buffered, since the endpoint is named by tenant configuration.floecat.delta-sharing.max-listing-bytes(system property) -- cap on everything one client lists, across pages and across listings, default 32 MiB. The other two bounds do not compose into a memory bound: ten thousand pages of thirty-two mebibytes is past any heap, and a reconcile pass keeps each schema's tables for the whole of its life, so a per-listing cap would not bound the pass. A table entry runs roughly three hundred bytes of JSON, nearer six hundred with a deep prefix, so the default admits above sixty thousand tables across a whole share where a large real share holds thousands, and the decoded records cost about as much again. Raise it for a genuinely larger share; exceeding it fails the reconcile asINVALID_RESPONSE, which is not retried and ends the pass rather than being recorded as one skipped schema.
Reconciler deployment modes¶
The reconciler runs in three shapes from the same artifact:
- All-in-one: default profile; public APIs, durable queue ownership, and local executor polling stay in one runtime.
- Control plane:
QUARKUS_PROFILE=reconciler-control; owns the queue, automatic enqueue, public reconcile APIs, and executor-control RPCs. - Executor plane:
QUARKUS_PROFILE=reconciler-executor; disables local queue ownership and automatic scheduling, then leases work remotely from the control plane over gRPC.
The durable queue is intentionally split into domains:
- canonical job-index state
- ready-queue state
- lease-coordination state
- canonical payload-artifact references on job rows
- projection/root-summary observability state
The control plane owns canonical job-state transitions and the derived job-index plus ready-queue mutations that move with them transactionally. Executors participate through the separate lease-coordination domain when they lease, renew, cancel, and complete work. Projection/root summary maintenance is best-effort observability only and does not participate in queue correctness.
Key reconciler mode flags live in service/src/main/resources/application.properties:
floecat.reconciler.job-queue.enabled
floecat.reconciler.worker.mode
floecat.reconciler.worker-affinity
reconciler.max-parallelism
floecat.reconciler.executor.remote-planner.enabled
floecat.reconciler.executor.remote-default.enabled
floecat.reconciler.executor.remote-snapshot-planner.enabled
floecat.reconciler.executor.remote-file-group.enabled
floecat.reconciler.executor.remote-snapshot-finalize.enabled
floecat.reconciler.executor.snapshot-finalize.enabled
floecat.reconciler.snapshot-finalize-publication.tick-every
floecat.reconciler.snapshot-finalize-publication.page-size
floecat.reconciler.snapshot-finalize-publication.max-parallelism
floecat.reconciler.authorization.header
floecat.reconciler.oidc.issuer
floecat.reconciler.oidc.client-id
floecat.reconciler.oidc.client-secret
floecat.reconciler.oidc.token-refresh-skew-seconds
floecat.reconciler.oidc.connect-timeout
floecat.reconciler.auto.execution-class
floecat.reconciler.execution-lane
floecat.reconciler.job-queue.enabled=false is the emergency admission and execution kill switch.
Set it with FLOECAT_RECONCILER_JOB_QUEUE_ENABLED=false and restart every control-plane and
executor instance. Disabled instances do not admit captures, run scheduled queue work, or serve
executor protocol calls. Existing jobs remain persisted; Get/List job inspection, finalized
snapshot status, cancellation, reconciler settings, and account-deletion cleanup remain available.
Async stats capture returns an uncapturable queue_disabled result and synchronous stats capture
returns FAILED. Re-enable the setting and restart to resume eligible persisted jobs. Queue
maintenance and reconcile-job GC are paused while the switch is disabled.
Recommended split deployment:
- Control plane:
QUARKUS_PROFILE=reconciler-control - Executor plane:
QUARKUS_PROFILE=reconciler-executor - Shared settings: same blob/kv backend, same reconciler OIDC worker principal configuration, executor nodes pointed at the control-plane gRPC host/port
- Control-plane-specific setting:
reconciler.max-parallelism=0 - Executor-plane-specific setting:
floecat.reconciler.worker.mode=remote
Reconciler versioned rollout¶
Set floecat.reconciler.worker-affinity to the job-tree contract version, currently
reconciler-v1, on both the control plane and every executor routed to it. Lease requests require
an exact match; blank or mismatched workers are rejected. Different affinities can share the same
durable blob/kv backend because filtered ready indexes, including pinned-executor indexes, are
partitioned by affinity, but their worker-control endpoints must remain separately routed.
For an upgrade from an unversioned deployment, keep the old control plane and workers running to drain their unversioned job trees. Start the new control plane and workers with the new matching affinity, route each worker fleet only to its matching control plane, and retire the old deployment after its active trees have drained. A new deployment does not adopt, migrate, or lease the old cohort's jobs.
Worker gRPC auth boundary:
- Remote reconcile workers authenticate to
ReconcileExecutorControlwith an explicit bearer token attached by the worker client itself. - In OIDC mode, that bearer token comes from the configured reconciler worker service principal.
- Internal user-context fanout still uses propagated request metadata where appropriate, but that is separate from reconcile worker control-plane authentication.
In the split model, the control plane owns top-level PLAN_CONNECTOR jobs and public reconcile
APIs, while executor-plane nodes primarily run child PLAN_TABLE, PLAN_VIEW, PLAN_SNAPSHOT,
EXEC_FILE_GROUP, and FINALIZE_SNAPSHOT_CAPTURE work. CaptureNow uses the same plan-plus-child
execution path. File-group workers submit one immutable artifact-bundle descriptor through
CommitLeasedFileGroupResult, which requires result_id so the control plane can enforce replay
safety across worker retries. Success durably completes the child job, then protects the bundle,
stages its stats/index target mappings without reading the bundle, and writes the prepared marker.
The RPC reports acceptance after staging completes; exact replay resumes incomplete staging.
For floecat.kv=dynamodb, the durable reconcile hot paths now use native queue-oriented storage
layouts rather than broad generic prefix scans:
- job-index queries are partitioned by their query slice
- ready rows are stored in due-ordered ready slices
- lease rows and lease-expiry scans are stored in dedicated lease partitions
- projection state is separate from queue ownership and is not part of lease/read repair
If PLAN_CONNECTOR jobs can be enqueued, at least one enabled executor must support that job kind.
Planner executors lease only the configured execution lane, and child planning, file-group, and
finalizer jobs preserve that lane. The snapshot-finalizer publication scheduler also requires an
exact lane match before publishing an accepted result. A blank configured lane matches only
unlabelled jobs; it is not a wildcard. The * token is reserved for internal queue scans and is
rejected in deployment configuration and job execution policies. The authenticated worker-control
protocol may carry * for an intentional internal scan, but normal executor polling substitutes
the deployment's configured lane before making that request.
Before renaming a lane or decommissioning its last Floecat instance, stop enqueueing new work on that lane and allow its accepted snapshot-finalizer intents to publish. Publication has no cross-lane fallback: an accepted intent remains owned by the exact lane recorded on its job.
To scale executors horizontally, add more executor-plane instances. They greedily lease eligible jobs from the shared durable queue, so no leader election is required at the executor layer.
- Logging – JSON console logs plus rotating files under
log/. Audit logs route gRPC request summaries tolog/audit.json; seedocs/log.md. - Metrics – Micrometer/Prometheus exporters expose gRPC, storage, and GC metrics at the
/q/metricsendpoint (see the telemetry hub contract indocs/telemetry/contract.md). - Tracing – OpenTelemetry (TraceContext propagator) is always enabled for MDC correlation
(
traceId/spanId). The OTLP exporter is built-in; activate thetelemetry-otlpprofile to ship spans to a collector. Seetelemetry-demo.mdfor the full Prometheus + Tempo + Loki + Grafana demo stack.
Account lifecycle¶
The default Floecat process serves every account. Deployments may bind a different
AccountScope/ownership implementation when they need account admission or lease coordination.
Query contexts and object caches are local optimizations; retention-based snapshot visibility and
account-scoped GC determine whether durable data can be used or collected. A pod can therefore be
restarted without a Floecat drain protocol: Core stops routing new work to the old pod and retries
queries whose local context was lost.
Telemetry hub configuration¶
The service uses the telemetry hub core + Micrometer backend. The following flags are available in
service/src/main/resources/application.properties (the telemetry-otlp profile toggles OTLP tracing/log exports):
telemetry.strict=false
telemetry.exporters=prometheus
telemetry.contract.version=v1
%dev.telemetry.strict=true
%test.telemetry.strict=true
%telemetry-otlp.telemetry.exporters=prometheus,otlp
These settings keep production lenient (dropped-tag counters are exposed) while blowing up early in dev/test when metrics violate their contract. Run the docgen tool if you change any metrics to regenerate the catalog:
mvn -pl telemetry-hub/tool-docgen -am process-classes
Documented metrics are listed in docs/telemetry/contract.md and generated JSON docs/telemetry/contract.json.
Configuration flags are documented per module (for example storage backend selection in
docs/storage-spi.md, GC cadence in docs/service.md,
Secrets Manager in docs/secrets-manager.md).
Telemetry exporter matrix¶
The telemetry.exporters flag (defined in service/src/main/resources/application.properties) tells the hub which backends to activate. Supported values are:
| Exporter | Description | Activation | Notes |
|---|---|---|---|
prometheus |
Micrometer Prometheus registry that exposes floecat.core.* and floecat.service.* metrics via /q/metrics. |
Enabled by default and controlled on the Micrometer side via quarkus.micrometer.export.prometheus.enabled=true. |
This exporter simply scrapes the Micrometer registry that Observability feeds. Keep telemetry.exporters set to include prometheus in dev/test to keep dashboards working. |
otlp |
OpenTelemetry exporter that forwards traces (and hub metrics if wired) to an OTLP collector via gRPC. | Enabled only with the telemetry-otlp profile (e.g., QUARKUS_PROFILE=telemetry-otlp mvn -pl service -am quarkus:dev). |
Configure the OTLP endpoint with quarkus.otel.exporter.otlp.endpoint (runtime property). The exporter itself is cdi (Quarkus built-in, build-time); do not set quarkus.otel.traces.exporter=otlp. Run the full telemetry demo stack (examples/telemetry/docker-compose.yml) to bring up the collector, Tempo, Loki, and Grafana. |
Drop an exporter by removing it from telemetry.exporters (or setting the property to the empty string), e.g., %dev.telemetry.exporters=prometheus keeps strict mode local without OTLP traffic. The hub simply skips wiring backends it isn’t asked for, so only the listed exporters get the meters/traces.
Spans emitted by the service carry custom attributes floecat.component and floecat.operation (set by GrpcTelemetryServerInterceptor), matching the component/operation labels on Prometheus metrics. The Grafana dashboard uses this bridge to build TraceQL links: clicking a series on an RPC panel opens Tempo Explore with span.floecat.operation = "<operation>". Standard OTel rpc.* attributes are also present on every gRPC span for ad-hoc queries. Logs expose the same values through floecat_component and floecat_operation MDC keys, plus traceId and spanId for trace ↔ log correlation. The Tempo datasource is provisioned with tracesToLogsV2 so you can click through from a trace span to the correlated Loki logs.
Metrics¶
Micrometer + Prometheus export is enabled by default. The scrape endpoint is:
GET http://<host>:<http-port>/q/metrics
Physical metadata-store reads use the existing floecat.core.store.* family with
component="pointer_store" or component="blob_store". For example, request rate and batch
fan-out can be inspected with:
sum by (component, operation) (rate(floecat_core_store_requests_total{component=~"pointer_store|blob_store"}[5m]))
sum by (component, operation) (rate(floecat_core_store_items_total{component=~"pointer_store|blob_store"}[5m]))
/ sum by (component, operation) (rate(floecat_core_store_requests_total{component=~"pointer_store|blob_store"}[5m]))
The first query is physical call rate; the second is addressed items per call. Store identifiers are intentionally absent from labels to keep cardinality bounded.
See docs/telemetry/overview.md for the naming, tagging, and contribution rules, and view the generated catalog (docs/telemetry/contract.md and docs/telemetry/contract.json) for the current set of metrics. Regenerate the catalog any time you add or modify a metric:
mvn -pl telemetry-hub/tool-docgen -am process-classes