Metadata Graph¶
Overview¶
Floecat’s query-facing services share a common metadata cache called the Metadata Graph. It sits between the pointer/blob repositories and any RPC that needs to inspect catalogs, namespaces, tables, or views. The graph provides:
- Immutable node models that can be safely reused across requests. Nodes are assembled through the
repository and
ObjectCache; serialized inputs are read throughBlobCacheAccesswhen the disk cache is enabled. There is no separate graph-level decoded-node cache. - Resource-ID → current-blob resolution served by the owner-managed planner index
(
PlanningPointerIndexbehindIndexedPointerStore), which is refreshed on write and does not expire, so a commit is visible to the replica that made it as soon as it lands. - Helper APIs for name resolution (Directory RPC parity) and coherent snapshot selection (Snapshot RPC parity).
- Extension points (
EngineHint) so planners/executors can attach engine-specific payloads without mutating the base metadata structures.
The graph reads through the service-wide caching disciplines described in
docs/caching.md.
┌────────────┐ ┌────────────────────┐ ┌───────────────────────┐
│ gRPC RPCs │ ---> │ MetadataGraph APIs │ ---> │ Repositories / RPCs │
│ (Query, │ │ - resolve() │ │ - Catalog/Table/View │
│ Planner, │ │ - catalog()/... │ │ - Directory/Snapshot │
│ Executors)│ │ - resolvedSnapshot│ │ - Storage backends │
└────────────┘ └────────────────────┘ └───────────────────────┘
Implementation Structure¶
Immutable node models live under core/metagraph/model, while the runtime helpers and
facade sit inside service/metagraph. The split looks like this:
core/metagraph/model/– Immutable node records (CatalogNode,NamespaceNode,TableNode,ViewNode,SystemViewNode) plus shared enums (GraphNodeKind,EngineKey,EngineHint,GraphNodeOrigin, etc.).service/repo/cache/– the planner index owns complete namespace and relation-name indexes; serialized immutable bodies are handled byBlobCacheAccess, while assembled metadata is held byObjectCache.service/metagraph/loader/–NodeLoaderwraps the catalog/namespace/table/view repositories to hydrate immutable nodes from protobuf metadata (metaForSafe+ pointer fetches).service/metagraph/resolver/–NameResolverhandles catalog/namespace/table/view lookups and lightweight pointer-backed ref listings, whileFullyQualifiedResolvermirrors DirectoryService’s ResolveFQ list/prefix semantics.service/metagraph/snapshot/–SnapshotHelperencapsulates snapshot selection and schema resolution, wrapping the SnapshotService RPC stub.service/cache/HintCache– resolves, caches, and attaches persisted user-relation hints for the exact requested engine.service/metagraph/hint/EngineHintPersistenceImplis the runtime SPI adapter that writes through this module.service/metagraph/overlay/–UserGraph(the Metadata Graph façade, seeservice/metagraph/overlay/user/UserGraph.java) composes the helpers above, exposes the public API, and keeps aCatalogGraphView-friendly view viaMetaGraph.SystemGraph(inoverlay/systemobjects/SystemGraph.java) consumesSystemNodeRegistrysnapshots so pg_catalog-style system tables/views merge with the user metadata when callers go through the composite graph view.
Node Model¶
Nodes live under core/metagraph/model and each implements GraphNode. They are Java records with
defensive copies to guarantee immutability.
CatalogNode– Lightweight display + connector/policy metadata. Optionally exposes namespace IDs for listing RPCs.NamespaceNode– Captures catalog ancestry, path segments, display name, and optional child IDs.TableNode– Holds logical schema JSON, partition keys, field-ID map, snapshot references (current, previous, resolved sets), optional stats summary, dependent view IDs, and engine hints. The persisted current snapshot is stored as a dedicated current-snapshot pointer resource.ViewNode– Stores SQL text, dialect, output columns, base relation IDs, creation search path, and optional owner.SystemViewNode– Reserved for virtual/system relations (e.g.,$files,$snapshots).
Common fields:
| Field | Description |
|---|---|
id() |
Stable ResourceId carrying account/kind/UUID. |
blobUri() |
Source blob the node was derived from (user nodes). Identical blob content always produces an identical node, so the URI is the node's content-stable identity. |
cacheIdentity() |
Content-stable cache identity: blobUri for blob-backed user nodes; defaults to String.valueOf(version()) for synthesized system nodes (user nodes do not retain the pointer version and return 0 from version()). |
engineHints() |
Map keyed by EngineHintKey(engineKind, engineVersion, payloadType) → opaque EngineHint. |
Engine Hints¶
EngineHint is a small immutable value with payloadType, opaque bytes, and optional metadata.
User-relation hints are persisted in one dedicated resource per (relation, engine kind, engine
version) and attached by MetaGraph; system-node hints are materialized directly from builtin
definitions. Consumers read the node map and do not know which source supplied it.
Version‑Specificity and Matching Semantics¶
Engine‑specific hint providers rely on EngineSpecificMatcher, which compares an engine’s
(engine_kind, engine_version) against each rule’s declared constraints. Version matching is
inclusive: min_version and max_version both participate in ≥ / ≤ comparisons. Versions are
compared using natural ordering: numeric segments compare by magnitude, while mixed alphanumeric
segments (e.g., 16.0beta2, 16.0rc1) follow prefix ordering where numeric segments always sort
after alphabetic suffixes. Pre‑release versions (alpha, beta, rc) therefore sort strictly
before the corresponding final release. Rules omitting engine_kind inherit the catalog file’s
engine kind. Hint providers compute a fingerprint per node and engine so cached hints remain
isolated across engine versions and hint revisions.
Builtin Nodes & Engine Filtering¶
Builtin SQL objects (types, functions, operators, casts, collations, aggregates) never hit the
pointer/blob repositories. Instead, SystemNodeRegistry (core/catalog) loads the pb/pbtxt catalogs once
per engine kind, materialises immutable relation nodes, and caches the result per
(engine_kind, engine_version). Catalog files live under resources/builtins and follow the
<engine_kind>.pb[pbtxt] naming convention. Each builtin definition can declare one or more
engine_specific rules (engine kind + min/max versions + optional properties). The registry filters
definitions using those rules so a planner that sets x-engine-kind=postgres and
x-engine-version=16.0 only sees builtin nodes that actually exist in that release. Callers that omit
either header simply receive an empty builtin bundle (the catalog files stay untouched), and
GetSystemObjects rejects the request until both headers are provided.
Each engine_specific block may also attach arbitrary key/value properties. When the registry
materialises a (engine_kind, engine_version) bundle it keeps only the rules that match the requested
engine/version, so the filtered catalog (and GetSystemObjects response) contains exactly the
entries that apply to the caller. Pbtxt authors rarely need to repeat the engine kind in every rule;
entries that omit it inherit the file’s engine kind. The registry materializes matching rules as
immutable node hints under publisher-defined payload types (the payload_type field in the catalog
rule), so planners request the hint whose payloadType matches the catalog payload. Documented
payload types live alongside the catalog definitions. Catalog
authors should prefer stable, namespaced strings (e.g., builtin.systemcatalog.function.semantic+json
or floe.type+proto) so that consumers can register decoders per payload family and avoid accidental
collisions. Catalog authors control override behavior by ordering engine_specific rules intentionally:
entries stay in pbtxt order (including overlay merges), and the mapper treats the last matching
rule as the one to publish.
The matcher applies all engine-specific constraints eagerly when materialising builtin bundles. For a
given (engine_kind, engine_version) pair, only the rules that match the naturally-ordered version
boundaries are retained. The BuiltinNodes returned by SystemNodeRegistry.nodesFor therefore already
represent the exact set applicable for that engine release. SystemGraph consumes those nodes to build
a _system GraphSnapshot that MetaGraph exposes via CatalogGraphView, so pg_catalog-style system
objects live alongside the user metadata when scanners run. SystemObjectsServiceImpl reuses the same
SystemNodeRegistry/SystemCatalogProtoMapper pipeline to answer GetSystemObjects() calls without
recomputing the catalog data, and because builtin catalogs are immutable per engine version the registry
keeps them entirely in memory until FloeCAT restarts.
Deterministic Hint Caching¶
HintCache owns the only dynamic hint cache. Its key includes account, relation, exact engine
identity, and the content-addressed hint blob URI. A pointer move therefore selects a new decoded
entry without invalidation, while a warm pointer and body lookup performs no store reads. The
resource also records the relation blob URI whose metadata produced the payload; a mismatch is a
safe miss and the runtime recomputes the hint. Hints are read only from this resource: relations
still carrying engine.hint.* properties from before it existed are a miss, so the runtime
re-derives and persists them on first use. New writes never rewrite the relation blob. Builtin
hints do not need this cache: SystemNodeRegistry already caches complete immutable
engine-version snapshots.
Graph APIs¶
The UserGraph façade (CDI @ApplicationScoped, see service/metagraph/overlay/user/UserGraph.java)
exposes the Metadata Graph APIs that higher layers call. Key methods:
| Method | Purpose |
|---|---|
Optional<GraphNode> resolve(ResourceId) |
Loads a node from cache or repository by ID/kind. |
Optional<CatalogNode> catalog(...) / namespace / table / view |
Typed convenience wrappers around resolve. |
ResourceId resolveName(String cid, NameRef ref) |
Mirrors DirectoryService semantics for planner RPCs (NameRef → ID). |
ResolveResult resolveTables(String cid, List<NameRef> list, int limit, String token) |
Resolves explicit table names (DirectoryService parity) with best-effort semantics. |
ResolveResult resolveTables(String cid, NameRef prefix, int limit, String token) |
Lists tables under a namespace prefix while enforcing Directory pagination contracts. |
ResolveResult resolveViews(String cid, List<NameRef> list, int limit, String token) |
Resolves explicit view names, returning canonical NameRefs and resource IDs. |
ResolveResult resolveViews(String cid, NameRef prefix, int limit, String token) |
Lists views below a prefix with next-page tokens and total counts. |
TablePin resolvedSnapshotFor(String cid, ResourceId tableId, SnapshotRef override, Optional<Timestamp> asOfDefault) |
Resolves and admits a snapshot selection (override → as-of → current). |
Engine Hint Retrieval¶
All tables and views participating in planning may embed engine‑specific hints. The Metadata Graph
attaches only the exact requested engine version through HintCache. Floecat-runtime still owns
hint validation, reuse, and computation through EngineMetadataDecorator; it reads the attached
maps from relation.node(), computes missing or stale payloads, and persists the result through
EngineHintPersistence. Engine versions cannot overwrite one another because they have different
pointers and blobs.
Internally resolve(ResourceId):
- Reads the pointer through
nodes.mutationMeta(id). There is no graph-level meta cache to probe first: the indexed store selects the complete owned partition or its durable fallback. - Loads the immutable serialized body through the repository's
BlobCacheAccessseam when disk caching is enabled. - Rehydrates the protobuf record (
Catalog,Namespace,Table,View) into the immutable node and stores the assembled relation/schema products inObjectCache. - A DDL publishes a new pointer and content identity, so the next resolution naturally uses the new object key; no graph-level invalidation is required.
Snapshot Selection Semantics¶
- Explicit snapshot ID overrides always win.
- Explicit AS-OF timestamps are resolved once to a concrete snapshot ID and retain the original timestamp only as provenance.
asOfDefaultis applied when no overrides exist (to supportBEGIN QUERY AS OF ...semantics).- Otherwise the graph calls
SnapshotService.GetSnapshot(SS_CURRENT)to discover the latest ID.
Name Resolution Semantics¶
resolveName first short-circuits when the NameRef embeds a ResourceId. Otherwise it performs the
catalog, namespace, and relation lookups through lightweight repository refs backed entirely by
pointer metadata. It throws the same ambiguity/unresolved error codes as
DirectoryService.Resolve*; graph callers get consistent NameRef → ResourceId translations without
depending on a secondary RPC hop or hydrating metadata blobs.
Fully Qualified (ResolveFQ*) Semantics¶
resolveTables/resolveViews mirror the ResolveFQ* RPCs. The helpers accept either a list selector
or a namespace prefix, apply input validation, paginate using Directory-compatible tokens, and
return canonical NameRef + ResourceId pairs. DirectoryService delegates to these helpers so
the graph defines the single source of truth for list/prefix resolution. Both exact selectors and
paged prefix selectors return lightweight refs directly from the complete pointer index; full
catalog, namespace, table, and view blobs are loaded only by APIs that actually return their
contents.
Usage Guidelines¶
- Always go through the graph for read paths instead of hitting repositories directly. This keeps cache hit rate predictable and ensures planner/executor code sees immutable snapshots.
- Nothing to invalidate after a mutation.
IndexedPointerStorecommits the durable mutation first, then publishes the result while holding the account read gate and the affected key's lock. Different planner keys can proceed in parallel; prefix and account-wide operations take the account write gate. Readers of a complete owned partition see the publication immediately; a handoff or restart rebuilds the partition, and a non-owner falls back to durable KV. Node entries never need eviction — they are content-keyed by blob URI. - Treat node instances as read-only. They are immutable records but they may still be shared across requests via the cache, so do not mutate maps or lists after retrieval.
- Attach engine hints sparingly. Hints should be small (think JSON blobs or compact protobufs) and versioned so planners/executors can safely down-level or up-level between releases.
Query Catalog Service¶
UserObjectsService.GetUserObjects streams UserObjectsBundleChunks directly from the metadata
graph. Each chunk carries a header, batched relation resolutions (RelationResolutions) and a final
summary, so planners can start binding as soon as the service resolves each relation. The service
shares the same QueryContext as the other query RPCs and relies on CatalogGraphView.resolve,
resolvedSnapshotFor, and view metadata stored in ViewNode to produce canonical names, pruned schemas,
and view definitions without issuing a second RPC batch.
Resolved tables/views also go through QueryInputResolver so their snapshot selections are merged into
QueryContext before the response hits the planner—QueryScanService.InitScan can therefore find
the same selections later in the lifecycle. Builtins remain behind GetSystemObjects; the information_schema/pg_catalog
relations are materialized in the engine-specific overlays for _system scans but do not appear in the RPC
response to avoid exposing synthetic tables twice.
Column decorations are surfaced per column via RelationInfo.columns[*] (ColumnResult), so a relation can
still resolve as FOUND while individual columns report COLUMN_STATUS_FAILED with structured failure reasons.
Metrics¶
Graph cache metrics are emitted through the shared CacheMetrics helper under the
floecat.core.cache.* metric family, distinguished by cache-name tag:
| Cache name | What is tracked |
|---|---|
graph-cache |
Node-load latency timer and load-failure counter, recorded by UserGraph around each node load. |
blob-cache |
Disk blob-cache hits, misses, mappings, corruption and sweep activity. |
Testing¶
MetadataGraphTest uses in-memory repository/snapshot/directory fakes to exercise cache behavior and
helper semantics without Mockito or bytecode agents. Any new helper should be covered there. Higher
level components (e.g., QueryInputResolverTest) rely on lightweight graph fakes to validate their
own logic while still mirroring real graph responses.
Future Work¶
- Add traversal helpers that expand view dependency trees into stable
RelationInfoproducts. - Surface resolved snapshot sets on
TableNodefor multi-table AS OF operations. - Provide SPI hooks so connectors can contribute engine hints lazily.