Skip to content

Metadata Graph

Overview

Floecat’s query-facing services share a common metadata cache called the Metadata Graph. It sits between the pointer/blob repositories and any RPC that needs to inspect catalogs, namespaces, tables, or views. The graph provides:

  • Immutable node models that can be safely reused across requests. Nodes are assembled through the repository and ObjectCache; serialized inputs are read through BlobCacheAccess when the disk cache is enabled. There is no separate graph-level decoded-node cache.
  • Resource-ID → current-blob resolution served by the owner-managed planner index (PlanningPointerIndex behind IndexedPointerStore), which is refreshed on write and does not expire, so a commit is visible to the replica that made it as soon as it lands.
  • Helper APIs for name resolution (Directory RPC parity) and coherent snapshot selection (Snapshot RPC parity).
  • Extension points (EngineHint) so planners/executors can attach engine-specific payloads without mutating the base metadata structures.

The graph reads through the service-wide caching disciplines described in docs/caching.md.

┌────────────┐      ┌────────────────────┐      ┌───────────────────────┐
│ gRPC RPCs  │ ---> │ MetadataGraph APIs │ ---> │ Repositories / RPCs   │
│ (Query,    │      │  - resolve()       │      │  - Catalog/Table/View │
│  Planner,  │      │  - catalog()/...   │      │  - Directory/Snapshot │
│  Executors)│      │  - resolvedSnapshot│      │  - Storage backends   │
└────────────┘      └────────────────────┘      └───────────────────────┘

Implementation Structure

Immutable node models live under core/metagraph/model, while the runtime helpers and facade sit inside service/metagraph. The split looks like this:

  • core/metagraph/model/ – Immutable node records (CatalogNode, NamespaceNode, TableNode, ViewNode, SystemViewNode) plus shared enums (GraphNodeKind, EngineKey, EngineHint, GraphNodeOrigin, etc.).
  • service/repo/cache/ – the planner index owns complete namespace and relation-name indexes; serialized immutable bodies are handled by BlobCacheAccess, while assembled metadata is held by ObjectCache.
  • service/metagraph/loader/ – NodeLoader wraps the catalog/namespace/table/view repositories to hydrate immutable nodes from protobuf metadata (metaForSafe + pointer fetches).
  • service/metagraph/resolver/ – NameResolver handles catalog/namespace/table/view lookups and lightweight pointer-backed ref listings, while FullyQualifiedResolver mirrors DirectoryService’s ResolveFQ list/prefix semantics.
  • service/metagraph/snapshot/ – SnapshotHelper encapsulates snapshot selection and schema resolution, wrapping the SnapshotService RPC stub.
  • service/cache/HintCache – resolves, caches, and attaches persisted user-relation hints for the exact requested engine. service/metagraph/hint/EngineHintPersistenceImpl is the runtime SPI adapter that writes through this module.
  • service/metagraph/overlay/ – UserGraph (the Metadata Graph façade, see service/metagraph/overlay/user/UserGraph.java) composes the helpers above, exposes the public API, and keeps a CatalogGraphView-friendly view via MetaGraph. SystemGraph (in overlay/systemobjects/SystemGraph.java) consumes SystemNodeRegistry snapshots so pg_catalog-style system tables/views merge with the user metadata when callers go through the composite graph view.

Node Model

Nodes live under core/metagraph/model and each implements GraphNode. They are Java records with defensive copies to guarantee immutability.

  • CatalogNode – Lightweight display + connector/policy metadata. Optionally exposes namespace IDs for listing RPCs.
  • NamespaceNode – Captures catalog ancestry, path segments, display name, and optional child IDs.
  • TableNode – Holds logical schema JSON, partition keys, field-ID map, snapshot references (current, previous, resolved sets), optional stats summary, dependent view IDs, and engine hints. The persisted current snapshot is stored as a dedicated current-snapshot pointer resource.
  • ViewNode – Stores SQL text, dialect, output columns, base relation IDs, creation search path, and optional owner.
  • SystemViewNode – Reserved for virtual/system relations (e.g., $files, $snapshots).

Common fields:

Field Description
id() Stable ResourceId carrying account/kind/UUID.
blobUri() Source blob the node was derived from (user nodes). Identical blob content always produces an identical node, so the URI is the node's content-stable identity.
cacheIdentity() Content-stable cache identity: blobUri for blob-backed user nodes; defaults to String.valueOf(version()) for synthesized system nodes (user nodes do not retain the pointer version and return 0 from version()).
engineHints() Map keyed by EngineHintKey(engineKind, engineVersion, payloadType) → opaque EngineHint.

Engine Hints

EngineHint is a small immutable value with payloadType, opaque bytes, and optional metadata. User-relation hints are persisted in one dedicated resource per (relation, engine kind, engine version) and attached by MetaGraph; system-node hints are materialized directly from builtin definitions. Consumers read the node map and do not know which source supplied it.

Version‑Specificity and Matching Semantics

Engine‑specific hint providers rely on EngineSpecificMatcher, which compares an engine’s (engine_kind, engine_version) against each rule’s declared constraints. Version matching is inclusive: min_version and max_version both participate in ≥ / ≤ comparisons. Versions are compared using natural ordering: numeric segments compare by magnitude, while mixed alphanumeric segments (e.g., 16.0beta2, 16.0rc1) follow prefix ordering where numeric segments always sort after alphabetic suffixes. Pre‑release versions (alpha, beta, rc) therefore sort strictly before the corresponding final release. Rules omitting engine_kind inherit the catalog file’s engine kind. Hint providers compute a fingerprint per node and engine so cached hints remain isolated across engine versions and hint revisions.

Builtin Nodes & Engine Filtering

Builtin SQL objects (types, functions, operators, casts, collations, aggregates) never hit the pointer/blob repositories. Instead, SystemNodeRegistry (core/catalog) loads the pb/pbtxt catalogs once per engine kind, materialises immutable relation nodes, and caches the result per (engine_kind, engine_version). Catalog files live under resources/builtins and follow the <engine_kind>.pb[pbtxt] naming convention. Each builtin definition can declare one or more engine_specific rules (engine kind + min/max versions + optional properties). The registry filters definitions using those rules so a planner that sets x-engine-kind=postgres and x-engine-version=16.0 only sees builtin nodes that actually exist in that release. Callers that omit either header simply receive an empty builtin bundle (the catalog files stay untouched), and GetSystemObjects rejects the request until both headers are provided.

Each engine_specific block may also attach arbitrary key/value properties. When the registry materialises a (engine_kind, engine_version) bundle it keeps only the rules that match the requested engine/version, so the filtered catalog (and GetSystemObjects response) contains exactly the entries that apply to the caller. Pbtxt authors rarely need to repeat the engine kind in every rule; entries that omit it inherit the file’s engine kind. The registry materializes matching rules as immutable node hints under publisher-defined payload types (the payload_type field in the catalog rule), so planners request the hint whose payloadType matches the catalog payload. Documented payload types live alongside the catalog definitions. Catalog authors should prefer stable, namespaced strings (e.g., builtin.systemcatalog.function.semantic+json or floe.type+proto) so that consumers can register decoders per payload family and avoid accidental collisions. Catalog authors control override behavior by ordering engine_specific rules intentionally: entries stay in pbtxt order (including overlay merges), and the mapper treats the last matching rule as the one to publish.

The matcher applies all engine-specific constraints eagerly when materialising builtin bundles. For a given (engine_kind, engine_version) pair, only the rules that match the naturally-ordered version boundaries are retained. The BuiltinNodes returned by SystemNodeRegistry.nodesFor therefore already represent the exact set applicable for that engine release. SystemGraph consumes those nodes to build a _system GraphSnapshot that MetaGraph exposes via CatalogGraphView, so pg_catalog-style system objects live alongside the user metadata when scanners run. SystemObjectsServiceImpl reuses the same SystemNodeRegistry/SystemCatalogProtoMapper pipeline to answer GetSystemObjects() calls without recomputing the catalog data, and because builtin catalogs are immutable per engine version the registry keeps them entirely in memory until FloeCAT restarts.

Deterministic Hint Caching

HintCache owns the only dynamic hint cache. Its key includes account, relation, exact engine identity, and the content-addressed hint blob URI. A pointer move therefore selects a new decoded entry without invalidation, while a warm pointer and body lookup performs no store reads. The resource also records the relation blob URI whose metadata produced the payload; a mismatch is a safe miss and the runtime recomputes the hint. Hints are read only from this resource: relations still carrying engine.hint.* properties from before it existed are a miss, so the runtime re-derives and persists them on first use. New writes never rewrite the relation blob. Builtin hints do not need this cache: SystemNodeRegistry already caches complete immutable engine-version snapshots.

Graph APIs

The UserGraph façade (CDI @ApplicationScoped, see service/metagraph/overlay/user/UserGraph.java) exposes the Metadata Graph APIs that higher layers call. Key methods:

Method Purpose
Optional<GraphNode> resolve(ResourceId) Loads a node from cache or repository by ID/kind.
Optional<CatalogNode> catalog(...) / namespace / table / view Typed convenience wrappers around resolve.
ResourceId resolveName(String cid, NameRef ref) Mirrors DirectoryService semantics for planner RPCs (NameRef → ID).
ResolveResult resolveTables(String cid, List<NameRef> list, int limit, String token) Resolves explicit table names (DirectoryService parity) with best-effort semantics.
ResolveResult resolveTables(String cid, NameRef prefix, int limit, String token) Lists tables under a namespace prefix while enforcing Directory pagination contracts.
ResolveResult resolveViews(String cid, List<NameRef> list, int limit, String token) Resolves explicit view names, returning canonical NameRefs and resource IDs.
ResolveResult resolveViews(String cid, NameRef prefix, int limit, String token) Lists views below a prefix with next-page tokens and total counts.
TablePin resolvedSnapshotFor(String cid, ResourceId tableId, SnapshotRef override, Optional<Timestamp> asOfDefault) Resolves and admits a snapshot selection (override → as-of → current).

Engine Hint Retrieval

All tables and views participating in planning may embed engine‑specific hints. The Metadata Graph attaches only the exact requested engine version through HintCache. Floecat-runtime still owns hint validation, reuse, and computation through EngineMetadataDecorator; it reads the attached maps from relation.node(), computes missing or stale payloads, and persists the result through EngineHintPersistence. Engine versions cannot overwrite one another because they have different pointers and blobs.

Internally resolve(ResourceId):

  1. Reads the pointer through nodes.mutationMeta(id). There is no graph-level meta cache to probe first: the indexed store selects the complete owned partition or its durable fallback.
  2. Loads the immutable serialized body through the repository's BlobCacheAccess seam when disk caching is enabled.
  3. Rehydrates the protobuf record (Catalog, Namespace, Table, View) into the immutable node and stores the assembled relation/schema products in ObjectCache.
  4. A DDL publishes a new pointer and content identity, so the next resolution naturally uses the new object key; no graph-level invalidation is required.

Snapshot Selection Semantics

  • Explicit snapshot ID overrides always win.
  • Explicit AS-OF timestamps are resolved once to a concrete snapshot ID and retain the original timestamp only as provenance.
  • asOfDefault is applied when no overrides exist (to support BEGIN QUERY AS OF ... semantics).
  • Otherwise the graph calls SnapshotService.GetSnapshot(SS_CURRENT) to discover the latest ID.

Name Resolution Semantics

resolveName first short-circuits when the NameRef embeds a ResourceId. Otherwise it performs the catalog, namespace, and relation lookups through lightweight repository refs backed entirely by pointer metadata. It throws the same ambiguity/unresolved error codes as DirectoryService.Resolve*; graph callers get consistent NameRef → ResourceId translations without depending on a secondary RPC hop or hydrating metadata blobs.

Fully Qualified (ResolveFQ*) Semantics

resolveTables/resolveViews mirror the ResolveFQ* RPCs. The helpers accept either a list selector or a namespace prefix, apply input validation, paginate using Directory-compatible tokens, and return canonical NameRef + ResourceId pairs. DirectoryService delegates to these helpers so the graph defines the single source of truth for list/prefix resolution. Both exact selectors and paged prefix selectors return lightweight refs directly from the complete pointer index; full catalog, namespace, table, and view blobs are loaded only by APIs that actually return their contents.

Usage Guidelines

  • Always go through the graph for read paths instead of hitting repositories directly. This keeps cache hit rate predictable and ensures planner/executor code sees immutable snapshots.
  • Nothing to invalidate after a mutation. IndexedPointerStore commits the durable mutation first, then publishes the result while holding the account read gate and the affected key's lock. Different planner keys can proceed in parallel; prefix and account-wide operations take the account write gate. Readers of a complete owned partition see the publication immediately; a handoff or restart rebuilds the partition, and a non-owner falls back to durable KV. Node entries never need eviction — they are content-keyed by blob URI.
  • Treat node instances as read-only. They are immutable records but they may still be shared across requests via the cache, so do not mutate maps or lists after retrieval.
  • Attach engine hints sparingly. Hints should be small (think JSON blobs or compact protobufs) and versioned so planners/executors can safely down-level or up-level between releases.

Query Catalog Service

UserObjectsService.GetUserObjects streams UserObjectsBundleChunks directly from the metadata graph. Each chunk carries a header, batched relation resolutions (RelationResolutions) and a final summary, so planners can start binding as soon as the service resolves each relation. The service shares the same QueryContext as the other query RPCs and relies on CatalogGraphView.resolve, resolvedSnapshotFor, and view metadata stored in ViewNode to produce canonical names, pruned schemas, and view definitions without issuing a second RPC batch. Resolved tables/views also go through QueryInputResolver so their snapshot selections are merged into QueryContext before the response hits the planner—QueryScanService.InitScan can therefore find the same selections later in the lifecycle. Builtins remain behind GetSystemObjects; the information_schema/pg_catalog relations are materialized in the engine-specific overlays for _system scans but do not appear in the RPC response to avoid exposing synthetic tables twice. Column decorations are surfaced per column via RelationInfo.columns[*] (ColumnResult), so a relation can still resolve as FOUND while individual columns report COLUMN_STATUS_FAILED with structured failure reasons.

Metrics

Graph cache metrics are emitted through the shared CacheMetrics helper under the floecat.core.cache.* metric family, distinguished by cache-name tag:

Cache name What is tracked
graph-cache Node-load latency timer and load-failure counter, recorded by UserGraph around each node load.
blob-cache Disk blob-cache hits, misses, mappings, corruption and sweep activity.

Testing

MetadataGraphTest uses in-memory repository/snapshot/directory fakes to exercise cache behavior and helper semantics without Mockito or bytecode agents. Any new helper should be covered there. Higher level components (e.g., QueryInputResolverTest) rely on lightweight graph fakes to validate their own logic while still mirroring real graph responses.

Future Work

  • Add traversal helpers that expand view dependency trees into stable RelationInfo products.
  • Surface resolved snapshot sets on TableNode for multi-table AS OF operations.
  • Provide SPI hooks so connectors can contribute engine hints lazily.