Type System Utilities¶
Overview¶
The core/types/ module implements Floecat's canonical logical type system. It bridges the logical
types declared in protobuf (types/types.proto) with Java helpers used by connectors, schema
mappers, statistics engines, and the execution scan bundle assembler.
Key classes: LogicalType, LogicalKind, LogicalTypeProtoAdapter, LogicalComparators,
LogicalCoercions, ValueEncoders, and MinMaxCodec.
Canonical Type Kinds¶
LogicalKind defines the complete set of canonical types shared across all table formats (Iceberg,
Delta, etc.) and SQL-facing components.
| Canonical Kind | Proto field number | Description |
|---|---|---|
BOOLEAN |
1 | Boolean true/false |
INT |
2 | 64-bit signed integer (all source sizes collapse) |
FLOAT |
4 | 32-bit IEEE-754 single-precision float |
DOUBLE |
5 | 64-bit IEEE-754 double-precision float |
DATE |
6 | Calendar date (no time, no timezone) |
TIME |
7 | Time of day (no date, no timezone) |
TIMESTAMP |
8 | Timezone-naive timestamp (local time) |
TIMESTAMPTZ |
13 | UTC-normalised timestamp |
STRING |
9 | UTF-8 text |
BINARY |
10 | Arbitrary byte sequence |
UUID |
11 | 128-bit universally unique identifier |
DECIMAL |
12 | Fixed-precision decimal (precision, scale) |
INTERVAL |
14 | Duration / period |
JSON |
15 | Semi-structured JSON text |
ARRAY |
16 | Ordered collection (carries an element type tree) |
MAP |
17 | Key-value map (carries key/value type trees) |
STRUCT |
18 | Named-field record (carries an ordered field list) |
VARIANT |
19 | Schema-flexible semi-structured value |
Integer collapsing¶
Every source-format integer size (TINYINT, SMALLINT, INT, INTEGER, BIGINT, LONG, INT8, INT4, INT2,
UINT8, UINT4, UINT2) collapses to canonical INT (64-bit signed). Source alias names can be
resolved via LogicalKind.fromName(String).
Timestamp semantics¶
TIMESTAMP— stores local time without UTC normalisation (IcebergwithoutZone(), Deltatimestamp_ntz).TIMESTAMPTZ— stores microseconds-since-epoch UTC (IcebergwithZone(), Deltatimestamp).
Note: The Floe spec decode matrix v1 has these two entries inverted. The implementation applies the semantically correct mapping and records the discrepancy in code comments.
Interval ranges¶
INTERVAL remains a single logical kind with an optional range:
- INTERVAL YEAR TO MONTH → range YEAR_TO_MONTH
- INTERVAL DAY TO SECOND → range DAY_TO_SECOND
- Plain INTERVAL → range UNSPECIFIED
Stats encoding (when present) uses ISO‑8601 duration strings. Engine‑native interval layouts are carried via overlays/hints, not Floecat core types.
Interval precisions follow ANSI SQL conventions:
- INTERVAL YEAR(p) TO MONTH → interval_leading_precision = p
- INTERVAL DAY(p) TO SECOND(s) → interval_leading_precision = p,
interval_fractional_precision = s
- INTERVAL(s) → normalized to DAY_TO_SECOND with interval_fractional_precision = s
Leading precision is non‑negative and connector‑defined; fractional precision is limited to 0..6 (microsecond scale) in Floecat encoders.
Precision + parsing¶
- Canonical
TIME,TIMESTAMP, andTIMESTAMPTZare microsecond precision. Inputs with higher precision are truncated to micros for stats encoding/comparison. - Numeric encodings are only accepted for
DATE(epoch days).TIME,TIMESTAMP, andTIMESTAMPTZrequire typed values or ISO‑8601 strings; numeric heuristics are not used. - Temporal precisions can be carried in the logical type string (e.g.
TIME(3),TIMESTAMP(6),TIMESTAMPTZ(0), range 0..6). When present, encoders truncate and emit exactly that many fractional digits. When absent, Floecat defaults to microsecond precision and ISO‑8601 formatting (no fixed width). TIMESTAMPexpects timezone‑naive inputs (noZor offset). By default, zoned strings are rejected. You can opt into conversion by setting:floecat.timestamp_no_tz.policy=CONVERT_TO_SESSION_ZONE(or envFLOECAT_TIMESTAMP_NO_TZ_POLICY)floecat.session.timezone=<IANA zone>(or envFLOECAT_SESSION_TIMEZONE) When enabled, zoned timestamps are converted into that session zone and stored as localTIMESTAMPvalues.
Complex types¶
ARRAY, MAP, and STRUCT carry a recursive type tree on LogicalType: element() +
elementNullable() for ARRAY, key()/value() + valueNullable() for MAP, and ordered
LogicalField entries (name, nullable, type) for STRUCT. VARIANT is self-describing and
never carries a tree. A bare container tag (LogicalType.of(LogicalKind.ARRAY)) remains valid and
represents the legacy non-parameterised form; use hasTypeTree() to distinguish the two.
The string grammar mirrors the tree:
type ::= scalar
| ARRAY "<" type ">"
| STRUCT "<" field ("," field)* ">"
| MAP "<" type "," type ">"
field ::= name ":" type
e.g. ARRAY<INT>, MAP<STRING, DOUBLE>, ARRAY<STRUCT<sku: STRING, quantities: ARRAY<INT>>>.
Struct field names are case-preserved; names that are not simple identifiers are double-quoted
with "" escaping. Nullability is not part of the grammar (parsing defaults to nullable).
On the wire, SchemaColumn.type is the single authoritative semantic type: a recursive
floecat.types.LogicalType message that preserves the full nested shape including
element/value/field nullability, which the string grammar cannot carry. (The former flat
logical_type string field is reserved; persisted columns written before the change are
re-reconciled from upstream metadata where an upstream exists, and view output columns with no
upstream are upgraded transparently on load via
LogicalTypeProtoAdapter.upgradeLegacyColumn, which recovers the legacy string from protobuf
unknown fields.) LogicalTypeProtoAdapter.toProto/fromProto convert
between the JVM model and the wire message; the string grammar remains for textual inputs
(generic schema declarations), docs, and debugging via LogicalTypeFormat.
Nested fields additionally surface as child SchemaColumn rows with their own canonical paths —
struct children as parent.child, list elements as parent[], map keys as parent.key, map
values as parent{} (e.g. address.city, items[].sku, tags{}.v). Schema construction and
the stats-side fieldId→path maps share one traversal, so the schema path set and the stats path
set are identical by construction; the child rows drive per-leaf stats and field-ID mapping.
ScalarStats.logical_type remains a flat canonical string: stats rows exist only for leaves, and
leaf types are never parameterised containers.
Architecture & Responsibilities¶
LogicalType/LogicalKind– Immutable representations of logical types.LogicalTypestores(kind, precision, scale, temporalPrecision, intervalRange, intervalLeadingPrecision, intervalFractionalPrecision)plus the nested type tree for complex kinds (element/elementNullable,key/value/valueNullable,fields).temporalPrecisionis optional (unset means default microsecond precision). Interval fields are optional and only apply toINTERVAL. CanonicalDECIMALsemantics areprecision ≥ 1and0 ≤ scale ≤ precisionwith no global precision ceiling in the core model. Connector-specific constraints apply (for example Iceberg/Delta cap precision at 38, while other sources may allow larger values). TIME/TIMESTAMP/TIMESTAMPTZ may carry a fractional‑second precision (0..6). All other kinds reject parameters.LogicalTypeProtoAdapter– Converts between the protobufai.floedb.floecat.types.LogicalTypewire message and the JVMLogicalType, preserving scalar parameters, recursive shape, and nested requiredness.LogicalCoercions– Coerces raw stat values to the canonical Java type for a given kind (e.g. anyNumber→LongforINT, string →LocalDateTimeforTIMESTAMP(timezone‑naive policy), string →InstantforTIMESTAMPTZ).LogicalComparators– ProvidesComparatorinstances for ordering values encoded as strings or byte buffers (used when building column stats).ValueEncoders/MinMaxCodec– Encode scalar values into canonical strings/bytes, enabling deterministic min/max statistics across connectors.
Type-Family Helpers¶
LogicalType exposes four predicate helpers for grouping kinds:
// Scalar numeric: INT, FLOAT, DOUBLE, DECIMAL
boolean isNumeric()
// Temporal: DATE, TIME, TIMESTAMP, TIMESTAMPTZ, INTERVAL
boolean isTemporal()
// Container: ARRAY, MAP, STRUCT, VARIANT
boolean isComplex()
// Everything that is not a container (alias for !isComplex())
boolean isScalar()
Example:
LogicalType t = LogicalType.decimal(10, 2);
t.isNumeric(); // true
t.isDecimal(); // true
t.isTemporal(); // false
t.isComplex(); // false
t.isScalar(); // true
Public API / Surface Area¶
Most classes expose static helpers:
LogicalType t = LogicalType.decimal(38, 4);
String logicalType = LogicalTypeFormat.format(t);
String encoded = MinMaxCodec.encode(t, BigDecimal.valueOf(42));
LogicalTypeFormat.parse(String) and .format(LogicalType) convert between canonical logical type
strings and runtime objects. LogicalTypeProtoAdapter converts the runtime type to and from the
typed protobuf payload.
Arrow Mapping Contract¶
core/arrow helpers (especially ArrowSchemaUtil) are defined over Floecat logical types, not over
arbitrary engine-native type systems.
- Input is
SchemaColumn.type, a recursivefloecat.types.LogicalTypemessage; decode it withLogicalTypeProtoAdapter.columnType. Textual entry points (generic schema declarations) still accept canonical names and aliases viaLogicalKind.fromNamesemantics. - Integer aliases (
TINYINT,SMALLINT,INT,BIGINT,INT2/4/8,UINT2/4/8) all map to Arrow signed 64-bit (Int64) to preserve collapsed canonicalINTbehavior. - Unknown, null, or blank logical types fail fast with
IllegalArgumentException; they are not silently coerced toUtf8. JSONmaps to ArrowUtf8.UUIDmaps to ArrowFixedSizeBinary(16);BINARYmaps to ArrowBinary.DECIMALmaps to ArrowDecimal128when precision ≤ 38, andDecimal256when precision ≤ 76. Precision > 76 is rejected byArrowSchemaUtil.TIMEmaps to ArrowTime(MICROSECOND, 64),TIMESTAMPtoTimestamp(MICROSECOND, null), andTIMESTAMPTZtoTimestamp(MICROSECOND, "UTC").INTERVALand complex container types (ARRAY,MAP,STRUCT,VARIANT) are not supported in Arrow schema generation; they must be omitted or cast toSTRING/BINARY.
If external Flight providers want to reuse core/arrow, they should first map their source type
surface into Floecat logical types.
Record reflection writer (ArrowRecordWriters)¶
ArrowRecordWriters.fromRecordClass(...) builds an Arrow schema + writer directly from a Java
record (used by Flight system-table producers). Component types map to Arrow as above, with one
extra rule for decimals: a BigDecimal carries no precision/scale, so a BigDecimal component
must be annotated @ArrowDecimal(precision, scale) — an unannotated one is rejected. Precision
and scale are validated through LogicalType.decimal(...) (precision ≥ 1, 0 ≤ scale ≤ precision),
and the column uses Decimal128 for precision ≤ 38 and Decimal256 for 39..76.
Rounding is explicit. @ArrowDecimal(rounding = ...) controls how an input whose scale exceeds the
column scale is coerced; it defaults to RoundingMode.UNNECESSARY, so excess scale throws rather
than silently dropping precision. Padding a value to a larger scale is always lossless. Callers that
intend lossy rounding opt in explicitly, e.g.
@ArrowDecimal(precision = 18, scale = 3, rounding = RoundingMode.HALF_UP).
Source-Format Alias Lookup¶
LogicalKind.fromName(String candidate) resolves source-format type names to canonical kinds:
LogicalKind.fromName("bigint") // → INT
LogicalKind.fromName("float4") // → FLOAT
LogicalKind.fromName("double precision") // → DOUBLE
LogicalKind.fromName("timestamp with time zone") // → TIMESTAMPTZ
LogicalKind.fromName("ARRAY") // → ARRAY
The lookup is case-insensitive and collapses internal whitespace. Unknown names throw
IllegalArgumentException.
Important Internal Details¶
- Validation –
LogicalTypeconstructor enforces: forDECIMAL,precision ≥ 1and0 ≤ scale ≤ precision. There is no global DECIMAL precision cap in the core model; connectors enforce their own ceilings (for example Iceberg/Delta cap at 38) at schema-parse time. Non-decimal kinds reject precision/scale altogether. - Non-stats-orderable types –
INTERVAL,JSON, and complex kinds (ARRAY,MAP,STRUCT,VARIANT) have no meaningful min/max statistics.LogicalComparators.normalize()returnsnullandValueEncoders.encodeToStringthrows for JSON/complex kinds, so connectors should leave bounds unset.INTERVALencodings can be stored but are ignored by stats comparisons. - Comparators –
LogicalComparatorsprovides specialised comparators for lexical ordering of encoded min/max values so histogram builders can operate on encoded strings. - Encoders –
ValueEncodersnormalises values before storing them in stats to guarantee consistent lexical ordering across connectors.
Data Flow & Lifecycle¶
Connector reads Parquet/Delta/Iceberg schema
→ schema mapper emits recursive LogicalType trees (e.g. INT, TIMESTAMPTZ, ARRAY<STRUCT<...>>)
→ SchemaColumn.type stores the typed tree; nested nodes also surface as child rows by path
→ ValueEncoders encode per-column min/max/ndv bounds
→ StatsRepository stores encoded values (string or bytes)
→ Query lifecycle service converts stored logical type IDs into planner TypeSpecs via TypeRegistry
Configuration & Extensibility¶
The module is pure Java; no configuration is required. Extending the type system involves:
1. Adding a new LogicalKind enum entry and proto Kind field number.
2. Registering source-format aliases in the ALIASES map inside LogicalKind.
3. Updating LogicalCoercions, LogicalComparators, and ValueEncoders for the new kind.
4. Updating connector schema mappers (IcebergSchemaMapper, DeltaSchemaMapper) to emit the new
canonical name.
5. Ensuring downstream consumers (FloeTypeMapper, DeltaManifestMaterializer) handle the new
canonical name.
Examples & Scenarios¶
- Iceberg schema parsing –
IcebergTypeMappings.toLogical(Type)converts Iceberg types toLogicalType, recursing into LIST/MAP/STRUCT and preserving element/key/value types and nullability;IcebergSchemaMapperstores the result inSchemaColumn.type. This avoids the historic ambiguity where"timestamp"meant UTC-stored in Delta but non-UTC in Iceberg. IcebergTimestampNanoTypeuses the same timezone semantics (withZone()→TIMESTAMPTZ,withoutZone()→TIMESTAMP), and IcebergVariantTypemaps toVARIANT. - Delta schema parsing –
DeltaSchemaMapper.toLogicalType(DataType)(kernel path) andfallbackLogicalType(JsonNode)(JSON fallback) apply Delta-specific semantics:"timestamp"→TIMESTAMPTZ(UTC-stored),"timestamp_ntz"→TIMESTAMP(timezone-naive). Both recurse into arrays/maps/structs, preservingcontainsNull/valueContainsNull. - Statistics ingestion – NDV providers convert Parquet min/max values using
MinMaxCodecbefore storing them inScalarStats, ensuring planners can compare them without deserialising actual binary payloads. Connector planners canonicalize connector-native numeric temporal bounds to typed values before generic coercion (for example Iceberg TIME micros-of-day, TIMESTAMP micros, and TIMESTAMP_NANO nanos).
Cross-References¶
- Protobuf type definitions:
docs/proto.md - Query lifecycle service:
docs/service.md