Track native Variant support across scans, expressions and transport. Completed milestones stay checked; each linked issue owns its implementation and parity tests.
Current direction
#5978 remains the whole-value performance priority. Buffer reuse landed in #6342; removing the extra reconstruction pass awaits apache/arrow-rs#11260. #5519 is in progress in draft #6443. The next expression steps are #5429, then #5424 and #5425. Arrow/DataFusion upgrades and the associated cleanup in #5477 are deferred.
Progress
The merged scan foundation supports canonical and shredded input with spark.sql.variant.allowReadingShredded=true and spark.sql.variant.pushVariantIntoScan=false. #5519 addresses the latter restriction. #5407 retains the original prototype discussion; it was closed without merging.
Shared contract
Every implementation must preserve:
- Spark
VariantType as Struct<value: Binary, metadata: Binary>, in that child order.
ARROW:extension:name=arrow.parquet.variant on the outer Arrow Field, including expression results and FFI schemas. An unmarked lookalike struct is not Variant.
- SQL NULL as a null parent, distinct from a present Variant-null payload.
- Spark results, byte/layout requirements and errors at the boundary being added, with explicit fallback for unsupported consumers and unchanged Spark 3 behavior.
Spark does not define general Variant ordering or hashing; Variant keys are outside this roadmap. Iceberg Variant payloads may coexist with equality deletes on legal primitive keys.
Track native Variant support across scans, expressions and transport. Completed milestones stay checked; each linked issue owns its implementation and parity tests.
Current direction
#5978 remains the whole-value performance priority. Buffer reuse landed in #6342; removing the extra reconstruction pass awaits apache/arrow-rs#11260. #5519 is in progress in draft #6443. The next expression steps are #5429, then #5424 and #5425. Arrow/DataFusion upgrades and the associated cleanup in #5477 are deferred.
Progress
variant_getandtry_variant_getsupport for dynamic paths and nested targets #5426 — dynamic paths and nested targets.parse_jsonandtry_parse_json#5428 — JSON parsing.to_variant_object#5431 — object construction.schema_of_variantandschema_of_variant_agg#5427 — schema inference.variant_explodeandvariant_explode_outer#5432 — explode.variant_getandtry_variant_getsupport for dynamic paths and nested targets #5426.The merged scan foundation supports canonical and shredded input with
spark.sql.variant.allowReadingShredded=trueandspark.sql.variant.pushVariantIntoScan=false. #5519 addresses the latter restriction. #5407 retains the original prototype discussion; it was closed without merging.Shared contract
Every implementation must preserve:
VariantTypeasStruct<value: Binary, metadata: Binary>, in that child order.ARROW:extension:name=arrow.parquet.varianton the outer Arrow Field, including expression results and FFI schemas. An unmarked lookalike struct is not Variant.Spark does not define general Variant ordering or hashing; Variant keys are outside this roadmap. Iceberg Variant payloads may coexist with equality deletes on legal primitive keys.