Amazon S3 Tables added native support for the Apache Iceberg V3 specification, rolling out features across all AWS regions to eliminate costly workarounds for semi-structured and geospatial data. The upgrade introduces deletion vectors, row lineage, and new native data types like variant and nanosecond timestamps to petabyte-scale data lakes.
Upgrading to Format Version 3 Without Rewriting Data
Managing petabyte-scale tables has long forced data engineering teams into architectural compromises. When compliance mandates the deletion of 50,000 user records from a two-billion-row table, older specification iterations write out thousands of individual positional delete files. That overhead slows down subsequent read operations until a background compaction job runs. Amazon S3 Tables now allows operators to execute an in-place migration to the new standard.
Executing an upgrade requires a single command:
ALTER TABLE my_catalog.namespace.existing_table
SET TBLPROPERTIES ('format-version' = '3')
This operation happens atomically without rewriting existing data files. On the next compaction cycle, S3 Tables removes legacy delete files. New modifications deploy deletion vectors automatically. Row lineage fields initialize on the first data mutation following the upgrade. This is a one-way operation. The Apache Iceberg specification does not support downgrading from V3 to V2. Engine compatibility must be verified before making the transition.
Storing Semi-Structured Events With the Variant Data Type
Semi-structured events previously landed in object storage as raw JSON strings that every query engine had to parse on the fly. V3 changes this dynamic by introducing a native variant data type that stores semi-structured data in a columnar format. During writes, the engine shreds variant data into hidden columns and collects statistics.
At query time, those statistics enable file pruning that significantly reduces input-output operations compared to parsing raw JSON strings. A retail analytics team tracking user behavior across web and mobile applications can store different event payloads in a single table without predefined schemas:
CREATE TABLE my_catalog.namespace.clickstream (
event_id bigint,
event_time timestamp,
user_id string,
payload variant
)
USING iceberg
TBLPROPERTIES ('format-version' = '3')
Events with distinct shapes land without triggering schema evolution overhead. Querying the variant column directly occurs without read-time parsing functions. Using Amazon EMR Spark, engineers query specific keys directly:
SELECT
event_id,
user_id,
variant_get(payload, '$.action', 'string') AS action,
variant_get(payload, '$.amount', 'double') AS amount
FROM my_catalog.namespace.clickstream
WHERE variant_get(payload, '$.action', 'string') = 'purchase'
AND variant_get(payload, '$.amount', 'double') > 50.00
Deletion Vectors Replace Traditional Positional Delete Files
Deletion vectors replace traditional positional delete files with a compact binary format. Configuring merge-on-read mode enables this capability for write operations:
ALTER TABLE my_catalog.namespace.clickstream
SET TBLPROPERTIES (
'write.delete.mode' = 'merge-on-read',
'write.update.mode' = 'merge-on-read',
'write.merge.mode' = 'merge-on-read'
)
When running a compliance delete, the engine writes a single deletion vector file instead of rewriting underlying data files. S3 Tables compaction handles these binary files automatically during the next maintenance cycle. Meanwhile, row lineage automatically assigns _row_id and _last_updated_sequence_number to each record. Downstream pipelines query these fields to isolate changed rows without scanning the full table:
SELECT *, _row_id, _last_updated_sequence_number
FROM my_catalog.namespace.clickstream
WHERE _last_updated_sequence_number > 42
AWS Provides Native Apache Iceberg Support
AWS provides native Apache Iceberg support spanning ingestion, storage, cataloging, and analytics. Storage and automatic optimization occur within Amazon S3 Tables. Data ingestion runs through Amazon EMR Spark, management integrates with AWS Glue, and analytics execute via Amazon Redshift. Both S3 Tables and AWS Glue Data Catalog support the Iceberg REST Catalog API, enabling interoperability across different compute engines.
New V3 data types—including variant, geometry, geography, nanosecond timestamps, and unknown—require an engine built on Apache Spark 4.0 or later. This includes AWS Glue 6.0 or later and Amazon EMR release 8.1 or later. These specific data types are supported exclusively for tables using the Parquet file format. Columns of type variant, geometry, geography, or nanosecond timestamp cannot be included in a table’s sort order for compaction, though tables containing these columns still compact under sort and Z-order strategies when the sort order utilizes columns of other types.