Designing the S3 raw layout

The S3 layout you pick on day one becomes the contract every downstream reader depends on. Get it wrong and rewriting the layout means rewriting every Glue catalog table, every Snowflake external volume, and every Airflow DAG that points at it. Get it right and you forget about it.

Three S3 zones in the lakehouse

Raw is what came out of the source. Staging is post-extract, pre-Iceberg. Warehouse is the Iceberg tables Snowflake reads.

A working naming convention for raw: s3://bucket/project_name/raw/<entity>/year=YYYY/month=MM/day=DD/batch_ID.json. Hive-style partition keys keep Athena and Glue crawlers happy. Embedding the batch ID makes reruns idempotent.

Keep raw long enough that you can rebuild the warehouse from it. 30 to 90 days is a common window. Beyond that, S3 lifecycle rules transition to Glacier or expire. The point of raw is reproducibility, not permanent storage. Iceberg holds the curated truth.

Quiz: Quiz

Loading practice…