Designing the S3 raw layout
The S3 layout you pick on day one becomes the contract every downstream reader depends on. Get it wrong and rewriting the layout means rewriting every Glue catalog table, every Snowflake external volume, and every Airflow DAG that points at it. Get it right and you forget about it.
Three S3 zones in the lakehouse
A working naming convention for raw: s3://bucket/project_name/raw/<entity>/year=YYYY/month=MM/day=DD/batch_ID.json. Hive-style partition keys keep Athena and Glue crawlers happy. Embedding the batch ID makes reruns idempotent.
Keep raw long enough that you can rebuild the warehouse from it. 30 to 90 days is a common window. Beyond that, S3 lifecycle rules transition to Glacier or expire. The point of raw is reproducibility, not permanent storage. Iceberg holds the curated truth.
Quiz: Quiz
Loading practice…