KnowBite KnowBite

Data Engineering · · 1 min read

Cost-effective ETL with DuckDB and Amazon S3 Tables on AWS Glue

Learn how to pair DuckDB with AWS Glue 6.0 to run SQL-centric ETL on a single worker, reading Parquet from Amazon S3 and writing Apache Iceberg tables to Amazon S3 Tables. This post walks through a complete, runnable example and compares measured cost and runtime against an equivalent Apache Spark job on the same Glue runtime.

Read original on AWS Big Data Blog