Tag: apache-iceberg
All the articles with the tag “apache-iceberg”.
-
Data This Week #17
4 min readPulsar 5.0 Scalable Topics, Iceberg multi-engine query routing, Netflix's high-throughput graph abstraction, Iceberg partition evolution, AWS Glue automation, and community thoughts on Databricks' BI migration tool.
-
Data This Week #16
5 min readDatabricks as a unified platform, PostgreSQL pluggable storage engines, Iceberg 1.11.0 highlights, database egress costs, JDBC query caching, and dbt column-level lineage.
-
Data This Week #15
6 min readSpark Declarative Pipelines for financial lakehouses, ten AWS Glue & Iceberg fixes, MOR as an architectural shift, DuckDB's Quack protocol, SQL fraud patterns, Kafka checkpoint patterns, and the LLM-for-validation debate.
-
Data This Week #14
6 min readFlink CDC streaming ELT from MySQL to Kafka, the LLM engineer's stack map, Ursa's diskless Kafka fork, Iceberg write mechanics, Instacart's billion-product search, Jikkou 1.0, and the AI knowledge-base debate.
-
Data This Week #12
5 min readCold Postgres data to S3 lakehouse, Databricks Lakeflow Designer, vector databases & HNSW indexing, Salesforce migration best practices, SwiftLake for Iceberg, and data observability lessons.
-
Data This Week #11
5 min readIceberg cross-account migrations, DuckLake 1.0 metadata, IaC for data engineers, Redshift Iceberg writes, agent-data patterns, LARQL for LLM graph queries, and Dagster pricing debate.
-
Data This Week #10
5 min readData product lifecycle, semantic context layer for LLM agents, Netflix's Druid interval caching, Ursa Kafka storage engine, Iceberg v3 VARIANT type, and Ministack vs LocalStack.
-
Data This Week #9
5 min readDuckLake's 926x Iceberg speedup, Expedia's Trino Gateway for workload routing, Ontul unified SQL engine, PostgreSQL memory myths, and a 6-tier FFLIIP streaming lakehouse deep-dive.