Tag: apache-spark
All the articles with the tag “apache-spark”.
-
Data This Week #30
7 min readAmazon acquires DuckDB Labs, AngelList's self-updating Markdown semantic layer, managing AI coding costs at Databricks, federated MCP access for AI agents, Glue 6.0 real-time mode, Polars 2.0 RC, and community advice on extracting from legacy on-prem databases.
-
Data This Week #24
5 min readDatabricks bakes native AI and semantic primitives into Spark 4.2, Netflix rebuilds LLM serving with vLLM on Triton, multi-cloud lakehouse architecture on AWS for agentic AI, Kubernetes internals deep dive, and Debezium's log-based CDC architecture.
-
Data This Week #19
5 min readSlack's SSH-to-REST EMR migration, clusterless Iceberg Lakehouse with DuckDB, Iceberg v4 metadata proposals, Spark 4.0 on EMR GA, Apache Gravitino unified catalog, and Databricks Omnigent for AI agent orchestration.
-
Data This Week #13
5 min readSpark memory tuning, row-level validation tiers, Postgres RLS pitfalls, Stripe's sharding at 5M QPS, Aurora DSQL vs Postgres, Velero joins CNCF, and SQLGlot 5x faster with mypyc.
-
Data This Week #5
4 min readSpark DAG compilation deep dive, query federation with StarRocks, Pinterest's CDC migration, CyberArk AI with Iceberg, Databricks Zerobus Ingest, and data quality tooling debates.