Tag: kafka
All the articles with the tag “kafka”.
-
Data This Week #29
5 min readDuckDB v2.0's new PEG-based SQL parser, Cassandra's ACCORD consensus protocol for ACID transactions, migrating multilingual full-text search to PostgreSQL, stream deduplication techniques, the hidden joins inside Snowflake MERGE, and the community debate on enterprise data catalog alternatives.
-
Data This Week #20
5 min readKafka's nine-layer architecture breakdown, self-healing pipeline barriers, ClickHouse full-text search on object storage, Google Cloud Next '26 data infrastructure, watsonx.data semantic layer, Neo4j on Databricks, DuckDB internals, and handling messy Excel ETL.
-
Data This Week #15
6 min readSpark Declarative Pipelines for financial lakehouses, ten AWS Glue & Iceberg fixes, MOR as an architectural shift, DuckDB's Quack protocol, SQL fraud patterns, Kafka checkpoint patterns, and the LLM-for-validation debate.
-
Data This Week #14
6 min readFlink CDC streaming ELT from MySQL to Kafka, the LLM engineer's stack map, Ursa's diskless Kafka fork, Iceberg write mechanics, Instacart's billion-product search, Jikkou 1.0, and the AI knowledge-base debate.
-
Data This Week #10
5 min readData product lifecycle, semantic context layer for LLM agents, Netflix's Druid interval caching, Ursa Kafka storage engine, Iceberg v3 VARIANT type, and Ministack vs LocalStack.
-
Data This Week #8
4 min readPydantic for schema contracts, Databricks Vector Search pitfalls, stateless Kafka broker Tansu, Capital One's GenAI agent, RAG as a DE problem, and testing culture in data teams.
-
Data This Week #5
4 min readSpark DAG compilation deep dive, query federation with StarRocks, Pinterest's CDC migration, CyberArk AI with Iceberg, Databricks Zerobus Ingest, and data quality tooling debates.