TopicPartition
Podcast Description
A in-depth engineering podcast about Apache Kafka
Podcast Insights
Content Themes
Covers advanced topics related to Apache Kafka with a special emphasis on architectural innovations like KIP-1150, as well as discussions on data replication, batch coordination, and latency requirements. Episodes delve into specific concepts like diskless topics, caching strategies, and the implications of open-source contributions.

A in-depth engineering podcast about Apache Kafka
In this episode, we talk with ”Jordan has no life” about a variety of timely data engineering including:- the industry standardizing on columnar processing (eg Arrow) and Rust-based OSS engines- new file format proliferation (Vortex, Lance, Nimble, F3)- object storage as the backbone for modern data systems- the open core business model (e.g the Tabular acquisition and Iceberg's PMC dominance)Jordan is the creator of the popular YouTube channel 'Jordan Has No Life,' ex-Google data-pipelines engineer, now on a data team at a high-frequency-trading firm.We also touched on all the second-order effects from AI coding revolution:- English becoming the working abstraction for code- human reviewers becoming the bottleneck (esp in OSS)- focus on benchmarks, tests and design more than implementation- testing practices for the AI era – triaging flaky tests with AI, fuzz/chaos testing a-la TigerBeetleLast but not least, we touched on Jordan's two latest projects he's been working on: StreamFusion and IceStreamStreamFusion accelerates Apache Flink by rewriting Flink SQL operators in Rust and offloading columnar SQL-plan nodes to embedded DataFusion instances running on each task manager- StreamFusion: https://github.com/datafusion-contrib/StreamFusionIceStream addresses the Iceberg upsert problem by doing change-data-capture on an Iceberg table, maintaining a secondary index in Apache Paimon and asynchronously converting write-optimized equality deletes into read-optimized positional deletes- IceStream: https://github.com/jordepic/icestream——————————————————————–*TIMELINE*0:00 Opening0:17 the future of the Data Ecosystem: including new file formats, columnar batching & object storage8:00 Jordan's Background (HFT trading firm, Google)13:21 OSS vs proprietary systems20:09 Jordan's Personal Projects22:50 the future of engineering \w AI29:03 the asymmetry of code review & testing (\w AI again)33:49 jordan's StreamFusion project – accelerating Flink in Rust with DataFusion and Arrow50:16 jordan's IceStream project and the Iceberg Upsert Problem57:52 table formats (Apache Paimon & more)1:04:27 Iceberg examples in prod and the need for real-time streaming to it—————————————–*JORDAN*Jordan is the creator of the YouTube channel 'Jordan Has No Life' (@jordanhasnolife5163), which has several hundred videos on distributed systems and system-design interviews. He has recently pivoted toward more hands-on data-infrastructure projects.He previously worked as a software engineer at Google on marketing-data pipelines. For the last ~2 years he has been on a data team at a high-frequency-trading firm doing offline analysis on self-hosted/on-prem infrastructure (Kafka, Spark, Flink, Iceberg)You can find Jordan on:- Github: https://github.com/jordepic- Linkedin: https://linkedin.com/in/jordan-epstein-69b017177- Youtube: https://www.youtube.com/@jordanhasnolife5163—————————————–*TRANSCRIPT*Feed this into your favorite AI for summarization, or to prompt it specific questions:https://gist.github.com/stanislavkozlovski/5f786853b0a0f248d08d5edb42117635(or just send Gemini this video link and ask it)—————————————–*OTHER PLATFORMS*YouTube:https://youtu.be/5ybQIFb9hE8Apple:https://podcasts.apple.com/us/podcast/topicpartition/id1814926834General RSS:https://anchor.fm/s/104fd76e0/podcast/rss—————————————–If you found anything useful from this episode, please consider supporting our growth (so we can continue delivering valuable content).You can do this by simply liking the video. It takes 2 seconds to do, and recording/producing this takes us 8hrs+

Disclaimer
This podcast’s information is provided for general reference and was obtained from publicly accessible sources. The Podcast Collaborative neither produces nor verifies the content, accuracy, or suitability of this podcast. Views and opinions belong solely to the podcast creators and guests.
For a complete disclaimer, please see our Full Disclaimer on the archive page. The Podcast Collaborative bears no responsibility for the podcast’s themes, language, or overall content. Listener discretion is advised. Read our Terms of Use and Privacy Policy for more details.