Show HN: DuckDB for Kafka Stream Processing

1. mihevc ◴[08 Dec 25 18:38 UTC] No.46195958[source]▶

>>46195007 (OP) #

How does this compare to https://github.com/Query-farm/tributary ?

replies(2): >>46196154 #>>46196322 #

2. dm03514 ◴[08 Dec 25 18:53 UTC] No.46196154[source]▶

>>46195958 (TP) #

Oh yes!! I've seen this a couple times. I am far from an expert in tributary so please take with a grain of salt.

Based on the tributary documentation, I understand that tributary embeds kafka consumers into duckdb. This makes duckdb the main process that you run to perform consumption. I think that this makes creating stream processing POCs very accessible. It looks like it is quite easy to start streaming data into duckdb. What I don't see is a full story around Devops, operations, testing, configuration as code etc.

SQLFlow is a service that embeds DuckDB as the storage and processing brains. Because of this, we're able to offer metrics, testing utilities, pipelines as code, and all the other DevOps utilities that are necessary to run a huge number of streaming instances 24x7. SQLFlow was created as a tool that I wish I had to for simple stream processing in production in high availability contexts :)

replies(1): >>46196283 #

3. mihevc ◴[08 Dec 25 19:05 UTC] No.46196283[source]▶

>>46196154 #

Nice! Thanks for the context, it's great to know!

4. rustyconover ◴[08 Dec 25 19:10 UTC] No.46196322[source]▶

>>46195958 (TP) #

The next major release of Tributary will support Avro, Protobuf and JSON along with the Schema Registry it will also bring the ability to write to Kafka with transactions.

But really you should get excited for DuckDB Labs to build out materialized views. Materialized views where you can ingest more streaming data to update aggregates. This way you could just keep pushing rows through aggregates from Kafka.

It is going to be a POWER HOUSE for streaming analytics.

Contact DuckDB Labs if you want to sponsor the work on materialized views: https://duckdb.org/roadmap

replies(2): >>46197772 #>>46200924 #

5. buremba ◴[08 Dec 25 21:18 UTC] No.46197772[source]▶

>>46196322 #

Exactly. I have also been playing with DuckDB for streaming use cases, but it feels hacky to issue micro-batching queries on streaming data in short intervals.

DuckDB has everything that streaming engines such as Flink have; it just needs to support managing intermediate aggregate states and scheduling the materialized views itself.

6. trueno ◴[09 Dec 25 03:27 UTC] No.46200924[source]▶

>>46196322 #

Is this to be used in an analytics application backend sort of scenario?

I am familiar with materialized views / dynamic tables from enterprise-grade cloud lake type offerings, but I've never quite understood where duckdb, though impressive, fits into everyones use case. I've toyed with it for personal things, it's very cool having a local instance of something akin to snowflake when it comes to processing and aggregating on Big Data™ but generally I don't see it used in operational settings. For application development people are generally tied to sqlite and postgres.

It all does seem really cool though, I guess I'm just not feeling creative enough to conjure up a stream-to-duckdb use case. Feel free to bombard me with cool ideas.