6 ms·
(not an expert in stream processing).. from the docs here https://sql-flow.com/docs/introduction/basics#output-sink https://sql-flow.com/docs/introduction/basic
by pulkitsh1234 9mo ago
(not an expert in stream processing).. from the docs here https://sql-flow.com/docs/introduction/basics#output-sink https://sql-flow.com/docs/introduction/basics#output-sink it seems like this works on "batches" of data, how is this different from batch processing ? Where is the "stream" here ?
- dm03514 9mo agoHa Yes! A pipeline assumes a "batch" of data, which is backed by an ephemeral duckdb in memory table. The goal is to provide SQL table semantics and implement pipelines in a way where the batch size can be toggled without a change to the pipeline logic. The stream is achieved by the continuous flow of data from Kafka. SQLFlow exposes a variable for batch size. Setting the batch size to 1 will make it so SQLFlow reads a kafka message, applies the processor SQL logic and then ensures it successfully commits the SQL results to the sink, one after another. SQLFlow provides at least once delivery guarantees. It will only commit the source message once it successfully writes to the pipeline output (sink). https://sql-flow.com/docs/operations/handling-errors https://sql-flow.com/docs/operations/handling-errors The batch table is just a convention which allows for seamless batch size configuration. If your throughput is low, or if you require message by message processing, SQLFlow can be toggled to a batch of 1. If you need higher throughput and can tolerate the latency, then the batch can be toggled higher.