Hi Rahul and Dr. Li,

I had a discussion with Professor Zou. The current implementation is
affecting the round-robin partitioner. What we're proposing instead is a
new, separate partitioner. I'll isolate the current changes so they don't
impact round-robin and create a separate PR within the next two weeks.

This has taken a bit longer as I've been working on another SIGMOD
submission, but I'll have this wrapped up within that timeline.

Thanks and Regards,
Jay

On Mon, Sep 14, 2026 at 9:03 PM Rahul P <[email protected]> wrote:

> Hi
>
> When is the optimization expected to be integrated?
>
> Thanks & Regards
> Rahul
>
>
> On Tue, 1 Sept, 2026, 08:13 Chen Li, <[email protected]> wrote:
>
>> Thank you, Jay,  It will be great to have your optimization integrated
>> into
>> the system.
>>
>> Chen
>>
>> On Mon, Aug 31, 2026 at 3:38 PM Jaykumar Rajeshbhai Tandel <
>> [email protected]>
>> wrote:
>>
>> > Hi Dr. Li and Rahul,
>> >
>> > Yes, we observed the same single-worker limitation on the CSV reader. In
>> > our case, though, it was not the bottleneck. For UDF-centric workloads,
>> we
>> > found the dominant costs were the distribution of work to downstream UDF
>> > workers and the per-tuple build/read/serialize for the Python UDF.
>> >
>> > For reference, our setup: two neural-network UDFs on CPU over a 2.25 GB
>> > dataset, on a machine with 48 cores and 125 GB of memory. The best
>> observed
>> > end-to-end time was 11 minutes 8 seconds. The two UDFs and a link to the
>> > dataset are attached.
>> >
>> > Two changes made the largest difference:
>> >
>> > *Use more workers.* Each Texera worker runs as a separate process, so
>> UDF
>> > work parallelizes well when the machine has cores to spare. We used 20
>> > workers for each UDF.
>> >
>> > *Use a smaller data transfer batch size.* A batch only ships once a
>> > buffer fills, and each sending worker keeps a separate buffer per
>> receiving
>> > worker, so a large batch size can delay downstream workers
>> considerably. We
>> > saw the best result at a 400 batch size if parallelization is needed.
>> >
>> > Rahul, with 3 workers and 27M records, this is worth checking. Texera
>> does
>> > not yet allocate resources for UDFs automatically, so worker counts and
>> > batch size have to be set manually, based on observations and the
>> resources
>> > available at your end.
>> >
>> > We are currently working on making that overlap independent of the batch
>> > size, so that pipelining also holds at larger batch sizes in multi-UDF
>> > workflows. We plan to raise a PR for the parallelining issue by Sept 9,
>> > and one for the per-tuple cost by Sept 15, so the wider community can
>> > benefit.
>> > Please feel free to reach out if you have more queries.
>> >
>> > Thanks and Regards,
>> > Jay.
>> >
>> >  creditcard_x16.csv
>> > <https://drive.google.com/open?id=12n5TRN64sJDFtsa-3igbGguQqYopo1AY>
>> >
>> >  UDF Code
>> > <
>> https://drive.google.com/open?id=1jgBexhIhxnoxx0GC_ZPMwI2eJ-ThtsDPnJfYjvQgrxg
>> >
>> >
>> > On Sun, Aug 30, 2026 at 9:27 PM Chen Li <[email protected]> wrote:
>> >
>> >> Rahul,
>> >>
>> >> Glad to know you were able to upload a 1GB file and process it.  We
>> know
>> >> the system has a lot of room for performance improvement.  You are very
>> >> welcome to do some investigation and identify bottlenecks.
>> >>
>> >> I am including two colleagues from ASU, Jaykumar Rajeshbhai Tandel <
>> >> [email protected]> and Jia Zou <[email protected]>.  Recently they are
>> >> investigating related issues, and may be able to chime in to
>> >> provide insights.
>> >>
>> >> We will be happy to work with you on an optimization project.
>> >>
>> >> Chen Li
>> >> PPMC
>> >> Apache Texera (Incubating)
>> >>
>> >> On Sun, Aug 30, 2026 at 8:14 AM Rahul P <[email protected]>
>> wrote:
>> >>
>> >>> Hi  Chen Li / Team
>> >>>
>> >>> Thanks for your valuable feedback.
>> >>>
>> >>> I have installed Texera on a remote Linux server and created the 1 GB
>> >>> test file on that server. Since I am accessing the Texera UI from my
>> >>> Windows machine, I manually used Texera's multipart upload API from
>> the
>> >>> Linux server to upload the file as a Texera dataset.
>> >>>
>> >>> During the upload, I initially encountered a file size limit issue,
>> >>> which I resolved by increasing the configured limit. After that, the
>> file
>> >>> was uploaded successfully.
>> >>>
>> >>> During execution, I observed that the *File Scan operator was reading
>> >>> the file in small chunks and writing the data in Iceberg format*. I
>> >>> increased the chunk size and restarted the execution.
>> >>>
>> >>> However, the *Python UDF operator*, which reads one line at a time,
>> >>> parses it as a CSV row, converts it into six structured fields, and
>> outputs
>> >>> the tuple, took around *11 hours* to process the 27 million records.
>> >>>
>> >>> I am using *3 worker nodes*, and the Python UDF was also running with
>> *3
>> >>> workers*.
>> >>>
>> >>> Is this expected behavior for Texera, or am I missing any
>> configuration
>> >>> or optimization that could improve the Python UDF performance?
>> >>>
>> >>> [image: image.png]
>> >>>
>> >>>
>> >>> Thanks & Regards
>> >>> Rahul
>> >>>
>> >>> On Fri, Aug 28, 2026 at 11:17 AM Chen Li <[email protected]> wrote:
>> >>>
>> >>>> Rahul,
>> >>>>
>> >>>> Sorry for the late reply.  Currently the CSVFileScan operator doesn't
>> >>>> support multiple workers to run in parallel.  So reading a large file
>> >>>> may
>> >>>> take time.  If you are interested in developing a parallel reader,
>> you
>> >>>> are
>> >>>> very welcome to do so and the team can provide support.
>> >>>>
>> >>>> The Python UDF operator allows a user to specify the number of
>> workers.
>> >>>>
>> >>>> The backend engine is doing in-memory computation.  Data transfer
>> >>>> between
>> >>>> operators can be pipelined.
>> >>>>
>> >>>> The backend storage layer supports large files using solutions such
>> as
>> >>>> MinIO, LakeFS, and Iceberg.
>> >>>>
>> >>>> The system had done a lot of optimization in uploading large files
>> >>>> efficiently (e.g., in the range of TBs).
>> >>>>
>> >>>> Our team will be happy to answer questions related to your particular
>> >>>> use
>> >>>> case.
>> >>>>
>> >>>> Thank you,
>> >>>>
>> >>>> Chen Li
>> >>>>
>> >>>>
>> >>>>
>> >>>>
>> >>>> On Thu, Aug 27, 2026 at 9:17 PM Rahul P <[email protected]>
>> >>>> wrote:
>> >>>>
>> >>>> > Hi Team
>> >>>> >
>> >>>> > Can I please get an update on this.
>> >>>> >
>> >>>> > Thanks & Regards
>> >>>> > Rahul
>> >>>> >
>> >>>> > On Thu, Aug 27, 2026 at 12:09 PM Rahul P <[email protected]
>> >
>> >>>> wrote:
>> >>>> >
>> >>>> > > Hi Team,
>> >>>> > >
>> >>>> > > I have installed Texera and am able to execute workflows
>> >>>> successfully. I
>> >>>> > > would like to understand how large-scale data processing is
>> handled
>> >>>> in
>> >>>> > > Texera.
>> >>>> > >
>> >>>> > > If the input data size is very large, will Texera process it
>> >>>> reliably
>> >>>> > > without crashing, or are there any known limitations related to
>> file
>> >>>> > size,
>> >>>> > > memory, or cluster resources?
>> >>>> > >
>> >>>> > > I also could not find any option to configure parallelism for the
>> >>>> file
>> >>>> > > fetch/file scan operator. How can parallelism be configured for
>> this
>> >>>> > > operator?
>> >>>> > >
>> >>>> > > Thanks & Regards
>> >>>> > > Rahul
>> >>>> > >
>> >>>> >
>> >>>>
>> >>>
>>
>

Reply via email to