Ma77Ball opened a new issue, #8564:
URL: https://github.com/apache/texera/issues/8564
### Feature Summary
> **In one sentence:** lay the foundation of columnar execution, the Arrow
wire format and the opt-in operator contract, and prove it end to end with two
operators (the CSV scan and the filter).
Parent: #8556 (opt-in columnar execution). This is PR 1 of the stacked
series and pairs with apache/texera#8558.
**What is this?**
Before any operator can be made faster, two things must exist: a way to
*carry* a batch of columns between workers, and a *contract* that lets an
operator opt in to reading those columns. This issue adds both, then wires up
the two simplest operators (scan and filter) so the whole path can be tested
honestly.
Think of it as laying one lane of a new highway and driving two cars down
it, while the old road stays open and default.
---
### Proposed Solution or Design
**The two new primitives.**
- **`ColumnarFrame`**: a data payload that carries one Arrow batch (the raw
column bytes) plus a row count. It rides alongside the existing per-row
`DataFrame`.
- **`ColumnarOperatorExecutor` / `ColumnarResult`**: the opt-in contract. An
operator can implement `processColumnarBatch(...)` and answer `Emit` (I made a
new batch), `Consumed` (I absorbed it), or `Unsupported` (decode it and run the
row path).
```mermaid
%%{init: {'theme':'dark', 'themeVariables':
{'background':'#000000','lineColor':'#0F766E'}}}%%
flowchart LR
SCAN[CSV scan: read file straight into Arrow columns] -->
WIRE[ColumnarFrame on the wire]
WIRE --> FILT[filter: keep rows on the column, no per-row decode]
FILT --> TERM[terminal]
classDef default fill:#000,color:#fff,stroke:#888,stroke-width:1px
```
**How the engine routes a batch.** The `DataProcessor` (the per-worker loop
that runs the operator) checks whether the operator understands columns; if
not, it decodes to rows. `OutputManager`, the partitioner, and the input side
carry the batch end to end.
```mermaid
%%{init: {'theme':'dark', 'themeVariables':
{'background':'#000000','lineColor':'#000000'}}}%%
flowchart TD
F{ColumnarFrame and<br/>operator opts in?} -->|yes| C[processColumnarBatch]
F -->|no| R[decode to rows, row path]
classDef default fill:#000,color:#fff,stroke:#888,stroke-width:1px
style C stroke:#1B7F3B
style R stroke:#B0451E
```
**What lands here.** `ColumnarFrame`, the contract, `ArrowUtils`
(serialize/deserialize a batch), engine plumbing (`DataProcessor`,
`OutputManager`, partitioner, input side), the Arrow-producing
`CSVScanSourceOpExec`, and the Arrow-consuming `SpecializedFilterOpExec`. All
behind a flag; default off.
| | Row path (default) | Columnar path (flag on) |
| --- | --- | --- |
| scan output | one `Tuple` per row | Arrow batch of columns |
| filter | evaluate per row | keep-mask over a column |
| between workers | N envelopes | 1 Arrow buffer |
Verified: the Arrow round-trip is lossless across all types, the vectorized
filter matches the row filter exactly, and the existing `DataProcessingSpec`
passes with the flag on and off.
---
*High-level overview. Part of #8556.*
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]