furquan39 opened a new issue, #66298:
URL: https://github.com/apache/doris/issues/66298

   ### Search before asking
   
   - [x] I had searched in the 
[issues](https://github.com/apache/doris/issues?q=is%3Aissue) and found no 
similar issues.
   
   ### Version
   
   Apache Doris 4.1.0-rc03. All file:line references below are at tag 
`4.1.0-rc03`.
   
   ### Environment
   
   We operate Doris 4.1.0-rc03 on Kubernetes with Ceph-backed storage.
   
   Related: "[Enhancement] Compaction scheduler retries permanently-failing 
tablets every 5s at full I/O cost with no backoff or failure counter" 
(companion issue filed together with this one; link to be added once both are 
open). It covers the retry policy that amplifies this bug into an I/O storm; 
this issue is about the segment sizing defect itself.
   
   ### What's Wrong?
   
   **Summary:** Vertical compaction plans output segment size in compressed 
bytes, but the string column engine enforces a hard limit on uncompressed 
bytes. For data that compresses well, the planned segment deterministically 
exceeds the uncompressed limit and compaction fails with E-3113 forever.
   
   **Mechanism:**
   
   1. `ColumnStr<T>` uses uint32 offsets. `ColumnStr<T>::check_chars_length` 
(`be/src/core/column/column_string.h:64-78`, `MAX_STRING_SIZE = 0xffffffff`) 
throws `STRING_OVERFLOW_IN_VEC_ENGINE` (E-3113) when the accumulated 
uncompressed character bytes of a single column exceed 4 GiB. During compaction 
merges this is reached via `insert_range_from` 
(`be/src/core/column/column_string.cpp:186`).
   
   2. Vertical compaction decides how many rows go into each output segment 
from **compressed** sizes: `Compaction::get_avg_segment_rows` computes 
`vertical_compaction_max_segment_size / average compressed row size` 
(`compaction.cpp:319-335`). So the segment cap is denominated in compressed 
bytes, while the limit in (1) is denominated in uncompressed bytes. Nothing in 
the merge path tracks accumulated uncompressed string length against 
`MAX_STRING_SIZE`.
   
   3. The mismatch is compression-ratio dependent. At compression ratio `r`, a 
segment capped at `S` compressed bytes carries roughly `S * r` uncompressed 
string bytes, so the overflow condition is `S * r > 4 GiB`. Equivalently, the 
safe compressed cap is `4 GiB / r`. Since `r` is a property of the data and is 
unbounded, **no static compressed-size cap is universally safe**. Our corpus 
(high-duplication JSON, `r` around 100x) overflowed even with the cap lowered 
to 256 MiB, which matches the law: 256 MiB * 100 is far above 4 GiB.
   
   4. VARIANT tables are the most exposed: `VariantColumnWriterImpl` buffers an 
entire output segment in memory before writing 
(`variant_column_writer_impl.cpp:1478-1487`), so the full segment's 
uncompressed data for a subcolumn must pass through a single `ColumnStr` and 
the 4 GiB check is applied to the whole segment at once.
   
   **Verbatim failure from our reproduction** (logged at `tablet.cpp:1889`):
   
   ```
   [E-3113] string column length is too large: total_length=4297424560, 
element_number=901728, rows=900736
   ```
   
   Note `total_length = 4297424560` is only about 2.4 MB above `0xffffffff`: 
the merge marches straight into the uint32 ceiling because nothing bounds it in 
uncompressed units.
   
   ### What You Expected?
   
   Compaction plans output segments so that the uncompressed string-column 
invariant is never violated: a merge either succeeds, or cuts a new output 
segment before the accumulated uncompressed string bytes reach the uint32 cap. 
A static compressed-size knob cannot express this expectation, because the safe 
value depends on the per-table compression ratio, which is a property of the 
data and not known in advance.
   
   ### How to Reproduce?
   
   1. Create a table with a VARIANT column (long plain STRING columns also 
work, VARIANT just hits it soonest because of the whole-segment buffering 
above).
   2. Load a corpus that compresses very well. Ours is high-duplication JSON 
with an observed compression ratio around 100x; we reproduced at 145 GB corpus 
scale on a dev cluster.
   3. Let cumulative/base compaction pick up the tablet. `get_avg_segment_rows` 
plans segments from compressed row size, the merge accumulates more than 4 GiB 
of uncompressed string bytes into one output column, and `check_chars_length` 
throws E-3113.
   4. Lowering `vertical_compaction_max_segment_size` does not fix it for 
high-ratio data. We still overflowed at a 256 MiB cap. Only a cap below `4 GiB 
/ r` would avoid it, and `r` is not known in advance and varies per tablet.
   
   ### Anything Else?
   
   **What it caused in production:** The failure is deterministic: every retry 
re-plans the same merge and fails at the same point, so the affected tablet 
never compacts again.
   
   Combined with the retry policy this becomes an I/O storm: failed compactions 
are retried on a fixed 5 s cooldown with no failure counter and no backoff, so 
the deterministic failure above is re-attempted at full merge I/O cost 
indefinitely. On one 3.14 GB tablet with **zero ingest** this wrote 51.9 GiB to 
storage in 67 minutes with zero durable progress, degrading everything else on 
the Ceph-backed storage. The retry policy itself, the full measurements, and 
the operator mitigation tradeoff (disabling auto compaction stops the storm but 
also disables the emergency force-compaction escape, so the tablet walks toward 
-235 ingest stalls) are detailed in the related issue above.
   
   **Status on master:** The reproduction is on 4.1.0-rc03, and the relevant 
code is verified unchanged on current master as of 2026-07-30: the uint32 
uncompressed limit and the compressed-byte segment sizing in 
`get_avg_segment_rows` are both still present (the compaction anchors above 
live under `be/src/storage/` on master after the PR #61107 restructure, which 
moved `be/src/olap/*` to `be/src/storage/*`). One difference to note for log 
matching: PR #63183 (merged 2026-05-18 on master) reworded the 
`STRING_OVERFLOW_IN_VEC_ENGINE` message, so master's error text differs; the 
quoted text above is verbatim from 4.1.0-rc03. The underlying limit and the 
compressed-vs-uncompressed sizing mismatch are unchanged on master.
   
   **Prior reports of the same 4 GiB cap:** #34971 and #19919 hit the same 
limit via other code paths and were closed stale without a fix; #49537 is also 
related. None of them addresses the compaction segment-sizing mismatch reported 
here.
   
   **Suggested direction:** We may be missing context on why segment sizing 
uses compressed bytes, so please treat these as suggestions:
   
   1. Size output segments by **estimated uncompressed** bytes instead of 
compressed bytes in `get_avg_segment_rows` (`compaction.cpp:319-335`), or
   2. Track the accumulated uncompressed string length per output column during 
the vertical merge and cut a new segment before `MAX_STRING_SIZE` is reached, 
independent of the configured compressed-size cap. This would make the 
invariant self-enforcing regardless of compression ratio.
   
   Option 2 seems more robust to us since it does not depend on estimating the 
ratio, and it would also bound the whole-segment memory buffering in 
`VariantColumnWriterImpl` (`variant_column_writer_impl.cpp:1478-1487`) as a 
side effect.
   
   Happy to provide full logs, tablet metadata, or run patched builds against 
our reproduction corpus.
   
   ### Are you willing to submit PR?
   
   - [ ] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's Code of Conduct
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to