[
https://issues.apache.org/jira/browse/AVRO-4295?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18102847#comment-18102847
]
ASF subversion and git services commented on AVRO-4295:
-------------------------------------------------------
Commit 20fb487ce19dfe138ac06e93f42099a6ee2f0e31 in avro's branch
refs/heads/AVRO-4295-csharp-available-bytes from Ismaël Mejía
[ https://gitbox.apache.org/repos/asf?p=avro.git;h=20fb487ce1 ]
AVRO-4295: [csharp] Grow array on demand instead of preallocating count
ReadArray resized the backing array to the full declared block count up front
(ResizeArray(ref result, i + n)) before reading any element. On a non-seekable
stream EnsureCollectionAvailable cannot bound the count against the remaining
bytes, so a large declared count could drive a multi-gigabyte Array.Resize
before the truncated input is detected.
Preallocate only a bounded amount (MaxCollectionPrealloc) up front and grow the
array geometrically on demand as elements are read. Blocks no larger than the
cap keep the original single-resize fast path; a hostile count now fails with a
bounded AvroException after a small allocation. Adds a non-seekable-stream
regression test. (ReadMap already grows via Dictionary.Add, so it is
unaffected.)
Assisted-by: GitHub Copilot:claude-opus-4.8
> [csharp] Bound allocation when decoding length-prefixed values and collections
> ------------------------------------------------------------------------------
>
> Key: AVRO-4295
> URL: https://issues.apache.org/jira/browse/AVRO-4295
> Project: Apache Avro
> Issue Type: Sub-task
> Components: csharp
> Affects Versions: 1.11.5, 1.12.1
> Reporter: Ismaël Mejía
> Assignee: Ismaël Mejía
> Priority: Major
> Labels: pull-request-available
> Time Spent: 6h
> Remaining Estimate: 0h
>
> A bytes or string value is encoded as a length prefix followed by that many
> bytes of data, and an array or map block is encoded as an element count
> followed by that many items. A malicious or truncated input can declare a
> very large length or count while carrying little or no actual data, causing a
> large allocation before the shortfall is noticed. When the source can report
> how many bytes remain, reject a declared length (or a collection block count)
> that exceeds the bytes actually available before allocating for it. Companion
> to AVRO-4241 (Java).
> BinaryDecoder.RemainingBytes() reports the bytes still readable for a
> seekable stream; ReadBytes/ReadString consult it directly
> (EnsureAvailableBytes), while DefaultReader.ReadArray/ReadMap consult it via
> MinBytesPerElement(). EnsureCollectionAvailable tracks the cumulative count
> and enforces the limits, and the schema-resolution Skip path for arrays and
> maps is bounded the same way. Negative and out-of-range counts are rejected
> before the int cast.
> Zero-byte elements (null, or a record with only zero-byte fields) consume no
> input, so the available-bytes check cannot bound their count: a tiny payload
> such as {"type":"array","items":"null"} declaring a block count of
> 200,000,000 would otherwise drive an unbounded allocation. In addition to the
> available-bytes check this therefore caps the cumulative count of zero-byte
> elements (default 10,000,000), applies a structural cap (int.MaxValue - 8,
> further clamped to the runtime's maximum array length) to every
> non-zero-byte-element collection (which also covers collections read from a
> source that cannot report the bytes remaining), and bounds the array/map skip
> paths. When set, the AVRO_MAX_COLLECTION_ITEMS environment variable caps both
> limits.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)