[ 
https://issues.apache.org/jira/browse/AVRO-4295?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18102847#comment-18102847
 ] 

ASF subversion and git services commented on AVRO-4295:
-------------------------------------------------------

Commit 20fb487ce19dfe138ac06e93f42099a6ee2f0e31 in avro's branch 
refs/heads/AVRO-4295-csharp-available-bytes from Ismaël Mejía
[ https://gitbox.apache.org/repos/asf?p=avro.git;h=20fb487ce1 ]

AVRO-4295: [csharp] Grow array on demand instead of preallocating count

ReadArray resized the backing array to the full declared block count up front
(ResizeArray(ref result, i + n)) before reading any element. On a non-seekable
stream EnsureCollectionAvailable cannot bound the count against the remaining
bytes, so a large declared count could drive a multi-gigabyte Array.Resize
before the truncated input is detected.

Preallocate only a bounded amount (MaxCollectionPrealloc) up front and grow the
array geometrically on demand as elements are read. Blocks no larger than the
cap keep the original single-resize fast path; a hostile count now fails with a
bounded AvroException after a small allocation. Adds a non-seekable-stream
regression test. (ReadMap already grows via Dictionary.Add, so it is 
unaffected.)

Assisted-by: GitHub Copilot:claude-opus-4.8


> [csharp] Bound allocation when decoding length-prefixed values and collections
> ------------------------------------------------------------------------------
>
>                 Key: AVRO-4295
>                 URL: https://issues.apache.org/jira/browse/AVRO-4295
>             Project: Apache Avro
>          Issue Type: Sub-task
>          Components: csharp
>    Affects Versions: 1.11.5, 1.12.1
>            Reporter: Ismaël Mejía
>            Assignee: Ismaël Mejía
>            Priority: Major
>              Labels: pull-request-available
>          Time Spent: 6h
>  Remaining Estimate: 0h
>
> A bytes or string value is encoded as a length prefix followed by that many 
> bytes of data, and an array or map block is encoded as an element count 
> followed by that many items. A malicious or truncated input can declare a 
> very large length or count while carrying little or no actual data, causing a 
> large allocation before the shortfall is noticed. When the source can report 
> how many bytes remain, reject a declared length (or a collection block count) 
> that exceeds the bytes actually available before allocating for it. Companion 
> to AVRO-4241 (Java).
> BinaryDecoder.RemainingBytes() reports the bytes still readable for a 
> seekable stream; ReadBytes/ReadString consult it directly 
> (EnsureAvailableBytes), while DefaultReader.ReadArray/ReadMap consult it via 
> MinBytesPerElement(). EnsureCollectionAvailable tracks the cumulative count 
> and enforces the limits, and the schema-resolution Skip path for arrays and 
> maps is bounded the same way. Negative and out-of-range counts are rejected 
> before the int cast.
> Zero-byte elements (null, or a record with only zero-byte fields) consume no 
> input, so the available-bytes check cannot bound their count: a tiny payload 
> such as {"type":"array","items":"null"} declaring a block count of 
> 200,000,000 would otherwise drive an unbounded allocation. In addition to the 
> available-bytes check this therefore caps the cumulative count of zero-byte 
> elements (default 10,000,000), applies a structural cap (int.MaxValue - 8, 
> further clamped to the runtime's maximum array length) to every 
> non-zero-byte-element collection (which also covers collections read from a 
> source that cannot report the bytes remaining), and bounds the array/map skip 
> paths. When set, the AVRO_MAX_COLLECTION_ITEMS environment variable caps both 
> limits.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to