[ 
https://issues.apache.org/jira/browse/SPARK-58505?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Wenchen Fan reassigned SPARK-58505:
-----------------------------------

    Assignee: Ganesha S

> Attach per-consumer memory breakdown to UNABLE_TO_ACQUIRE_MEMORY task errors
> ----------------------------------------------------------------------------
>
>                 Key: SPARK-58505
>                 URL: https://issues.apache.org/jira/browse/SPARK-58505
>             Project: Spark
>          Issue Type: Improvement
>          Components: Spark Core
>    Affects Versions: 4.2.0
>            Reporter: Ganesha S
>            Assignee: Ganesha S
>            Priority: Major
>              Labels: pull-request-available
>
> h3. Problem
> When a task cannot acquire execution memory, Spark throws a 
> SparkOutOfMemoryError with error class UNABLE_TO_ACQUIRE_MEMORY:
> {code:java}
> Unable to acquire 8388608 bytes of memory, got 2097152.{code}
> This message, which propagates into the TaskFailedReason and surfaces on the 
> driver and in the Spark UI, reports only the requested and received byte 
> counts. It does not indicate which operator was holding the memory, so it 
> gives no direct signal about
> the cause of the OOM.
> `TaskMemoryManager.showMemoryUsage()` does compute a per-MemoryConsumer 
> breakdown, but it writes that breakdown only to the executor logs. Recovering 
> it after a failure means locating the correct executor's logs and correlating 
> by task-attempt id, which is often impractical (logs rotated/aggregated, 
> executor lost) and is not accessible to the driver or to 
> programmatic/automated diagnosis.
> h3. Proposal
> Attach the same per-consumer attribution that showMemoryUsage() already logs 
> to the `UNABLE_TO_ACQUIRE_MEMORY` error itself, so it travels with the task 
> failure reason to the driver and the UI. The message becomes, for example:
> {code:java}
> Unable to acquire 8388608 bytes of memory, got 2097152.
>   Memory used by task 4211 grouped by consumer:
>     org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@1a2b: 
> 456.0 MiB
>     org.apache.spark.unsafe.map.BytesToBytesMap@3c4d: 12.0 MiB
>     (not attributed to a specific consumer): 3.0 MiB{code}
> Consumers are listed largest-first (the most likely culprit surfaces first), 
> followed by the bytes not attributable to any specific consumer.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to