L. C. Hsieh created SPARK-60071:
-----------------------------------

             Summary: Use identity equality for HybridRowQueue
                 Key: SPARK-60071
                 URL: https://issues.apache.org/jira/browse/SPARK-60071
             Project: Spark
          Issue Type: Bug
          Components: SQL
    Affects Versions: 5.0.0
            Reporter: L. C. Hsieh


HybridRowQueue is a case class, so two queues created in the same task with the 
same TaskMemoryManager, temp dir, number of fields, SerializerManager and 
lockFree flag are equal. Because HybridRowQueue is a MemoryConsumer (via 
HybridQueue), this breaks memory management:

1. TaskMemoryManager tracks consumers in a HashSet. A second queue that equals 
an already registered one is never added, so it is never offered for spilling 
when another consumer needs memory, and its memory is not attributed in the 
memory usage breakdown on OOM.
2. HybridQueue.spill(size, trigger) returns 0 when trigger == this, which uses 
equals. A queue therefore refuses to spill when an equal but distinct queue 
triggers the spill.

This can happen in a real query, e.g. when several partitions evaluated by a 
Python UDF node are processed in one task (coalesce), or when both sides of a 
sort-merge join contain Python UDF eval nodes with the same input width.

The fix is to give HybridRowQueue identity equality by making it a regular 
class.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to