DanielLeens opened a new issue, #11808:
URL: https://github.com/apache/seatunnel/issues/11808

   ### Search before asking
   
   - [x] I searched existing issues and PRs and found related ClassLoader work 
(#8288, #10669, #10678), but I did not find an issue for the `deployTask()` 
failure path leaking already-acquired classloaders.
   
   ### What happened
   
   While reviewing the current `seatunnel-engine` source on `dev`, I found a 
classloader leak path in `TaskExecutionService.deployTask(...)`.
   
   Current source chain:
   
   1. `TaskExecutionService.deployTask(TaskGroupImmutableInformation)` iterates 
through every serialized task in the task group.
   2. For each entry, it may:
      - download / resolve connector jars
      - call `classLoaderService.getClassLoader(jobId, jars)`
      - deserialize the task with that classloader
   3. If a later task in the same loop throws during jar resolution, 
classloader creation, or task deserialization, the outer `catch (Throwable t)` 
returns `TaskDeployState.failed(t)`.
   4. In that failure path, there is no compensation that releases the 
classloaders already acquired for earlier tasks in the same deployment attempt.
   
   So a partial deployment failure can leave entries behind in 
`DefaultClassLoaderService.classLoaderCache` / reference counts even though the 
`TaskGroupContext` was never successfully published.
   
   ### SeaTunnel Version
   
   Current `dev` branch source as of 2026-08-14.
   
   ### Reproduction / Evidence
   
   This report is based on static source analysis of the current engine code.
   
   Relevant methods/classes:
   
   - `org.apache.seatunnel.engine.server.TaskExecutionService#deployTask`
   - 
`org.apache.seatunnel.engine.core.classloader.DefaultClassLoaderService#getClassLoader`
   - 
`org.apache.seatunnel.engine.core.classloader.DefaultClassLoaderService#releaseClassLoader`
   
   ### Why this is different from existing ClassLoader work
   
   PR #10678 focuses on the normal release / shutdown lifecycle and explicit 
`URLClassLoader.close()` handling.
   
   This issue is about a different gap: when deployment fails before the task 
group is fully constructed, some classloaders have already been acquired, but 
the failure path never calls `releaseClassLoader(...)` for them.
   
   ### Expected behavior
   
   If `deployTask(...)` fails after acquiring one or more task classloaders, 
the method should release all classloaders / jar references acquired during the 
current attempt before returning failure.
   
   ### Why this matters
   
   Under repeated bad submissions, broken jars, incompatible task payloads, or 
partial deserialization failures, this path can gradually increase retained 
classloader references, metaspace pressure, and JAR handle usage.
   
   ### Possible fix direction
   
   - Track `taskId -> jars` (or an ordered list of acquired jar sets) during 
the current deployment attempt.
   - In the outer failure path, release every classloader acquired before the 
failure.
   - Add a regression test where a later task in the same task group fails 
after an earlier task has already acquired its classloader.
   
   ### Are you willing to submit a PR?
   
   - [ ] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's Code of Conduct
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to