jerryshao opened a new issue, #13372:
URL: https://github.com/apache/gravitino/issues/13372

   ### Version
   
   main branch
   
   ### Describe what's wrong
   
   A job's staging directory is created as 
`<gravitino.job.stagingDir>/<metalake>/<templateName>/job-<id>`, using the 
template name at the time the job runs. `JobManager` never stores this path. 
When it later cleans a job up, it rebuilds the path from 
`job.jobTemplateName()`, both in `cleanUpStagingDirs()` (after 
`gravitino.job.stagingDirKeepTimeInMs`) and in `deleteJobTemplate()`. 
`job.jobTemplateName()` is read by joining the job-template table, so after 
`alterJobTemplate(rename)` it returns the **new** name.
   
   The rebuilt path then points to a directory that doesn't exist, and 
`FileUtils.deleteDirectory` quietly does nothing. The job entity is still 
deleted, so nothing references the old directory anymore and it is never 
removed. This includes everything in it: downloaded executables, scripts, jars, 
and `output.log` / `error.log`.
   
   Renaming a metalake probably has the same effect, because the path is also 
rebuilt from the current metalake name. I haven't verified that case.
   
   ### Error message and/or stacktrace
   
   None. The cleanup reports success.
   
   ### How to reproduce
   
   Verified on main with the `JobManager` / `LocalJobExecutor` test harness of 
`TestJobManagerMultiNode`:
   
   1. Run a job from template `echo_job` and wait for it to finish. 
`<stagingDir>/<metalake>/echo_job/job-<id>` exists.
   2. Rename the template to `renamed_echo`.
   3. Let the job expire (its `finishedAt` is older than 
`gravitino.job.stagingDirKeepTimeInMs`) and run `cleanUpStagingDirs()`. The job 
entity is deleted, but `<stagingDir>/<metalake>/echo_job/job-<id>` is still 
there.
   4. Or: run a job from `renamed_echo`, rename the template again to 
`renamed_twice`, and delete it with `deleteJobTemplate()`. The template and the 
job entity are deleted, but `<stagingDir>/<metalake>/renamed_echo/job-<id>` is 
still there.
   
   ### Additional context
   
   Found while working on multi-node job output retrieval (#12716), part of 
epic #12667. The output index added for that work 
(`<stagingDir>/.job-output-index`) is only removed once its job's staging 
directory is gone, so these leaked jobs also keep their index files.
   
   Possible fix: stop rebuilding the path from names that can change. Either:
   - derive it from data fixed at run time, e.g. the job's persisted 
`runtimeJobTemplate`, whose executable lives in the job's staging directory; or
   - make the staging layout depend only on immutable ids, keeping cleanup of 
the existing layout for one retention period after upgrade.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to