[ 
https://issues.apache.org/jira/browse/FLINK-40627?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18120980#comment-18120980
 ] 

Robert Metzger commented on FLINK-40627:
----------------------------------------

Is this file storage mechanism a cache to avoid redownloading the same file 
over and over again, or a necessity to load the jar using a classloader?

I suspect it was meant for both, but its not properly implemented. We could 
ditch the caching ability, and make this purely a local file storage. Then we 
could just use a hash with a timestamp in the directory name?

If we want this to be a cache, we need to decide based on the HTTP headers if 
we need to redownload? Is that a common thing? I'm leaning towards not caching.

For the cleanup: I wonder if there's a way to delete those files, e.g. when the 
job reached a terminal execution state, or maybe already once the job has been 
submitted, because the jobgraph / jar is in the blob store? I don't know the 
exact semantics there

> Partially fetched user artifacts are reused on subsequent starts
> ----------------------------------------------------------------
>
>                 Key: FLINK-40627
>                 URL: https://issues.apache.org/jira/browse/FLINK-40627
>             Project: Flink
>          Issue Type: Bug
>          Components: Client / Job Submission
>            Reporter: Dale Lane
>            Assignee: Dale Lane
>            Priority: Minor
>
> {{ArtifactFetchManager.fetchArtifact}} returns any file already present at 
> the target path without re-fetching ("Already fetched user artifacts are 
> kept"). {{HttpArtifactFetcher}} and {{FsArtifactFetcher}} write directly to 
> that final path via {{{}FileUtils.copyToFile{}}}, so if a transfer ends part 
> way through, a truncated file is left at the final path. Nothing cleans up 
> files from the artifacts directory.
> A transfer can end part way through in several ways, all of which leave a 
> truncated file:
>  * fetch fails with an exception (e.g. connection reset)
>  * Job Manager process is killed mid-transfer (e.g. OOMKilled, liveness-probe 
> kill in Kubernetes)
>  * an HTTP server without a {{Content-Length}} ends the response early (the 
> fetch then reports success, and the job fails on its first start as well as 
> every later one)
> Any later start that sees the same directory reuses the truncated file and 
> never retries the download. The job fails on every start with no fetch 
> traffic:
> {{org.apache.flink.client.program.ProgramInvocationException: Error while 
> opening jar file '.../job.jar'}}
> {{Caused by: java.io.IOException: Error while opening jar file '.../job.jar'}}
> {{Caused by: java.util.zip.ZipException: zip END header not found}}
> This affects:
>  * Native Kubernetes Application Mode: the artifacts dir 
> ({{{}<user.artifacts.base-dir>/<namespace>/<cluster-id>{}}}) is an 
> {{{}emptyDir{}}}, which survives Job Manager container restarts within a pod, 
> so the pod is stuck until it is deleted. With a persistent volume or 
> {{hostPath}} mounted at {{{}base-dir{}}}, it is stuck permanently.
>  * Standalone Application Mode with a persistent 
> {{{}user.artifacts.base-dir{}}}.
> Clearing it today needs manual deletion of the file (or pod deletion for 
> {{{}emptyDir{}}}).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to