[ 
https://issues.apache.org/jira/browse/FLINK-40627?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121214#comment-18121214
 ] 

Dale Lane commented on FLINK-40627:
-----------------------------------

We do need a local file at some point, because PackagedProgram uses a File to 
build the classloader and to upload to the blob server. And we need to keep it 
around until the main() call has finished, in case it has more to submit.

The FLINK-28915 commits that added remote fetching had an exists() check to 
reuse an existing jar from the start, and included descriptions about this 
supporting reuse of the jar across a container restart. That is a good benefit 
to have, avoiding the time needed to restart while waiting for the fetch again 
when that happens.

But I don't think we've got much managing it more formally as a cache. I don't 
know of anything that cleans up unused jars or anything like that, or anything 
like we've been talking about of along the lines of invalidating an already 
downloaded file based on metadata.

Assuming we want to keep reuse ({_}and IMO fast restarts makes that worth 
it{_}), I think we could either go for simple ({_}keep the reuse to restarts 
only, storing files based on source URI, with no extra re-validation{_}). Or we 
could make it a smarter cache ({_}that before reusing an existing file based 
solely on exists() first does a check on ETag/Last-Modified headers for the 
HTTP fetching and status for the FS fetching, and uses that to decide whether 
to get a new file{_}).

Whether it's worth the more sophisticated cache depends on how often people use 
emptyDir (only useful for container restarts, and not used for redeploys and 
pod restarts) vs persistent storage. If emptyDir is the norm (which is what 
I've mostly seen), then maybe a fancy cache is overkill?

> Partially fetched user artifacts are reused on subsequent starts
> ----------------------------------------------------------------
>
>                 Key: FLINK-40627
>                 URL: https://issues.apache.org/jira/browse/FLINK-40627
>             Project: Flink
>          Issue Type: Bug
>          Components: Client / Job Submission
>            Reporter: Dale Lane
>            Assignee: Dale Lane
>            Priority: Minor
>
> {{ArtifactFetchManager.fetchArtifact}} returns any file already present at 
> the target path without re-fetching ("Already fetched user artifacts are 
> kept"). {{HttpArtifactFetcher}} and {{FsArtifactFetcher}} write directly to 
> that final path via {{{}FileUtils.copyToFile{}}}, so if a transfer ends part 
> way through, a truncated file is left at the final path. Nothing cleans up 
> files from the artifacts directory.
> A transfer can end part way through in several ways, all of which leave a 
> truncated file:
>  * fetch fails with an exception (e.g. connection reset)
>  * Job Manager process is killed mid-transfer (e.g. OOMKilled, liveness-probe 
> kill in Kubernetes)
>  * an HTTP server without a {{Content-Length}} ends the response early (the 
> fetch then reports success, and the job fails on its first start as well as 
> every later one)
> Any later start that sees the same directory reuses the truncated file and 
> never retries the download. The job fails on every start with no fetch 
> traffic:
> {{org.apache.flink.client.program.ProgramInvocationException: Error while 
> opening jar file '.../job.jar'}}
> {{Caused by: java.io.IOException: Error while opening jar file '.../job.jar'}}
> {{Caused by: java.util.zip.ZipException: zip END header not found}}
> This affects:
>  * Native Kubernetes Application Mode: the artifacts dir 
> ({{{}<user.artifacts.base-dir>/<namespace>/<cluster-id>{}}}) is an 
> {{{}emptyDir{}}}, which survives Job Manager container restarts within a pod, 
> so the pod is stuck until it is deleted. With a persistent volume or 
> {{hostPath}} mounted at {{{}base-dir{}}}, it is stuck permanently.
>  * Standalone Application Mode with a persistent 
> {{{}user.artifacts.base-dir{}}}.
> Clearing it today needs manual deletion of the file (or pod deletion for 
> {{{}emptyDir{}}}).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to