Thanks for the detailed info, Carlos. Let me give your suggestions a try and I will see what the details show.
- ian On Sun, Aug 30, 2026 at 9:18 PM Carlos Ernesto Alvarez Berumen <[email protected]> wrote: > > Hi Ian, > > I recently deployed v1.2.0-incubating on our Kubernetes/RKE2 cluster. I > checked the exact release source and our completed Job. > > In the 1.2.0 template at commit c3aa6e4, the commands execute in this order: > > set -e > apk add --no-cache curl ca-certificates wget > wget -q https://dl.min.io/client/mc/release/linux-amd64/mc -O > /usr/local/bin/mc > chmod +x /usr/local/bin/mc > ... > echo "Waiting for Lakekeeper health endpoint..." > > One important distinction: seeing package-installation output does not > necessarily mean apk add returned successfully. If its final OK message is > absent, apk itself remains a possible failure point. > > If apk add did return successfully, the pod was not externally terminated, > and “Waiting…” never appeared, then the remaining script-level failure > points in that interval are wget and chmod. > > The absence of a download error does not rule out wget. It uses -q, which GNU > Wget documents as disabling its output. Because this is a simple command > under set -e, a nonzero exit would terminate the script immediately. > > For comparison, our deployment reports: > > Job: texera-lakekeeper-init > Succeeded: 1 > Completion: 2026-08-21T19:51:25Z > Pod exit code: 0 > > Its logs continue from the completed Alpine installation to: > > Waiting for Lakekeeper health endpoint... > ... > Lakekeeper initialization sequence completed successfully! > > Therefore, we did not reproduce the failure on Jails, but the release > script has a diagnostic blind spot around the mc download. > > Since the Job has backoffLimit: 3, it may have created multiple pod > attempts. I would inspect all of them: > > kubectl get pods -n <namespace> \ > -l job-name=<release>-lakekeeper-init > > kubectl get pod -n <namespace> <failed-pod> \ > -o jsonpath='{.status.containerStatuses[0].state.terminated.reason}{" > exit="}{.status.containerStatuses[0].state.terminated.exitCode}{" > message="}{.status.containerStatuses[0].state.terminated.message}{"\n"}' > > kubectl logs -n <namespace> <failed-pod> > kubectl describe pod -n <namespace> <failed-pod> > > For one diagnostic run, I would replace only the download block with: > > apk add --no-cache curl ca-certificates wget > > echo "Downloading MinIO client..." > if ! wget -S \ > https://dl.min.io/client/mc/release/linux-amd64/mc \ > -O /usr/local/bin/mc; then > echo "MinIO client download failed" > exit 1 > fi > > if [ ! -s /usr/local/bin/mc ]; then > echo "Downloaded MinIO client is missing or empty" > exit 1 > fi > > ls -l /usr/local/bin/mc > chmod +x /usr/local/bin/mc > > This should provide evidence distinguishing a download failure from a > subsequent file or permission failure. > > I would avoid enabling set -x for the entire script because a later mc > alias set command expands the MinIO credentials. If tracing is used, it > should be disabled before that command. > > If you share the termination reason, exit code, and non-quiet wget output, > I can compare them with our deployment. > > Best, > Carlos
