The GitHub Actions job "License Binary Checker" on texera.git/main has failed.
Run started by GitHub user github-merge-queue[bot] (triggered by 
github-merge-queue[bot]).

Head commit for run:
209fc4152e33eb7acd4fde7a33b8402962e508dc / ali risheh <[email protected]>
feat(computing-unit): out-of-pod LakeFS repository mount infrastructure (#6866)

### Abstract
This PR adds infrastructure to mount LakeFS repositories to any pod,
later it will be used for models and datasets be mounted on computing
unit pods. The goal of this PR to add `mounter` service and `S3 proxy`
to file service. We also moved computing unit prefix to configuration,
in the past it was "computing-unit" by default, we just moved it to
configuration to have one source of truth because mounter needs to know
computing unit pod name.

### What changes were proposed in this PR?

Perform the FUSE mount for dataset repositories **outside** the
(unprivileged) computing-unit pod — the infrastructure foundation of the
dataset-mounting feature (#6606).

- **`texera-mounter` DaemonSet** — a per-node privileged agent
(`bin/mounter/mounter.py` + tests, dockerfile, helm
daemonset/rbac/values) that runs GeeseFS on a pod's behalf. The
read-only mount reaches the CU pod via Kubernetes **mount propagation**,
so the pod that runs user code stays **unprivileged**.
- **Per-computing-unit isolation** — mounts land at
`<mount-root>/<cuid>/<repository>/<commit>`, and a CU pod's `hostPath`
volume is only its own `<cuid>` subtree. Two units mounting the same
version get two separate GeeseFS mounts, each authorized with its own
JWT. A pod watcher unmounts a CU's directories when its pod is deleted.
- **Unprivileged CU pod wiring** (`KubernetesClient`) — the propagation
volume and mount env, added only when the feature is switched on (see
below).
- **File-service JWT S3 proxy** (`S3ProxyServlet`) — fronts the LakeFS
S3 gateway: verifies the pod's JWT, checks the user's read access,
re-signs to LakeFS with credentials held only server-side. No global
credential enters the pod.

Foundation only — nothing triggers a mount yet; the platform integration
(engine client + per-CU mount API + UI + UDF bindings) comes in the
follow-up PR.

**Off by default (`mounter.enabled: false`).** A reviewer running this
branch on Talos could not
create a computing unit at all:

```
pods "computing-unit-1" is forbidden: violates PodSecurity "baseline:latest":
hostPath volumes (volume "texera-mounts")
```

The CU pod was being given a `hostPath` unconditionally, and both the
`baseline` and `restricted`
Pod Security Standards forbid `hostPath` — so on any cluster enforcing
either on the pool
namespace (Talos does so by default), *every* computing unit becomes
unschedulable, whether or not
anyone wants to mount a dataset. It went unnoticed locally because a
default minikube enforces
nothing.

Since no caller requests a mount until the follow-up PR, the feature is
now opt-in. `mounter.enabled`
gates the DaemonSet, its RBAC, the access-control-service identity and
token, and — through
`kubernetes.mounter-enabled` — the CU pod's `hostPath`, its mount and
its env. With the flag off the
chart renders no mounter object and the CU pod spec is byte-for-byte
what it was before this feature
existed. Enabling it requires a cluster that admits `hostPath` in the
pool namespace and a
privileged pod in the release namespace, so an operator opts in once
that is true for them.


<img width="1241" height="423" alt="Ali Texera-geeseFS (4)"
src="https://github.com/user-attachments/assets/c951d777-1285-48e1-be5a-082aa4b52967";
/>


### Mount request validation and the caller the mounter trusts

Addressing the review on request validation:

- **Every path component is validated, not just `repo`/`commit`.**
`cuid`, `repositoryName`
and `commitHash` are all joined into
`MOUNT_ROOT/<cuid>/<repo>/<commit>`, and the
directory is created **before** geesefs — and therefore LakeFS — ever
sees the request,
so "LakeFS rejects a bad repository" was never a defence for the *path*.
`cuid` must now
match `^[0-9]+$` (it is the computing unit's integer primary key);
`repositoryName` and
`commitHash` must each be a single safe segment,
`^[A-Za-z0-9][A-Za-z0-9._-]*$` — which
admits everything the platform actually sends (`dataset-<did>` and a hex
digest) while
rejecting a separator, a `..`, an absolute path, or a leading `-` that
geesefs might read
  as a flag. Anything else is a `400`, and nothing is created on disk.
- **`_remove_empty_dirs` had a prefix bug.** It tested
`path.startswith(MOUNT_ROOT/<cuid>)`,
so a sibling whose name merely began the same way — `<root>/7x` against
`<root>/7` — was
treated as a child and deleted. It now compares path segments
(`os.path.commonpath`).
The `stop_at` name and docstring were also wrong: the loop removes
`MOUNT_ROOT/<cuid>`
itself and stops at its parent. That is safe for a running pod — the
CU's hostPath volume
is `DirectoryOrCreate`, so the next mount recreates it — and both the
name and the
  docstring now say so.
- **The mounter API is safeguarded: only access-control-service can call
it.** The mounter is
a privileged, per-node DaemonSet, so before this merges it must not be
callable by anything
else. Two things enforce that, and both are decided by the
kube-apiserver rather than by the
  mounter:
  - **A dedicated ServiceAccount is the identity.** This PR adds
`access-control-service-service-account.yaml` and binds ACS to it. ACS
is already the
JWT-authenticating routing proxy that validates the user's token and
checks their
computing-unit access, so it is the right — and only — place for that
authorization; the
mounter stays a small service that mounts what one known caller asks
for.
- **An audience scopes the credential to the mounter.** ACS receives a
projected
`serviceAccountToken` bound to the audience `texera-mounter`, separate
from its ordinary
kube-apiserver token. `authenticate_caller` (`bin/mounter/mounter.py`)
submits it to the
`TokenReview` API and requires all three of: the token verifies,
`texera-mounter` is among
    the audiences the API server echoes back, and the username equals
`system:serviceaccount:<ns>:<acs-sa>`. Anything else is a `401` and no
mount happens.
The audience matters because without it a TokenReview validates against
the API server's
own audience — so a *generic* ACS token, one that leaked into a log or a
crash dump, would
be accepted. Bound to `texera-mounter`, only the credential minted for
the mounter works.

The audience and the allowed caller are fixed in
`templates/base/_helpers.tpl` and are
deliberately **not** settable in `values.yaml`: both sides of the
contract must agree, and
widening the allow-list is a security decision, not a deployment
preference. `/healthz` stays
open because the kubelet probes it and holds no token for this audience;
`/mount` and
`/mounts` are both gated. This is deliberately not a `NetworkPolicy` — a
NetworkPolicy is
silently unenforced on CNIs that do not implement it (EKS's VPC CNI has
it off by default)
and is commonly bypassed by hostPort traffic, whereas `TokenReview`
holds regardless of how
the request arrived. The mounter is also reachable only in-cluster: it
has no Service and no
Ingress, so nothing routes to it from the gateway. The path validation
above holds
independently of all this, so a malformed `cuid` is refused whether or
not the caller
  authenticates.

### Any related issues, documentation, discussions?

Closes #6862 · part of #6606.

### How was this PR tested?

- `sbt FileService/compile ComputingUnitManagingService/compile` green.
- The mounter has its own **pytest suite** (`bin/mounter/tests`, now 93
tests) covering mount, the pod-deletion reaper, dead-mount self-heal,
and — new in this revision — request validation, including the reported
`cuid=5/../8`, `cuid=../..` and absolute-`cuid` escapes, the equivalent
`repositoryName`/`commitHash` escapes, and the sibling-directory
deletion bug, each asserted both at the function level and end to end
over the mounter's real HTTP surface (`400` + no `geesefs` invocation +
nothing created on disk). This suite **is** run in CI: the `build /
infra` job runs `pytest bin/` on ubuntu and macos. The proxy's
request-parsing helpers are unit-tested (`S3ProxyServletSpec`).
- Validated **end-to-end on a single-node minikube**: a Python UDF read
a ~2 GB sharded PyTorch model from a propagated mount via `torch.load`
with **bit-exact** output; the proxy's JWT authorization was exercised
for both an authorized user (200) and an unauthorized repository (403 +
refused to mount).

> **Note on patch coverage:** most of this PR is inherently
integration/IO code — the S3 proxy's request **forwarding + re-signing**
(needs a live LakeFS gateway) and the Kubernetes **pod wiring** (fabric8
has no mock server in this repo). codecov's patch % therefore reads low
even though the behavior is covered by the end-to-end validation above
and by the mounter's pytest suite — which runs in CI under `build /
infra` but is not instrumented by codecov, being a `bin/` script rather
than a build module. We'd appreciate reviewers weighing the
patch-coverage signal in that light.

### Was this PR authored or co-authored using generative AI tooling?

Generated-by: Claude Opus 4.8

---------

Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>

Report URL: https://github.com/apache/texera/actions/runs/34114740384

With regards,
GitHub Actions via GitBox

Reply via email to