viukpe opened a new issue, #74205: URL: https://github.com/apache/airflow/issues/74205
### Under which category would you file this issue? Providers ### Apache Airflow version 3.0.6 ### What happened and how to reproduce it? ### Apache Airflow version Reproduced on Airflow 3.0.6 (official AWS MWAA 3.0.6 image, `apache-airflow-providers-ssh==4.1.3`). Not limited to that version: the faulty `select()` read loop in `SSHHook.exec_ssh_client_command` is byte-for-byte identical in every released SSH provider from 3.14.0 (shipped with Airflow 2.10) through the latest release 6.1.0 (2026-09-29), and on `main`. It has never been changed. So the bug is present on the current latest provider and latest Airflow 3.x; it is not fixed anywhere. > Related prior report: [Discussion #30085](https://github.com/apache/airflow/discussions/30085) describes the same `select()` failure on Airflow 2.x. This issue files it as a bug, explains the mechanism, and shows why Airflow 3.x makes it far easier to hit. ### What happened `SSHOperator` tasks fail intermittently with: ``` ValueError: filedescriptor out of range in select() File ".../airflow/providers/ssh/hooks/ssh.py", line 470, in exec_ssh_client_command readq, _, _ = select([channel], [], [], cmd_timeout) ``` SSH connects and authenticates successfully and the remote command starts; the failure is purely in the output read loop. ### Root cause `SSHHook.exec_ssh_client_command` reads the command channel with stdlib `select.select()`: ```python from select import select ... readq, _, _ = select([channel], [], [], cmd_timeout) ``` `select.select()` is backed by a fixed-width `fd_set` bitmap of `FD_SETSIZE` (1024) bits. It cannot represent a file-descriptor whose number is >= 1024, and raises `ValueError: filedescriptor out of range` when asked to. This ceiling is compile-time and is unrelated to `RLIMIT_NOFILE` (`ulimit -n`) — raising the open-files limit does not help. Whether the error fires depends solely on the FD *number* assigned to the SSH channel, which depends on how many descriptors the task process already holds. The provider code has never changed, so this is not version-specific to the provider; it is triggered by the surrounding process holding a high FD count. ### Why Airflow 3.x makes this far easier to hit Under the Airflow 3.x Task Execution SDK (`apache-airflow-task-sdk`), the task process holds more open descriptors than the Airflow 2.x `LocalTaskJob` model did. Because descriptors are assigned lowest-free-first, those extra open FDs push the SSH channel's number higher, so even a **single** `SSHOperator` task can cross 1024. On Airflow 2.x the same task kept its FD numbers below 1024, so the identical line never tripped. This matches an earlier Airflow 2 report (#30085), where the user reached the same failure by inflating the process FD count via extreme config (`parallelism=1800`, `max_active_tasks_per_dag=200`, a ~350-task DAG, large dagbag settings). Same ceiling, same line — reached there by scale/config, reached here by the 3.x runtime's baseline FD footprint. ### Reproduction (official MWAA 3.0.6 image, provider ssh 4.1.3) A single task that (1) pre-opens dummy FDs in its own process to push the next FD number past 1024, then (2) calls `exec_ssh_client_command`: - High FD count (~1,146 open, highest FD 1144): SSH connects + authenticates, command starts, task fails with the exact error at `ssh/hooks/ssh.py` `exec_ssh_client_command`. - Low FD count (~100 open): the same task succeeds. - `RLIMIT_NOFILE` was 1,048,576 in **both** runs — confirming `ulimit` is unrelated. The only variable that flips pass/fail is the FD number assigned to the channel. Additional findings from the same harness: - `get_pty=True` vs `get_pty=False` makes no difference at a high FD count — both fail. `get_pty` only controls the remote PTY request; it does not change the local `select()` read path. (This refutes a commonly suggested "set get_pty=True" workaround.) - Running the same command via the system `ssh` client (which uses `poll()`), under the same high FD count, succeeds. ### Workaround We worked around this by replacing the `SSHOperator` with a `BashOperator` that invokes the system `ssh` client: ```python BashOperator( task_id="t01_CaPreSessionSQL", bash_command="ssh -o StrictHostKeyChecking=no -i /tmp/your_key.pem user@remote_host '/path/to/remote_script.ksh'", ) ``` The system `ssh` client uses `poll()` internally, so it has no `FD_SETSIZE` ceiling. We verified this on the same MWAA 3.0.6 image under the same high FD count that fails `SSHOperator` — it succeeds. This is a workaround, not a fix; the provider itself should not fail at high FD counts. ### Suggested fix Replace the `select.select()` read loop in `exec_ssh_client_command` with a `selectors`-based loop (`selectors.DefaultSelector`, i.e. `epoll`/`poll`), which has no `FD_SETSIZE` ceiling. The loop body is otherwise unchanged (same `recv`/`recv_stderr`/exit handling and per-line logging). Change the import `from select import select` to `import selectors`, then: ```python timedout = False selector = selectors.DefaultSelector() selector.register(channel, selectors.EVENT_READ) try: while not channel.closed or channel.recv_ready() or channel.recv_stderr_ready(): events = selector.select(timeout=cmd_timeout) if cmd_timeout is not None: timedout = not events for key, _ in events: recv = key.fileobj if recv.recv_ready(): output = stdout.channel.recv(len(recv.in_buffer)) agg_stdout += output for line in output.decode("utf-8", "replace").strip("\n").splitlines(): self.log.info(line) if recv.recv_stderr_ready(): output = stderr.channel.recv_stderr(len(recv.in_stderr_buffer)) agg_stderr += output for line in output.decode("utf-8", "replace").strip("\n").splitlines(): self.log.warning(line) if ( stdout.channel.exit_status_ready() and not stderr.channel.recv_stderr_ready() and not stdout.channel.recv_ready() ) or timedout: stdout.channel.shutdown_read() try: stdout.channel.close() except Exception: self.log.warning("Ignoring exception on close", exc_info=True) break finally: try: selector.unregister(channel) except (KeyError, ValueError): pass selector.close() ``` I validated this exact change against the real `apache-airflow-providers-ssh` 4.1.3 `SSHHook` inside the official MWAA 3.0.6 container: with the change applied, a channel whose fd number is >= 1024 no longer raises, stdout/stderr aggregation is unchanged, and the `cmd_timeout` path still raises `AirflowException` as before. I'm willing to submit the PR (with unit tests covering stdout/stderr aggregation, the high-fd case, and timeout parity). ### Related - Discussion #30085 (same error on Airflow 2.x, reached via high-concurrency config) ### What you think should happen instead? SSHOperator should read the command channel without a 1024 file-descriptor ceiling. The read loop in SSHHook.exec_ssh_client_command should use selectors (epoll/poll), which has no FD_SETSIZE limit, instead of select.select(). The task should run successfully regardless of the FD number assigned to the SSH channel. ### Operating System Amazon Linux (AWS MWAA 3.0.6 base image) ### Deployment None ### Apache Airflow Provider(s) _No response_ ### Versions of Apache Airflow Providers apache-airflow-providers-ssh==4.1.3 (the version pinned by the MWAA 3.0.6 constraints). The same select() read loop is present in every SSH provider release from 3.14.0 through 6.1.0 and on main. ### Official Helm Chart version Not Applicable ### Kubernetes Version _No response_ ### Helm Chart configuration _No response_ ### Docker Image customizations _No response_ ### Anything else? Reproduced on the official AWS MWAA 3.0.6 Docker image. The failure is intermittent in practice because the FD number assigned to the SSH channel varies run to run near the 1024 boundary. Related: Discussion #30085 (same error on Airflow 2.x) and Discussion #54335. ### Are you willing to submit PR? - [x] Yes I am willing to submit a PR! ### Code of Conduct - [x] I agree to follow this project's [Code of Conduct](https://github.com/apache/airflow/blob/main/CODE_OF_CONDUCT.md) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
