viukpe opened a new issue, #74205:
URL: https://github.com/apache/airflow/issues/74205

   ### Under which category would you file this issue?
   
   Providers
   
   ### Apache Airflow version
   
   3.0.6
   
   ### What happened and how to reproduce it?
   
   ### Apache Airflow version
   
   Reproduced on Airflow 3.0.6 (official AWS MWAA 3.0.6 image, 
`apache-airflow-providers-ssh==4.1.3`).
   
   Not limited to that version: the faulty `select()` read loop in 
`SSHHook.exec_ssh_client_command` is byte-for-byte identical in every released 
SSH provider from 3.14.0 (shipped with Airflow 2.10) through the latest release 
6.1.0 (2026-09-29), and on `main`. It has never been changed. So the bug is 
present on the current latest provider and latest Airflow 3.x; it is not fixed 
anywhere.
   
   > Related prior report: [Discussion 
#30085](https://github.com/apache/airflow/discussions/30085) describes the same 
`select()` failure on Airflow 2.x. This issue files it as a bug, explains the 
mechanism, and shows why Airflow 3.x makes it far easier to hit.
   
   ### What happened
   
   `SSHOperator` tasks fail intermittently with:
   
   ```
   ValueError: filedescriptor out of range in select()
     File ".../airflow/providers/ssh/hooks/ssh.py", line 470, in 
exec_ssh_client_command
       readq, _, _ = select([channel], [], [], cmd_timeout)
   ```
   
   SSH connects and authenticates successfully and the remote command starts; 
the failure is purely in the output read loop.
   
   ### Root cause
   
   `SSHHook.exec_ssh_client_command` reads the command channel with stdlib 
`select.select()`:
   
   ```python
   from select import select
   ...
   readq, _, _ = select([channel], [], [], cmd_timeout)
   ```
   
   `select.select()` is backed by a fixed-width `fd_set` bitmap of `FD_SETSIZE` 
(1024) bits. It cannot represent a file-descriptor whose number is >= 1024, and 
raises `ValueError: filedescriptor out of range` when asked to. This ceiling is 
compile-time and is unrelated to `RLIMIT_NOFILE` (`ulimit -n`) — raising the 
open-files limit does not help.
   
   Whether the error fires depends solely on the FD *number* assigned to the 
SSH channel, which depends on how many descriptors the task process already 
holds. The provider code has never changed, so this is not version-specific to 
the provider; it is triggered by the surrounding process holding a high FD 
count.
   
   ### Why Airflow 3.x makes this far easier to hit
   
   Under the Airflow 3.x Task Execution SDK (`apache-airflow-task-sdk`), the 
task process holds more open descriptors than the Airflow 2.x `LocalTaskJob` 
model did. Because descriptors are assigned lowest-free-first, those extra open 
FDs push the SSH channel's number higher, so even a **single** `SSHOperator` 
task can cross 1024. On Airflow 2.x the same task kept its FD numbers below 
1024, so the identical line never tripped.
   
   This matches an earlier Airflow 2 report (#30085), where the user reached 
the same failure by inflating the process FD count via extreme config 
(`parallelism=1800`, `max_active_tasks_per_dag=200`, a ~350-task DAG, large 
dagbag settings). Same ceiling, same line — reached there by scale/config, 
reached here by the 3.x runtime's baseline FD footprint.
   
   ### Reproduction (official MWAA 3.0.6 image, provider ssh 4.1.3)
   
   A single task that (1) pre-opens dummy FDs in its own process to push the 
next FD number past 1024, then (2) calls `exec_ssh_client_command`:
   
   - High FD count (~1,146 open, highest FD 1144): SSH connects + 
authenticates, command starts, task fails with the exact error at 
`ssh/hooks/ssh.py` `exec_ssh_client_command`.
   - Low FD count (~100 open): the same task succeeds.
   - `RLIMIT_NOFILE` was 1,048,576 in **both** runs — confirming `ulimit` is 
unrelated.
   
   The only variable that flips pass/fail is the FD number assigned to the 
channel.
   
   Additional findings from the same harness:
   - `get_pty=True` vs `get_pty=False` makes no difference at a high FD count — 
both fail. `get_pty` only controls the remote PTY request; it does not change 
the local `select()` read path. (This refutes a commonly suggested "set 
get_pty=True" workaround.)
   - Running the same command via the system `ssh` client (which uses 
`poll()`), under the same high FD count, succeeds.
   
   ### Workaround
   
   We worked around this by replacing the `SSHOperator` with a `BashOperator` 
that invokes the system `ssh` client:
   
   ```python
   BashOperator(
       task_id="t01_CaPreSessionSQL",
       bash_command="ssh -o StrictHostKeyChecking=no -i /tmp/your_key.pem 
user@remote_host '/path/to/remote_script.ksh'",
   )
   ```
   
   The system `ssh` client uses `poll()` internally, so it has no `FD_SETSIZE` 
ceiling. We verified this on the same MWAA 3.0.6 image under the same high FD 
count that fails `SSHOperator` — it succeeds. This is a workaround, not a fix; 
the provider itself should not fail at high FD counts.
   
   ### Suggested fix
   
   Replace the `select.select()` read loop in `exec_ssh_client_command` with a 
`selectors`-based loop (`selectors.DefaultSelector`, i.e. `epoll`/`poll`), 
which has no `FD_SETSIZE` ceiling. The loop body is otherwise unchanged (same 
`recv`/`recv_stderr`/exit handling and per-line logging). Change the import 
`from select import select` to `import selectors`, then:
   
   ```python
   timedout = False
   selector = selectors.DefaultSelector()
   selector.register(channel, selectors.EVENT_READ)
   try:
       while not channel.closed or channel.recv_ready() or 
channel.recv_stderr_ready():
           events = selector.select(timeout=cmd_timeout)
           if cmd_timeout is not None:
               timedout = not events
           for key, _ in events:
               recv = key.fileobj
               if recv.recv_ready():
                   output = stdout.channel.recv(len(recv.in_buffer))
                   agg_stdout += output
                   for line in output.decode("utf-8", 
"replace").strip("\n").splitlines():
                       self.log.info(line)
               if recv.recv_stderr_ready():
                   output = 
stderr.channel.recv_stderr(len(recv.in_stderr_buffer))
                   agg_stderr += output
                   for line in output.decode("utf-8", 
"replace").strip("\n").splitlines():
                       self.log.warning(line)
           if (
               stdout.channel.exit_status_ready()
               and not stderr.channel.recv_stderr_ready()
               and not stdout.channel.recv_ready()
           ) or timedout:
               stdout.channel.shutdown_read()
               try:
                   stdout.channel.close()
               except Exception:
                   self.log.warning("Ignoring exception on close", 
exc_info=True)
               break
   finally:
       try:
           selector.unregister(channel)
       except (KeyError, ValueError):
           pass
       selector.close()
   ```
   
   I validated this exact change against the real 
`apache-airflow-providers-ssh` 4.1.3 `SSHHook` inside the official MWAA 3.0.6 
container: with the change applied, a channel whose fd number is >= 1024 no 
longer raises, stdout/stderr aggregation is unchanged, and the `cmd_timeout` 
path still raises `AirflowException` as before. I'm willing to submit the PR 
(with unit tests covering stdout/stderr aggregation, the high-fd case, and 
timeout parity).
   
   ### Related
   
   - Discussion #30085 (same error on Airflow 2.x, reached via high-concurrency 
config)
   
   
   
   ### What you think should happen instead?
   
   SSHOperator should read the command channel without a 1024 file-descriptor 
ceiling. The read loop in SSHHook.exec_ssh_client_command should use selectors 
(epoll/poll), which has no FD_SETSIZE limit, instead of select.select(). The 
task should run successfully regardless of the FD number assigned to the SSH 
channel.
   
   
   ### Operating System
   
   Amazon Linux (AWS MWAA 3.0.6 base image)
   
   ### Deployment
   
   None
   
   ### Apache Airflow Provider(s)
   
   _No response_
   
   ### Versions of Apache Airflow Providers
   
   apache-airflow-providers-ssh==4.1.3 (the version pinned by the MWAA 3.0.6 
constraints). The same select() read loop is present in every SSH provider 
release from 3.14.0 through 6.1.0 and on main.
   
   
   ### Official Helm Chart version
   
   Not Applicable
   
   ### Kubernetes Version
   
   _No response_
   
   ### Helm Chart configuration
   
   _No response_
   
   ### Docker Image customizations
   
   _No response_
   
   ### Anything else?
   
   Reproduced on the official AWS MWAA 3.0.6 Docker image. The failure is 
intermittent in practice because the FD number assigned to the SSH channel 
varies run to run near the 1024 boundary. Related: Discussion #30085 (same 
error on Airflow 2.x) and Discussion #54335.
   
   
   ### Are you willing to submit PR?
   
   - [x] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://github.com/apache/airflow/blob/main/CODE_OF_CONDUCT.md)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to