aru-trackunit opened a new issue, #73238:
URL: https://github.com/apache/airflow/issues/73238

   ### Under which category would you file this issue?
   
   Providers
   
   ### Apache Airflow version
   
   3.3.1
   
   ### What happened and how to reproduce it?
   
   Executor KubernetesExecutor(parallelism=60) reported that the task instance 
<TaskInstance: xxxx scheduled__2026-09-16T07:00:00+00:00 map_index=5 [queued] 
ti_id=01a0a95a-37ed-7177-b212-4e25272f5c2a> finished with state failed, but the 
task instance's state attribute is queued. Learn more: 
https://airflow.apache.org/docs/apache-airflow/stable/troubleshooting.html#task-state-changed-externally
 Extra info: Pod failed because of None [airflow.task] loc=taskinstance.py:1841"
   
   On the screenshot below it says that first attempt is not visible, and it 
succeeded performing only 2nd attempt.
   
   <img width="790" height="306" alt="Image" 
src="https://github.com/user-attachments/assets/10befa80-9662-4799-b0a6-a683115f43fe";
 />
   
   ### What you think should happen instead?
   
   It would be valuable to support different retry policies based on the cause 
of a failure, particularly for application failures versus infrastructure or 
connectivity failures.
   
   Task authors may want deterministic application errors to fail immediately, 
since rerunning the same code with the same input is unlikely to succeed. They 
may also need to disable retries when an operation is non-idempotent or 
otherwise unsafe to repeat.
   
   However, disabling retries entirely is risky. A task can fail before the 
operation even begins—for example, because its worker pod cannot start—or 
because of a transient network problem or temporary third-party service outage. 
These failures may be safe and worthwhile to retry.
   
   Conversely, applying one retry policy to every failure can cause 
deterministic application errors to run repeatedly when they should fail 
immediately.
   
   Ideally, Airflow would allow different behavior for at least the following 
categories:
   
   Application failures that should fail immediately
   Infrastructure or pre-execution failures that should be retried without 
consuming the application retry budget
   Transient connectivity or third-party failures that should follow a 
configurable retry policy
   
   I am not sure what the ideal implementation would look like, but 
representing application errors, infrastructure failures, and third-party 
timeouts with the same generic failed state makes it difficult to choose a safe 
and effective retry strategy.
   
   Apologies for covering several concerns in one ticket, but I believe they 
are closely related: without knowing why a task failed, Airflow cannot 
determine—or allow users to configure—the appropriate retry behavior.
   
   Related:
   https://github.com/apache/airflow/issues/73164
   https://github.com/apache/airflow/issues/69052
   https://github.com/apache/airflow/pull/66405
   
   ### Operating System
   
   Debian GNU/Linux 12 (bookworm)
   
   ### Deployment
   
   Official Apache Airflow Helm Chart
   
   ### Apache Airflow Provider(s)
   
   cncf-kubernetes
   
   ### Versions of Apache Airflow Providers
   
   ```
   apache-airflow-providers-amazon==9.34.0
   apache-airflow-providers-cncf-kubernetes==10.21.0
   apache-airflow-providers-common-compat==1.18.0
   apache-airflow-providers-common-io==1.8.0
   apache-airflow-providers-common-sql==2.1.0
   apache-airflow-providers-databricks==7.18.1
   apache-airflow-providers-fab==3.8.0
   apache-airflow-providers-github==2.11.3
   apache-airflow-providers-hashicorp==4.8.0
   apache-airflow-providers-http==6.0.5
   apache-airflow-providers-microsoft-mssql==4.7.0
   apache-airflow-providers-mysql==6.6.1
   apache-airflow-providers-postgres==7.0.1
   apache-airflow-providers-sftp==6.0.1
   apache-airflow-providers-slack==9.10.2
   apache-airflow-providers-smtp==3.0.3
   apache-airflow-providers-ssh==6.0.1
   apache-airflow-providers-standard==1.17.0
   ```
   
   ### Official Helm Chart version
   
   1.22.0 (latest released)
   
   ### Kubernetes Version
   
   1.34.9
   
   ### Helm Chart configuration
   
   _No response_
   
   ### Docker Image customizations
   
   _No response_
   
   ### Anything else?
   
   _No response_
   
   ### Are you willing to submit PR?
   
   - [ ] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://github.com/apache/airflow/blob/main/CODE_OF_CONDUCT.md)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to