luc-pimentel opened a new issue, #74207:
URL: https://github.com/apache/airflow/issues/74207

   ### Under which category would you file this issue?
   
   Airflow Core
   
   ### Apache Airflow version
   
   main (as of 2026-10-04)
   
   ### What happened and how to reproduce it?
   
   With LocalExecutor, a task that fails and is retried makes the scheduler log 
a WARNING for the finished attempt, with no Dag or task identity:
   
   ```text
   [info   ] Received executor event with state success for task instance 
01a1085d-ee6c-... (coordinates=None)
   [warning] Discarding executor event for task instance 01a1085d-ee6c-... 
(coordinates=None): no matching task instance was returned; it may no longer 
exist or may be locked by another scheduler
   ```
   
   Neither clause applies: there is one scheduler, and the task instance exists 
with its next try number. Every "Received executor event" line, including those 
for tasks that succeed on the first try, also shows `coordinates=None`.
   
   ```python
   from datetime import timedelta
   
   from airflow.providers.standard.operators.bash import BashOperator
   from airflow.sdk import DAG
   
   with DAG(dag_id="retry_probe", schedule=None, catchup=False):
       BashOperator(
           task_id="fail_once",
           bash_command="if [ {{ ti.try_number }} -eq 1 ]; then exit 1; fi; 
echo ok",
           retries=1,
           retry_delay=timedelta(seconds=5),
       )
   ```
   
   Steps: `airflow standalone` (SQLite, LocalExecutor), trigger `retry_probe`, 
read the scheduler output. Try 1 fails, try 2 succeeds, and the run ends in 
success; the warning appears when the executor result of try 1 is processed.
   
   | Build | Received executor event | `coordinates=None` | Discarding WARNING |
   | --- | --- | --- | --- |
   | main | 3 (retry_probe try 1 and 2, plus a control Dag that succeeds) | 3 
of 3 | 1 (retry_probe try 1) |
   | 3.3.2 | 3 | n/a (keys are logged as `TaskInstanceKey(...)`) | 0, and no 
scheduler warnings at all |
   
   ### What you think should happen instead?
   
   A finished attempt that was replaced by a retry is the normal case, not a 
missing or locked row. #73916 added the `coordinates=` field and that warning 
so the retired attempt can still be identified, and its test asserts the 
coordinates appear in both lines. In practice they never do with LocalExecutor.
   
   The scheduler drains events every loop. `BaseExecutor.get_event_buffer` in 
`airflow-core/src/airflow/executors/base_executor.py` keeps the 
UUID-to-coordinates snapshot only while the key is in `running`, the task queue 
or the event buffer, and LocalExecutor never adds dispatched tasks to 
`running`, so the snapshot is dropped on the first drain after dispatch, before 
the worker result arrives. Since #73554 the retry allocates a new task instance 
UUID as soon as the failure is reported, so the result for the old UUID matches 
no row. On 3.3.2 the same situation matched the task instance and logged at 
INFO.
   
   Expected: either the retired attempt's coordinates are kept until its result 
is drained (so the line reads 
`coordinates=TaskInstanceKey(dag_id='retry_probe', ...)`), or an event for a 
known retired attempt is logged at INFO as before, or both.
   
   ### Operating System
   
   macOS 26.2 (host for 3.3.2); main run in Breeze (Linux container, Python 
3.10)
   
   ### Deployment
   
   Virtualenv installation
   
   ### Apache Airflow Provider(s)
   
   standard
   
   ### Versions of Apache Airflow Providers
   
   apache-airflow-providers-standard from source on main; 1.12.0 with 3.3.2
   
   ### Official Helm Chart version
   
   Not Applicable
   
   ### Kubernetes Version
   
   Not Applicable
   
   ### Helm Chart configuration
   
   Not Applicable
   
   ### Docker Image customizations
   
   Not Applicable
   
   ### Anything else?
   
   It happens every time. Two unit tests in 
`airflow-core/tests/unit/executors/test_local_executor.py` show the mechanism 
on main: after `queue_workload` and one `heartbeat` (with no worker spawned), 
`has_task` returns False for the dispatched task, and after one 
`get_event_buffer()` call followed by the worker result, 
`_drain_events_with_task_ids` returns no coordinates for it. I can attach them 
if useful. Executors that track `running` (Celery, Edge, Amazon, Kubernetes) 
should keep the coordinates; I did not run those.
   
   ---
   Drafted-by: Claude Code (Fable 5.1); reviewed by @luc-pimentel before posting
   
   ### Are you willing to submit PR?
   
   - [ ] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://github.com/apache/airflow/blob/main/CODE_OF_CONDUCT.md)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to