Dale Richardson created YUNIKORN-3370:
-----------------------------------------

             Summary: yunikorn-core services are not restartable in-process, so 
the shim cannot stop the core it starts without intermittent failures
                 Key: YUNIKORN-3370
                 URL: https://issues.apache.org/jira/browse/YUNIKORN-3370
             Project: Apache YuniKorn
          Issue Type: Bug
          Components: shim - kubernetes
            Reporter: Dale Richardson


Follow-up to YUNIKORN-3357 (surfaced by shim PR #1061). Blocks burning down the 
17 inherited core-service leakcheck exemptions in the shim.

yunikorn-core's service lifecycle is one-way: {{EventSystemImpl.Stop()}} niles 
its channel and early-returns on a {{stopped}} flag that is never cleared, and 
{{StartServiceWithPublisher}} starts its handler unconditionally, so a start 
after a stop leaks a handler; process-global config callbacks and 
{{UserGroupCache}} are likewise torn down and reused. Because of this the 
shim's {{MockScheduler.stop()}} cannot call {{coreContext.StopAll()}} to clean 
up the in-process core without introducing intermittent test failures.

Evidence: adding {{coreContext.StopAll()}} to {{MockScheduler.stop()}} produced 
a ~25% {{TestAssumePodError}} flake at {{-count>1}} (3/12 and 3/10, versus 0/35 
without it).

Proposed fix: make the core services restartable / their {{Stop()}} idempotent 
(clear {{{}stopped{}}}, guard the start, deregister process-global callbacks) 
so a component can be stopped and re-started in one process. Once done, the 
shim's 17 inherited core-service exemptions can be removed. Related core-side 
event-system defects are tracked in YUNIKORN-3363.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to