Dale Richardson created YUNIKORN-3370:
-----------------------------------------
Summary: yunikorn-core services are not restartable in-process, so
the shim cannot stop the core it starts without intermittent failures
Key: YUNIKORN-3370
URL: https://issues.apache.org/jira/browse/YUNIKORN-3370
Project: Apache YuniKorn
Issue Type: Bug
Components: shim - kubernetes
Reporter: Dale Richardson
Follow-up to YUNIKORN-3357 (surfaced by shim PR #1061). Blocks burning down the
17 inherited core-service leakcheck exemptions in the shim.
yunikorn-core's service lifecycle is one-way: {{EventSystemImpl.Stop()}} niles
its channel and early-returns on a {{stopped}} flag that is never cleared, and
{{StartServiceWithPublisher}} starts its handler unconditionally, so a start
after a stop leaks a handler; process-global config callbacks and
{{UserGroupCache}} are likewise torn down and reused. Because of this the
shim's {{MockScheduler.stop()}} cannot call {{coreContext.StopAll()}} to clean
up the in-process core without introducing intermittent test failures.
Evidence: adding {{coreContext.StopAll()}} to {{MockScheduler.stop()}} produced
a ~25% {{TestAssumePodError}} flake at {{-count>1}} (3/12 and 3/10, versus 0/35
without it).
Proposed fix: make the core services restartable / their {{Stop()}} idempotent
(clear {{{}stopped{}}}, guard the start, deregister process-global callbacks)
so a component can be stopped and re-started in one process. Once done, the
shim's 17 inherited core-service exemptions can be removed. Related core-side
event-system defects are tracked in YUNIKORN-3363.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]