[
https://issues.apache.org/jira/browse/GROOVY-12318?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18109353#comment-18109353
]
ASF GitHub Bot commented on GROOVY-12318:
-----------------------------------------
paulk-asert opened a new pull request, #2844:
URL: https://github.com/apache/groovy/pull/2844
The shared parser DFA cache had no ceiling out of the box. The GC canary is
only observed on the parse path, so under sustained pressure the cache grows
unchecked between observations: a long-lived daemon parsing across many modules
ends up spending its time collecting rather than parsing. On a Grails groovydoc
build (12 cores, 5G heap) the groovydoc tasks took 1784s against 98s on the
last release without the regression, and at a tight heap the same growth
surfaces as OutOfMemoryError instead.
Bound the cache by what it actually holds. groovy.antlr4.cache.size gives a
ceiling on DFA states retained across the shared ATN, defaulting to 20000;
exceeding it drops the cache. Calibration: parsing 6849 real Groovy sources
with no ceiling grows the cache to ~179,000 states / ~13.5M ATNConfigs, and
states track configs at a stable ~1:75, so the default bounds it at roughly
1.5M configs. The Grails groovydoc tasks take 128s with this ceiling; a
multi-module parse at a constrained heap needs 44M where both 5.1.0 and 5.1.1
need 64-65M.
Size is the default ceiling rather than the existing parse counter because
it costs nothing while the cache is small: a project that never reaches the
limit never clears and keeps a fully warm cache. The counter clears on a fixed
cadence whether or not there is anything worth dropping, which costs ~87% on
parse-only workloads against ~13% for the ceiling, so it reverts to opt-in
(threshold default back to 0). Unlike the GC canary the ceiling does not depend
on SoftReference policy, so it behaves the same on SubstrateVM where that
policy differs; and unlike the cleaner thread removed by GROOVY-12142 it starts
nothing, so it cannot pin a container's class loader.
Only one thread clears per crossing, guarded by a CAS. Without the guard
every concurrent parser sees the same over-limit count and queues its own clear
on the fair write lock, turning a single crossing into a herd of
clear-and-rebuild cycles that blocks every reader; that churned the cache badly
enough to exhaust a 5G heap across parallel groovydoc tasks while passing every
single-threaded benchmark.
> bound the parser DFA cache by size
> ----------------------------------
>
> Key: GROOVY-12318
> URL: https://issues.apache.org/jira/browse/GROOVY-12318
> Project: Groovy
> Issue Type: Improvement
> Reporter: Paul King
> Priority: Major
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)