[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17945733#comment-17945733
]
ASF subversion and git services commented on SOLR-17497:
Commit 849ec5b7c52eb6bf43c2b5b69eb8b2c0d55ab416 in solr's branch
refs/heads/add-pki-caching from Sanjay Dutt
[ https://gitbox.apache.org/repos/asf?p=solr.git;h=849ec5b7c52 ]
SOLR-17448 SOLR-17497: IndexFetcher, catch exception instead of bubbling up
uncaught (#2800)
(cherry picked from commit cc30093c5ee988555389b50cf2333edf743bb50f)
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
>Reporter: Sanjay Dutt
>Assignee: Sanjay Dutt
>Priority: Major
> Fix For: main (10.0), 9.8
>
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has b
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17923454#comment-17923454
]
Houston Putman commented on SOLR-17497:
---
Should we resolve this with a fix version?
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has been observed, that after restarting the PULL replica. The
> recovery process started and after fetching all the files info from the NRT,
> the replication aborted and logged "User aborted replication"
>
> {code:java}
> o.a.s.h.IndexFetcher User aborted Replication =>
> org.apache.solr.handler.IndexFetcher$ReplicationHandlerException: User
> aborted replication at
> org.apache.solr.handler.IndexFetch
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17895334#comment-17895334
]
ASF subversion and git services commented on SOLR-17497:
Commit 4c211ba5be43c63d3b9a7ccf5810b4aa73960e80 in solr's branch
refs/heads/branch_9_7 from Sanjay Dutt
[ https://gitbox.apache.org/repos/asf?p=solr.git;h=4c211ba5be4 ]
SOLR-17448 SOLR-17497: IndexFetcher, catch exception instead of bubbling up
uncaught (#2800)
(cherry picked from commit cc30093c5ee988555389b50cf2333edf743bb50f)
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has been observed, that after restarting the PULL replica. The
> recovery process star
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17895320#comment-17895320
]
Jason Gerlowski commented on SOLR-17497:
Hey [~sanjaydutt] - I'm going to backport your change above to branch_9_7 as
well, unless you've got any objections? Was doing some "beast" runs this
morning and noticed a big improvement in branch_9x with your change, so I'd
love see branch_9_7 get that benefit as well with a 9.7.1 release coming up...
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has been observed, that after restarting the PULL replica. The
> recovery process started and after fetching all the files info from the NRT,
>
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17894647#comment-17894647
]
ASF subversion and git services commented on SOLR-17497:
Commit 9a9d922088c982508f17e537f5db7bf407a1b037 in solr's branch
refs/heads/branch_9x from Sanjay Dutt
[ https://gitbox.apache.org/repos/asf?p=solr.git;h=9a9d922088c ]
SOLR-17448 SOLR-17497: IndexFetcher, catch exception instead of bubbling up
uncaught (#2800)
(cherry picked from commit cc30093c5ee988555389b50cf2333edf743bb50f)
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has been observed, that after restarting the PULL replica. The
> recovery process start
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17894641#comment-17894641
]
ASF subversion and git services commented on SOLR-17497:
Commit cc30093c5ee988555389b50cf2333edf743bb50f in solr's branch
refs/heads/main from Sanjay Dutt
[ https://gitbox.apache.org/repos/asf?p=solr.git;h=cc30093c5ee ]
SOLR-17448 SOLR-17497: IndexFetcher, catch exception instead of bubbling up
uncaught (#2800)
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has been observed, that after restarting the PULL replica. The
> recovery process started and after fetching all the files info from the NRT,
> the replication
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17893601#comment-17893601
]
David Smiley commented on SOLR-17497:
-
FYI since SOLR-17448 isn't released yet, any fixes can have commits referencing
that issue.
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has been observed, that after restarting the PULL replica. The
> recovery process started and after fetching all the files info from the NRT,
> the replication aborted and logged "User aborted replication"
>
> {code:java}
> o.a.s.h.IndexFetcher User aborted Replication =>
> org.apache.solr.handler.IndexFetcher$ReplicationHandlerException: User
> aborted replic
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17893025#comment-17893025
]
Sanjay Dutt commented on SOLR-17497:
Sorry, Initially I had no idea what's going on so i shared whatever I can found
here in this JIRA. There is only one exception that is relevant –
AlreadyClosedException. The other one "User aborted Replication" is expected
and observed whenever the replication is aborted. Even when you run
org.apache.solr.cloud.TestPullReplica.testKillPullReplica, you will see this
exception in the logs and that's fine IMO.
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has been observed, that after restarting th
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17893024#comment-17893024
]
David Smiley commented on SOLR-17497:
-
I'm confused; is this one JIRA issue about two different exception?
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has been observed, that after restarting the PULL replica. The
> recovery process started and after fetching all the files info from the NRT,
> the replication aborted and logged "User aborted replication"
>
> {code:java}
> o.a.s.h.IndexFetcher User aborted Replication =>
> org.apache.solr.handler.IndexFetcher$ReplicationHandlerException: User
> aborted replication at
> org.apache.so
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17892798#comment-17892798
]
Sanjay Dutt commented on SOLR-17497:
Yeah you are right, I have to look into this subject (execute vs submit) more
and how this whole things works.
+1 to this.
{quote}For the case here in IndexFetcher, as long as the exception that is
thrown is logged, I think we should suppress its propagation further.
{quote}
I was also looking why we are getting "User aborted replication" messages.
RecoveryStrategy in case of PULL replicas cancel the replication. Here is the
explanation from the old JIRA.
https://issues.apache.org/jira/browse/SOLR-10233
{quote}
h3. Passive replica dies (or is unreachable)
Replica won’t be query-able. On restart, replica will recover from the leader,
following the same flow as _realtime_ replicas: set state to DOWN, then
RECOVERING, and finally ACTIVE. _Passive_ replicas will use a different
{{RecoveryStrategy}} implementation, that omits *preparerecovery,* and peer
sync attempt, it will jump to replication . If the leader didn't change, or if
the other replicas are of type “append”, replication should be incremental.
Once the first replication is done, passive replica will declare itself active
and start serving traffic.
{quote}
*RecoveryStrategy.java*
{noformat}
log.info("Stopping background replicate from leader process");
zkController.stopReplicationFromLeader(coreName);
replicate(zkController.getNodeName(), core, leaderprops);{noformat}
My own theory:
# RecoveryStrategy cancel replication.
# FileFetcher#fetchPackets throws ReplicationHandlerException
{code:java}
if (stop) {
stop = false;
aborted = true;
throw new ReplicationHandlerException("User aborted replication");
}{code}
# FileFetcher#fetch runs finally block where the sync is executed in async
{code:java}
fsyncService.submit(() -> {
try {
file.sync();
} catch (IOException e) {
fsyncException = e;
} catch (InterruptedException e) {
throw new RuntimeException(e);
}
});{code}
# At the same time the control gets back to fetchLatestIndex that performs
cleanup and closed the directory
{code:java}
finally {
if (!cleanupDone) {
cleanup(solrCore, tmpIndexDir, indexDir, deleteTmpIdxDir, tmpTlogDir,
successfulInstall);
}
}{code}
And basically there is race condition between step 3 and 4 that's what I
believe. Not able to reproduce on my system yet.
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
> Security Level: Public(Default Security Level. Issues are Public)
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17892780#comment-17892780
]
David Smiley commented on SOLR-17497:
-
bq. execute throws exception rather than suppressing it
Maybe I'm nitpicking but execute() definitely isn't throwing the exception;
it's impossible that it could even do so since the Runnable that does throw an
exception happens asynchronously after execute() returns. The change from
before is that the thrown exception (from the Runnable) is no longer captured
into a Future; it bubbles up to the Thread uncaughtExceptionHandler where our
test infrastructure notices it and reports it via
com.carrotsearch.randomizedtesting.UncaughtExceptionError. CC [~andreybozhko].
For the case here in IndexFetcher, as long as the exception that is thrown is
logged, I think we should suppress its propagation further.
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
> Security Level: Public(Default Security Level. Issues are Public)
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17892684#comment-17892684
]
Sanjay Dutt commented on SOLR-17497:
{code:java}
@Test
public void test(){
ExecutorService fsyncService =
ExecutorUtil.newMDCAwareSingleThreadExecutor(new
SolrNamedThreadFactory("fsyncService"));
try {
fsyncService.submit(() -> {
throw new AlreadyClosedException("Directory is already closed!");
});
} catch (Exception e) {
System.out.println(e);
} finally {
fsyncService.shutdown();
}
}{code}
In [https://github.com/apache/solr/pull/2707], we have basically replaced
ExecutorService#submit with ExecutorService#execute, and now execute throws
exception rather than suppressing it. Same can be tested with the above example
where running it won't fail, on the other hand If you use execute it will fail
immediately.
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
> Security Level: Public(Default Security Level. Issues are Public)
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17892144#comment-17892144
]
Sanjay Dutt commented on SOLR-17497:
!Screenshot 2024-10-23 at 6.01.02 PM.png!
One more thing, the spike in flaky tests starts from October 13th, which
initially led me to look for commits around that date. However, I later
realized I was looking in the wrong place. The issue had actually started
earlier. We were facing problems with the Jenkins node between October 2nd and
12th, and during that period, we weren't running builds as frequently as usual.
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
> Security Level: Public(Default Security Level. Issues are Public)
>Reporter: Sanjay Dutt
>Priority: Major
> Attachments: Screenshot 2024-10-23 at 6.01.02 PM.png
>
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where i
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17892135#comment-17892135
]
David Smiley commented on SOLR-17497:
-
That's an astute observation/theory! I guess, give it a shot and see if you're
right. Hopefully you can reliably reproduce the bug.
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
> Security Level: Public(Default Security Level. Issues are Public)
>Reporter: Sanjay Dutt
>Priority: Major
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=javabin&version=2} rid=null-11348 hits=2 status=0 QTime=0
> 616607 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node4
> (https://127.0.0.1:38207/solr/pull_replica_test_kill_pull_replica_shard1_replica_p2/)
> has all 2 docs{code}
>
> *Where is the issue then?*
> In the logs it has been observed, that after restarting the PULL replica. The
> recovery process started and after fetching all the files info from the NRT,
> the replication aborted and logged "User aborted replication"
>
> {code:java}
> o.a.s.h.IndexFetcher User aborted Replication =>
> org.apache.solr.handler.IndexFetcher$
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17892132#comment-17892132
]
Sanjay Dutt commented on SOLR-17497:
I wonder If this is related to this [https://github.com/apache/solr/pull/2707]
Because in above PR below code got changed in IndexFetcher
{code:java}
fsyncService.execute(
() -> {
try {
file.sync();
} catch (IOException e) {
fsyncException = e;
}
});
}{code}
The error that we are currently seeing is
org.apache.lucene.store.AlreadyClosedException sub class of Runtime exception
and thus catch block cannot handle it.
Also yes that it could be very much possible that this block has been throwing
exception for long time even before this change, however because of this code
change now those are not suppressed anymore. There are more test case case
failing with a same message "Capture an uncaught exception"
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
> Security Level: Public(Default Security Level. Issues are Public)
>Reporter: Sanjay Dutt
>Priority: Major
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Request webapp=/solr path=/select
> params={q=*:*&wt=java
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17891110#comment-17891110
]
Sanjay Dutt commented on SOLR-17497:
FAILED:
org.apache.solr.cloud.api.collections.TestLocalFSCloudBackupRestore.test
{code:java}
Error Message:
com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an uncaught
exception in thread: Thread[id=11418, name=fsyncService-7131-thread-1,
state=RUNNABLE, group=TGRP-TestLocalFSCloudBackupRestore]
Stack Trace:
com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an uncaught
exception in thread: Thread[id=11418, name=fsyncService-7131-thread-1,
state=RUNNABLE, group=TGRP-TestLocalFSCloudBackupRestore]
at
__randomizedtesting.SeedInfo.seed([1B41BE625EF5CDC:89E0243C8B133124]:0)
Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
closed
at __randomizedtesting.SeedInfo.seed([1B41BE625EF5CDC]:0)
at
app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50){code}
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
> Security Level: Public(Default Security Level. Issues are Public)
>Reporter: Sanjay Dutt
>Priority: Major
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
> (link may stop working in the future)
>
> Last step where they are making sure that all the active replicas should have
> two documents each has logged a info which is another proof that it completed
> successfully.
>
> {code:java}
> 616575 INFO
> (TEST-TestPullReplica.testKillPullReplica-seed#[F30CC837FDD0DC28]) [n: c: s:
> r: x: t:] o.a.s.c.TestPullReplica Replica core_node3
> (https://127.0.0.1:35647/solr/pull_replica_test_kill_pull_replica_shard1_replica_n1/)
> has all 2 docs 616606 INFO (qtp1091538342-13057-null-11348)
> [n:127.0.0.1:38207_solr c:pull_replica_test_kill_pull_replica s:shard1
> r:core_node4 x:pull_replica_test_kill_pull_replica_shard1_replica_p2
> t:null-11348] o.a.s.c.S.Re
[jira] [Commented] (SOLR-17497) Pull replicas throws AlreadyClosedException
[
https://issues.apache.org/jira/browse/SOLR-17497?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17891109#comment-17891109
]
Sanjay Dutt commented on SOLR-17497:
More test cases failing with same exception, and again it's somewhat similar
configuration one shard, one NRT and two PULL.
FAILED:
org.apache.solr.cloud.TestPullReplicaWithAuth.testPKIAuthWorksForPullReplication
{code:java}
Stack Trace:
com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an uncaught
exception in thread: Thread[id=9567, name=fsyncService-5965-thread-1,
state=RUNNABLE, group=TGRP-TestPullReplicaWithAuth]
at
__randomizedtesting.SeedInfo.seed([1B41BE625EF5CDC:548EE8CAC23F01EA]:0)
Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
closed
at __randomizedtesting.SeedInfo.seed([1B41BE625EF5CDC]:0)
at
app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
at
app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
at
app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
at
app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
at
app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
at
app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$0(ExecutorUtil.java:380)
at
[email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
at
[email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
at [email protected]/java.lang.Thread.run(Thread.java:829){code}
> Pull replicas throws AlreadyClosedException
> -
>
> Key: SOLR-17497
> URL: https://issues.apache.org/jira/browse/SOLR-17497
> Project: Solr
> Issue Type: Task
> Security Level: Public(Default Security Level. Issues are Public)
>Reporter: Sanjay Dutt
>Priority: Major
>
> Recently, a common exception (org.apache.lucene.store.AlreadyClosedException:
> this Directory is closed) seen in multiple failed test cases.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
> FAILED:
> org.apache.solr.cloud.SplitShardWithNodeRoleTest.testSolrClusterWithNodeRoleWithPull
> FAILED: org.apache.solr.cloud.TestPullReplica.testAddDocs
>
>
> {code:java}
> com.carrotsearch.randomizedtesting.UncaughtExceptionError: Captured an
> uncaught exception in thread: Thread[id=10271,
> name=fsyncService-6341-thread-1, state=RUNNABLE,
> group=TGRP-SplitShardWithNodeRoleTest]
> at
> __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4:E5DB3E97188A8EB9]:0)
> Caused by: org.apache.lucene.store.AlreadyClosedException: this Directory is
> closed
> at __randomizedtesting.SeedInfo.seed([3F7DACB3BC44C3C4]:0)
> at
> app//org.apache.lucene.store.BaseDirectory.ensureOpen(BaseDirectory.java:50)
> at
> app//org.apache.lucene.store.ByteBuffersDirectory.sync(ByteBuffersDirectory.java:237)
> at
> app//org.apache.lucene.tests.store.MockDirectoryWrapper.sync(MockDirectoryWrapper.java:214)
> at
> app//org.apache.solr.handler.IndexFetcher$DirectoryFile.sync(IndexFetcher.java:2034)
> at
> app//org.apache.solr.handler.IndexFetcher$FileFetcher.lambda$fetch$0(IndexFetcher.java:1803)
> at
> app//org.apache.solr.common.util.ExecutorUtil$MDCAwareThreadPoolExecutor.lambda$execute$1(ExecutorUtil.java:449)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
> at
> [email protected]/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
> at [email protected]/java.lang.Thread.run(Thread.java:829)
> {code}
>
> Interesting thing about these test cases is that they all share same kind of
> setup where each has one shard and two replicas – one NRT and another is PULL.
>
> Going through one of the test case execution step.
> FAILED: org.apache.solr.cloud.TestPullReplica.testKillPullReplica
>
> Test flow
> 1. Create a collection with 1 NRT and 1 PULL replica
> 2. waitForState
> 3. waitForNumDocsInAllActiveReplicas(0); // *Name says it all*
> 4. Index another document.
> 5. waitForNumDocsInAllActiveReplicas(1);
> 6. Stop Pull replica
> 7. Index another document
> 8. waitForNumDocsInAllActiveReplicas(2);
> 9. Start Pull Replica
> 10. waitForState
> 11. waitForNumDocsInAllActiveReplicas(2);
>
> As per the logs the whole sequence executed successfully. Here is the link to
> the logs:
> [https://ge.apache.org/s/yxydiox3gvlf2/tests/task/:solr:core:test/details/org.apache.solr.cloud.TestPullReplica/testKillPullReplica/1/output]
