This is an automated email from the ASF dual-hosted git repository.
kou pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/arrow.git
The following commit(s) were added to refs/heads/main by this push:
new 4e61c97e3a GH-51011: [C++][CI] Poll for GCS testbench readiness
instead of one 10s attempt (#51012)
4e61c97e3a is described below
commit 4e61c97e3a3d000c999d2f7e3671d5fcd97153b2
Author: Jonas Dedden <[email protected]>
AuthorDate: Tue Sep 29 08:06:48 2026 +0300
GH-51011: [C++][CI] Poll for GCS testbench readiness instead of one 10s
attempt (#51012)
### Rationale for this change
`arrow-gcsfs-test` sometimes fails on CI with `Could not start GCS emulator
'storage-testbench' (failed to listen)`. The testbench does listen, it is just
slow to answer: it runs under Werkzeug's reloader, which binds the socket in
the parent and serves from a child that re-imports grpcio, protobuf and flask
first.
`GcsTestbench` passes its 10s startup budget to `LimitedTimeRetryPolicy` as
well, so the readiness loop only makes one attempt. And since stderr is
discarded, the CI log can't tell a slow start from a crash.
### What changes are included in this PR?
In `gcsfs_test.cc`:
* 5s per attempt, 60s overall, so the loop actually polls
* keep the testbench's stderr
* error message says `(did not become ready)` instead of `(failed to
listen)`
`util::Process::IgnoreStderr` is now unused; I can remove it if preferred.
### Are these changes tested?
By `arrow-gcsfs-test` itself. I couldn't reproduce the CI flake on demand,
so this widens the margin and makes the next failure diagnosable.
### Are there any user-facing changes?
No.
* GitHub Issue: #51011
Authored-by: Jonas Dedden <[email protected]>
Signed-off-by: Sutou Kouhei <[email protected]>
---
cpp/src/arrow/filesystem/gcsfs_test.cc | 11 +++++++----
1 file changed, 7 insertions(+), 4 deletions(-)
diff --git a/cpp/src/arrow/filesystem/gcsfs_test.cc
b/cpp/src/arrow/filesystem/gcsfs_test.cc
index e174638b53..a8637377aa 100644
--- a/cpp/src/arrow/filesystem/gcsfs_test.cc
+++ b/cpp/src/arrow/filesystem/gcsfs_test.cc
@@ -77,7 +77,7 @@ class GcsTestbench : public ::testing::Environment {
}
server_process->SetArgs({"--port", port_});
- server_process->IgnoreStderr();
+ // Keep stderr so that startup failures are visible in the test log.
status = server_process->Execute();
if (!status.ok()) {
error += ": " + status.ToString();
@@ -85,8 +85,11 @@ class GcsTestbench : public ::testing::Environment {
return;
}
+ // The testbench may accept connections well before it answers requests,
+ // so poll with a per-attempt budget shorter than the overall one.
auto testbench_is_running = [&server_process, this]() {
- auto ready_timeout = std::chrono::seconds(10);
+ auto ready_timeout = std::chrono::seconds(60);
+ auto attempt_timeout = std::chrono::seconds(5);
std::chrono::time_point<std::chrono::steady_clock> end =
std::chrono::steady_clock::now() + ready_timeout;
while (server_process->IsRunning() && std::chrono::steady_clock::now() <
end) {
@@ -95,7 +98,7 @@ class GcsTestbench : public ::testing::Environment {
.set<gcs::RestEndpointOption>("http://127.0.0.1:" + port_)
.set<gc::UnifiedCredentialsOption>(gc::MakeInsecureCredentials())
.set<gcs::RetryPolicyOption>(
- gcs::LimitedTimeRetryPolicy(ready_timeout).clone()));
+ gcs::LimitedTimeRetryPolicy(attempt_timeout).clone()));
auto metadata = client.GetBucketMetadata("nonexistent");
if (metadata.status().code() == google::cloud::StatusCode::kNotFound) {
return true;
@@ -105,7 +108,7 @@ class GcsTestbench : public ::testing::Environment {
};
if (!testbench_is_running()) {
- error += " (failed to listen)";
+ error += " (did not become ready)";
error_ = std::move(error);
return;
}