Messages by Thread
-
[I] Support anonymous S3 access through the CometS3CredentialProvider SPI [datafusion-comet]
via GitHub
-
[PR] fix: key each scan's planning data by its plan node in a native block [datafusion-comet]
via GitHub
-
[I] spark.comet.batchSize below 8192 makes CometConf fail to initialize on every executor [datafusion-comet]
via GitHub
-
[I] An input ArrowArrayStream that native never takes is never released [datafusion-comet]
via GitHub
-
Re: [I] Iceberg scan fails with a hard-wired 10s OpenDAL io_timeout that cannot be configured [datafusion-comet]
via GitHub
-
[I] Cancelling a task doesn't stop its native plan until the plan produces its next batch [datafusion-comet]
via GitHub
-
Re: [I] macOS aarch64 flake: SIGBUS in _pthread_tsd_cleanup after ParquetReadFromFakeHadoopFsSuite [datafusion-comet]
via GitHub
-
[I] S3 credential refreshes are not coalesced, and bucket region detection has no timeout [datafusion-comet]
via GitHub
-
Re: [I] Native Iceberg scan and write build a new FileIO, and a new storage client, for every task [datafusion-comet]
via GitHub
-
[I] A native plan with no JVM input returns truncated output without an error when the Tokio runtime shuts down [datafusion-comet]
via GitHub
-
[I] Blocking JVM calls from native plans hold Tokio workers, delaying other plans and I/O [datafusion-comet]
via GitHub
-
[I] Comet starts a single Tokio worker on standalone executors when spark.executor.cores is unset [datafusion-comet]
via GitHub
-
Re: [I] Delegate int/float/boolean to decimal cast arms to arrow safe cast [datafusion-comet]
via GitHub
-
Re: [I] Support GROUPS window frame units [datafusion-comet]
via GitHub
-
Re: [I] Upgrade to Spark 4.2.0-preview5 [datafusion-comet]
via GitHub
-
Re: [I] Explore some optimization ideas for JVM shuffle [datafusion-comet]
via GitHub
-
Re: [I] Triage the `spark.sql.legacy.*` configs dropped from the session-wide fallback in #4799 [datafusion-comet]
via GitHub
-
Re: [I] ci: next set of jobs to move from the PR tier to the merge queue tier [datafusion-comet]
via GitHub
-
Re: [I] Add a guard test that every CometNativeExec constructor parameter participates in equals [datafusion-comet]
via GitHub
-
Re: [I] `translate` falls back to Spark by default instead of using the codegen dispatcher like the other string functions [datafusion-comet]
via GitHub
-
Re: [I] `to_csv` never runs inside Comet by default, unlike `to_json` / `from_csv` / `schema_of_csv` [datafusion-comet]
via GitHub
-
Re: [I] Comet native broadcast fails under spark.kryo.registrationRequired=true [datafusion-comet]
via GitHub
-
Re: [I] Comet cache decode ignores column projection, making narrow reads of wide cached relations slower than Spark [datafusion-comet]
via GitHub
-
Re: [I] cast_map_to_map drops entries null buffer and ignores target sorted flag [datafusion-comet]
via GitHub
-
Re: [I] ANSI mode test coverage: re-enable stale ignored tests and add missing cases [datafusion-comet]
via GitHub
-
Re: [I] Cast from float/double to decimal should return NULL for NaN/Infinity under ANSI mode [datafusion-comet]
via GitHub
-
Re: [I] [EPIC] Optimize native scalar expressions used in TPC-DS [datafusion-comet]
via GitHub
-
Re: [I] Performance optimizations for native in-memory cache (follow-on to #4591) [datafusion-comet]
via GitHub
-
Re: [I] unix_timestamp support level disagrees with documented TimestampNTZ tz-conversion divergence [datafusion-comet]
via GitHub
-
Re: [I] Natively support time-window grouping expressions: window, session_window, window_time [datafusion-comet]
via GitHub
-
Re: [I] Deliberately opt eligible Incompatible expressions into codegen dispatch with test coverage [datafusion-comet]
via GitHub
-
Re: [I] [Doc] decode does not appear in auto-generated compatibility docs [datafusion-comet]
via GitHub
-
Re: [I] CometCaseConversionBase gates compat inside convert() instead of getSupportLevel [datafusion-comet]
via GitHub
-
Re: [I] Implement tiered CI approach [datafusion-comet]
via GitHub
-
Re: [I] Pending PR filter excludes PRs due to skipped iceberg CI checks [datafusion-comet]
via GitHub
-
Re: [I] Support higher-order array functions via JVM UDF bridge [datafusion-comet]
via GitHub
-
Re: [I] Add support for Spark 4.2.0-preview4 [datafusion-comet]
via GitHub
-
Re: [I] perf: extend field-major processing to nested struct fields [datafusion-comet]
via GitHub
-
Re: [I] Add compatible timezone support for date_format expression [datafusion-comet]
via GitHub
-
Re: [I] [Feature] Support Spark expression: timestamp_diff [datafusion-comet]
via GitHub
-
Re: [I] [Feature] Support Spark expression: timestamp_add [datafusion-comet]
via GitHub
-
Re: [I] [Feature] Support Spark expression: make_ym_interval [datafusion-comet]
via GitHub
-
Re: [I] Avoid `CometColumnarExchange` when next query stage requires `CometColumnarToRow` [datafusion-comet]
via GitHub
-
Re: [I] Review names of configuration settings [datafusion-comet]
via GitHub
-
Re: [I] [EPIC] Add support for all Map functions [datafusion-comet]
via GitHub
-
Re: [I] Add support for `COUNT(DISTINCT expr, expr1, ...)` [datafusion-comet]
via GitHub
-
Re: [I] [EPIC] Improve performance of TPC-DS queries [datafusion-comet]
via GitHub
-
Re: [PR] introduce optional rle reads from parquet [datafusion]
via GitHub
-
Re: [PR] feat: support Spark 4.1 TIME type and expressions via codegen dispatch [datafusion-comet]
via GitHub
-
[PR] fix: count buffers shared between sort input batches once [datafusion]
via GitHub
-
Re: [I] Bug triage results: 2026-09-21 [datafusion-comet]
via GitHub
-
Re: [I] Bug triage results: 2026-09-14 [datafusion-comet]
via GitHub
-
Re: [I] Bug triage results: 2026-08-17 [datafusion-comet]
via GitHub
-
Re: [I] Bug triage results: 2026-08-31 [datafusion-comet]
via GitHub
-
Re: [I] Bug triage results: 2026-08-11 [datafusion-comet]
via GitHub
-
Re: [I] Bug triage results: 2026-08-03 [datafusion-comet]
via GitHub
-
[PR] build(deps): bump github/codeql-action/analyze from 4.37.9 to 4.38.2 [datafusion-python]
via GitHub
-
[PR] feat: Prune ListingTable file groups for range-key filters [datafusion]
via GitHub
-
[PR] fix: respect table timestamp units in Parquet statistics [datafusion]
via GitHub
-
Re: [I] perf: bypass Arrow FFI for broadcast exchange reads [datafusion-comet]
via GitHub
-
[I] Accelerated mapInArrow reads Python output at the declared types without checking them [datafusion-comet]
via GitHub
-
Re: [I] Dropping a GlobalRef in a detached thread [datafusion-comet]
via GitHub
-
Re: [I] Native columnar-to-row conversion is much slower than the JVM implementation for small batches [datafusion-comet]
via GitHub
-
Re: [I] Allocate Comet's parquet reader buffers from ArrowUtils.rootAllocator to enable zero-copy PyArrow UDF runner [datafusion-comet]
via GitHub
-
Re: [I] Large-offset Arrow vectors from PyArrow UDFs cannot be serialized for broadcast or collect [datafusion-comet]
via GitHub
-
Re: [I] Optimize JVM columnar-to-row conversion [datafusion-comet]
via GitHub
-
Re: [I] [EPIC] Consistent handling of invalid UTF-8 in native StringType (ingress policy) [datafusion-comet]
via GitHub
-
[I] Remove AlignedArrowStreamReader and fix the Native to JVM section of ffi.md [datafusion-comet]
via GitHub
-
[I] Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results [datafusion-comet]
via GitHub
-
[PR] fix: validate spark.comet.shuffle.jvm.batchSize and spark.comet.exec.memoryPool [datafusion-comet]
via GitHub
-
Re: [PR] feat: Use NativeType in get_example_type, information schema [datafusion]
via GitHub
-
Re: [PR] feat: reuse repeated exchanges in the distributed planner (ReuseExchange analog) [datafusion-ballista]
via GitHub
-
Re: [I] Task input metrics are unreliable when a native block mixes a native scan with a JVM input [datafusion-comet]
via GitHub
-
Re: [PR] feat: add partition-aware metrics snapshots [datafusion]
via GitHub
-
[PR] fix: [branch-1.0] build the Spark 3.4 and 3.5 jars for Java 11 on any JDK [datafusion-comet]
via GitHub
-
[I] Discussion: move the expensive CI jobs to a merge queue, like Comet [datafusion-ballista]
via GitHub
-
Re: [PR] fix(executor): drain Flight before shutdown cleanup [datafusion-ballista]
via GitHub
-
Re: [PR] fix: Not all rows are accounted in RowCursorStream [datafusion]
via GitHub
-
[PR] ci: [branch-1.1] run every tier on release-branch pull requests, and the full suite before an RC (#6218) [datafusion-comet]
via GitHub
-
[I] Spark 3.4 and 3.5 release jars require Java 17 since 0.11.0, though the docs list Java 11 [datafusion-comet]
via GitHub
-
Re: [PR] feat: add `ExecutionPlan::as_probe_side` hook for CollectLeft join swaps [datafusion]
via GitHub
-
Re: [PR] fix(scheduler): read a swapped broadcast partitioned on the probe side [datafusion-ballista]
via GitHub
-
[I] Incorrect ORDER BY results with mismatched Parquet timestamp units [datafusion]
via GitHub
-
Re: [I] Negative `INTERVAL` offsets in RANGE window frames are accepted and panic at execution [datafusion]
via GitHub
-
[PR] fix: Reject negative RANGE window frame offsets [datafusion]
via GitHub
-
Re: [I] Native Iceberg write gate misses a custom location provider supplied by TableOperations [datafusion-comet]
via GitHub
-
Re: [PR] build(deps): bump github/codeql-action/analyze from 4.37.9 to 4.38.0 [datafusion-python]
via GitHub
-
[PR] build(deps): bump astral-sh/setup-uv from 10.0.1 to 10.2.0 [datafusion-python]
via GitHub
-
Re: [PR] build(deps): bump astral-sh/setup-uv from 10.0.1 to 10.1.0 [datafusion-python]
via GitHub
-
[PR] build(deps): bump taiki-e/install-action from 2.85.5 to 2.87.20 [datafusion-python]
via GitHub
-
Re: [PR] build(deps): bump taiki-e/install-action from 2.85.5 to 2.87.14 [datafusion-python]
via GitHub
-
Re: [PR] build(deps): bump github/codeql-action/init from 4.37.9 to 4.38.0 [datafusion-python]
via GitHub
-
[PR] build(deps): bump github/codeql-action/init from 4.37.9 to 4.38.2 [datafusion-python]
via GitHub
-
[PR] chore: Add 1.1.0 changelog [WIP] [datafusion-comet]
via GitHub
-
Re: [I] Add AQE to DataFusion [datafusion]
via GitHub
-
[PR] feat(physical-plan): add staged execution boundary contract [datafusion]
via GitHub
-
[I] TPC-H SF1000 results for Ballista main (DataFusion 55.1.0): 1.45x of Spark + Comet, 15% of time in planning (#2497) [datafusion-ballista]
via GitHub
-
[PR] perf(scheduler): share the file statistics cache across sessions [datafusion-ballista]
via GitHub