andygrove opened a new pull request, #6453: URL: https://github.com/apache/datafusion-comet/pull/6453
## Which issue does this PR close? Closes #6452. ## Rationale for this change The cache scan's plan string dumped its `CachedRDDBuilder`, and with it the whole cached plan, into the middle of every plan that read the cache, breaking `EXPLAIN` and `EXPLAIN FORMATTED`. The issue has an example. ## What changes are included in this PR? `CometInMemoryTableScanExec` overrides `stringArgs` to print the Spark `InMemoryTableScanExec` it replaces, the way other Comet operators leave their `originalPlan` out. That shows the table's name when it has one, the attributes read and any pruning predicates, for example `CometInMemoryTableScan Scan In-memory table named_t [k#1L]`. Spark's own scan also lists the `InMemoryRelation` as an inner child, so the cached plan is drawn as an indented subtree below it. This PR does not do that, because `ExtendedExplainInfo` walks `innerChildren`, so the cached plan's operators and fallback reasons would then count toward every query that reads the cache. That can be decided separately. ## How are these changes tested? A new test in `CometInMemoryCacheSuite` checks the scan's line, including its pruning predicates, and checks that neither the tree string nor `EXPLAIN FORMATTED` prints the `CachedRDDBuilder` or the serializer. It fails without the override. `CometInMemoryCacheSuite`, `CometInMemoryCacheKryoSuite` and `CometInMemoryCachePruningSuite` pass on Spark 3.4 and 4.1. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
