KaoWeiCheng opened a new issue, #8667:
URL: https://github.com/apache/hop/issues/8667

   ## Apache Hop version?
   
   2.19.0 (the same code is still on `main` as of 2026-09-29)
   
   ## Java version?
   
   OpenJDK 21.0.12
   
   ## Operating system
   
   Docker
   
   ## What happened?
   
   When a running pipeline is stopped (Hop GUI stop, `hop/stopPipeline`, 
`hop/stopWorkflow`, ...), `TableInput.stopRunning()`
   is supposed to cancel the running statement via `data.db.cancelQuery()`. In 
practice this code is never reached:
   
   `Pipeline.stopTransform()` marks the transform as stopped **before** calling 
`stopRunning()`:
   
   ```java
   // engine/src/main/java/org/apache/hop/pipeline/Pipeline.java
   public void stopTransform(TransformMetaDataCombi combi, boolean safeStop) {
     ITransform rt = combi.transform;
     rt.setStopped(true);          // <-- 1
     rt.setSafeStopped(safeStop);
     rt.resumeRunning();
     try {
       rt.stopRunning();           // <-- 2
     } ...
   ```
   
   and `TableInput.stopRunning()` returns early when the transform is already 
stopped:
   
   ```java
   // plugins/transforms/tableinput/.../TableInput.java
   public synchronized void stopRunning() throws HopException {
     if (this.isStopped() || data.isDisposed()) {
       return;                     // always taken when called from 
Pipeline.stopTransform()
     }
     setStopped(true);
     if (data.db != null && data.db.getConnection() != null && 
!data.isCanceled) {
       data.db.cancelQuery();      // never reached
       data.isCanceled = true;
     }
   }
   ```
   
   `DatabaseJoin.stopRunning()` has the same `isStopped()` guard, so it looks 
affected too (by reading the code; I did not test it).
   For comparison, `DatabaseLookup`, `DynamicSqlRow`, `ExecSqlRow` and 
`ExecSql` do not have this guard and do call `cancelQuery()`.
   
   **Effect:** the transform stops emitting rows right away, but `dispose()` 
then calls `ResultSet.close()` on a result set that
   has only been partly read. With the Microsoft SQL Server JDBC driver, 
closing the result set reads (and throws away) all of
   the remaining rows before it returns. So:
   
   - the query keeps running on the database until the whole result set has 
been sent, and stopping saves no database load;
   - the pipeline stays in `Halting` until that finishes, and only then becomes 
`Stopped`.
   
   **Measurements** (SQL Server, mssql-jdbc 12.8.1 — the version shipped in 
`lib/jdbc` — one `SELECT` returning ~1,020,000 rows):
   
   | | stop → query gone from `sys.dm_exec_requests` |
   |---|---|
   | Hop 2.19.0 pipeline, `stopWorkflow` ~4 s in (Table Input → Text File 
Output) | 3.2 s; `Halting` → `Stopped` took 3.0–4.5 s over 3 runs |
   | Plain JDBC, read N rows then `rs.close()` (what Hop does now), N = 20k / 
200k / 700k | 5.7 s / 5.1 s / 2.3 s |
   | Plain JDBC, read N rows then `stmt.cancel()` + `rs.close()` (what 
`stopRunning()` intends), same N | 31 ms / 203 ms / 31 ms |
   
   While Hop was `Halting`, the query's request in `sys.dm_exec_requests` kept 
switching between `ASYNC_NETWORK_IO` and
   `runnable`, which means the client was still reading rows.
   
   I only tested SQL Server. Other drivers that read the rest of the result set 
on close (or that cannot close a streaming
   result set early) are probably affected in the same way; drivers that close 
cheaply would only lose the server-side cancel.
   
   **Steps to reproduce**
   
   1. Create a pipeline: Table Input (`SELECT` returning a few million rows 
from any large table) → Dummy (or Text File Output).
   2. Run it (Hop GUI, or Hop Server via `hop/asyncRun` / `hop/startPipeline`).
   3. Stop it a couple of seconds after rows start flowing.
   4. Look at the database's active sessions (SQL Server: 
`sys.dm_exec_requests`; PostgreSQL: `pg_stat_activity`): the
      `SELECT` is still running and only goes away after the full result set 
has been read. The pipeline stays in
      `Halting` until then.
   
   **Expected:** the statement is cancelled (`Statement.cancel()`) when the 
pipeline is stopped, and the pipeline reaches
   `Stopped` within a moment, as for Database Lookup / Exec SQL.
   
   **Suggested fix** (I checked this with a locally built `TableInput.class`: 
the bytecode shows `stopRunning()` now reaches
   `cancelQuery()`. I did not deploy it to Hop Server to measure the result end 
to end):
   
   ```java
   public synchronized void stopRunning() throws HopException {
     if (data.isDisposed()) {
       return;
     }
     setStopped(true);
     if (data.db != null && data.db.getConnection() != null && 
!data.isCanceled) {
       data.db.cancelQuery();
       data.isCanceled = true;
     }
   }
   ```
   
   The same change applies to `DatabaseJoin.stopRunning()`. `isCanceled` 
already prevents a second cancel.
   Another option is for `Pipeline.stopTransform()` to call `stopRunning()` 
before `setStopped(true)`, but that affects every
   transform, so the local fix looks safer.
   
   ## Issue Priority
   
   Priority: 2
   
   ## Issue Component
   
   Component: Transforms
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to