EhsanFarasat opened a new issue, #2586:
URL: https://github.com/apache/age/issues/2586

   **Describe the bug**
   In 1.8.0, every Cypher `DELETE` / `DETACH DELETE` statement costs about 1 ms 
per edge label in the graph, even when the statement matches nothing. The 
number of rows or edges doesn't change it. With 150 edge labels, a `DETACH 
DELETE` that matches no vertex takes ~165 ms on 1.8.0 and ~1 ms on 1.7.0. 
Reads, `CREATE` and `SET` are not affected.
   
   The cost is in `process_edges_by_index()`, which was added in #2351:
   
   * `check_for_connected_edges()` calls it twice for every edge label, once 
for the `start_id` pass and once for the `end_id` pass.
   * Each call walks `vertex_id_htab` with `hash_seq_init()` / 
`hash_seq_search()`.
   * That table is created with `DELETE_VERTEX_HTAB_SIZE` = 1,000,000, so every 
walk visits about 1M buckets, even when the table holds zero vertices or one.
   
   Rebuilding 1.8.0 with `DELETE_VERTEX_HTAB_SIZE` set to 1024 brings the 
150-label case from ~164 ms to ~0.55 ms (table below).
   
   * 
https://github.com/apache/age/blob/PG18/v1.8.0-rc0/src/backend/executor/cypher_delete.c#L577-L678
 (`process_edges_by_index`, `hash_seq_init` at L603)
   * 
https://github.com/apache/age/blob/PG18/v1.8.0-rc0/src/backend/executor/cypher_delete.c#L739-L744
 (two calls per edge label)
   * 
https://github.com/apache/age/blob/PG18/v1.8.0-rc0/src/include/executor/cypher_utils.h#L49
 (`DELETE_VERTEX_HTAB_SIZE 1000000`)
   
   **How are you accessing AGE (Command line, driver, etc.)?**
   - psql for the repro below. Our application uses Python (asyncpg).
   
   **What data setup do we need to do?**
   ```pgsql
   LOAD 'age';
   SET search_path = ag_catalog, "$user", public;
   
   SELECT create_graph('delete_repro');
   
   -- 150 edge labels; they can all be em
   DO $$
   BEGIN
     FOR i IN 1..150 LOOP
       PERFORM create_elabel('delete_repr
     END LOOP;
   END $$;
   
   SELECT * FROM cypher('delete_repro', $
     CREATE (:Person {name: 'a'})-[:E1]->(:Person {name: 'b'})
   $$) AS (a agtype);
   ```
   
   **What is the necessary configuration info needed?**
   - Nothing special. Reproduced with the_PG18_1.8.0` Docker image 
(PostgreSQL18.6) on default settings, and with AGE built from the 
`PG18/v1.8.0-rc0` tag.
   
   **What is the command that caused the error?**
   ```pgsql
   \timing on
   -- matches no vertex
   SELECT * FROM cypher('delete_repro', $$
     MATCH (n:Person {name: 'nobody'}) DE
   $$) AS (a agtype);
   -- deletes one vertex and its one edge
   SELECT * FROM cypher('delete_repro', $$
     MATCH (n:Person {name: 'a'}) DETACH
   $$) AS (a agtype);
   ```
   ```
   -- 1.8.0 (apache/age:release_PG18_1.8.
   Time: 166.609 ms
   Time: 165.141 ms
   
   -- 1.7.0 (apache/age:release_PG18_1.7.
   Time: 4.658 ms   (first Cypher statement in the session)
   Time: 1.325 ms
   ```
   
   **Expected behavior**
   A `DELETE` that matches nothing, or det get slower as the number of 
edgelabels grows. 1.7.0 does the same statements in about 1 ms with 150 edge 
labels.
   
   **Environment (please complete the following information):**
   - Version: 1.8.0 (`PG18/v1.8.0-rc0`) oion from 1.7.0 (`PG18/v1.7.0-rc0`).
   - Docker on Linux arm64, 4 vCPUs.
   - `process_edges_by_index()` and `DELEnchanged on `master` as of 2026-10-01.
   
   **Additional context**
   
   Average latency of a `DETACH DELETE` m, 50 runs each). All three builds 
arefrom source on the same PostgreSQL 18.6:
   
   | Edge labels | 1.7.0 | 1.8.0 | 1.8.0 with `DELETE_VERTEX_HTAB_SIZE 1024` |
   |---|---|---|---|
   | 10 | 0.19 ms | 11.7 ms | 0.10 ms |
   | 50 | 0.25 ms | 54.5 ms | 0.21 ms |
   | 150 | 1.19 ms | 163.7 ms | 0.55 ms |
   
   On 1.8.0, deleting one vertex with one edge, and a plain `MATCH ()-[e:E1 {w: 
'nope'}]->() DELETE e`
   matching nothing, show the same per-laabels).
   
   What we ruled out:
   - **PostgreSQL version:** 1.7.0 built on 18.6 performs the same as the stock 
1.7.0 image on 18.1.
   - **The `_age_cache_invalidate` trigge label table doesn't change the timing.
   - **The rest of the delete path:** dropping the `end_id` indexes on the edge 
label tables, which makes
   `check_for_connected_edges()` fall bacrings the same statement down to ~2 ms.
   
   Correctness with the smaller size: a ` with three edges in different 
labels,including a self-loop, removes exactly those edges, and unrelated edges 
stay. The local dynahash grows
   on demand, so the constant is only an
   
   Possible fixes, any of which would hel
   - Create `vertex_id_htab` with a small initial size and let it grow.
   - Skip `check_for_connected_edges()` wcss->vertex_id_htab) == 0`. This 
covers edge-only `DELETE` and deletes that match nothing.
   - Collect the vertex ids into an arrayrray in both index passes.
   
   Impact for us: our graph has ~120 edgeot ~130 ms slower after upgrading. 
Our672-test suite went from 16 s to 5 min 44 s. We're staying on 1.7.0 in 
production until this is fixed.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to