vakaobr opened a new issue, #2588:
URL: https://github.com/apache/age/issues/2588

   ### Summary
   
   Under sustained concurrent Cypher writes, AGE intermittently fails to 
resolve a label name and raises an undefined-table error for a relation whose 
name is **not a label at all**. The name is made of fragments of the data being 
written: agtype property keys, property values, and occasionally a single 
control byte equal to a label id.
   
   The graph's label catalog is correct and unchanged throughout, so this does 
not look like catalog corruption. It looks like the label name is being read 
through a pointer that no longer refers to the label.
   
   ### Version
   
   - Apache AGE **1.7.0** (`PG17/v1.7.0-rc0`)
   - PostgreSQL **17.10** (Debian 17.10-1.pgdg12+1)
   - Linux x86_64, AGE loaded via `shared_preload_libraries`
   
   ### What happens
   
   The error surfaces from a `CREATE` of an edge:
   
   ```
   ERROR:  relation "mygraph.\x01" does not exist
   ERROR:  relation "mygraph.<entity name fragment>" does not exist
   ERROR:  relation 
"mygraph.dfile_pathsource_idcreated_atdescriptionentity_type<entity 
name><filename><chunk id>" does not exist
   ERROR:  relation "mygraph.<text fragment><SEP><text fragment><binary bytes>" 
does not exist
   ```
   
   The third example is the clearest. `file_path`, `source_id`, `created_at`, 
`description` and `entity_type` are **the property keys of the edge being 
created**, concatenated, followed by a property value and an identifier. `\x01` 
is `chr(1)`, which is the `id` of `_ag_label_vertex` in `ag_catalog.ag_label` 
for this graph.
   
   So the resolver appears to be producing a name either from the raw label id 
or from memory adjacent to the agtype payload, instead of from the label 
catalog.
   
   ### The catalog is fine
   
   Queried during and after the failures, unchanged throughout:
   
   ```sql
   SELECT name, kind, id FROM ag_catalog.ag_label l
     JOIN ag_catalog.ag_graph g ON g.graphid = l.graph
    WHERE g.name = 'mygraph';
   
         name        | kind | id
   ------------------+------+----
    _ag_label_vertex | v    |  1
    _ag_label_edge   | e    |  2
    base             | v    |  3
    DIRECTED         | e    |  4
   ```
   
   Four labels, exactly as expected. `information_schema.tables` for the graph 
schema likewise shows four tables. No label named anything like the strings in 
the errors has ever existed.
   
   ### Workload shape
   
   The write is an idempotent edge upsert. Edge properties are inlined in the 
`CREATE` clause because `SET r += {...}` does not persist edge properties on 
AGE (the endpoint ids are parameterised):
   
   ```sql
   SELECT r FROM cypher('mygraph', $$
     MATCH (source:base {entity_id: $src_id})
     WITH source
     MATCH (target:base {entity_id: $tgt_id})
     WITH source, target
     OPTIONAL MATCH (source)-[old:DIRECTED]-(target)
     DELETE old
     WITH source, target
     CREATE (source)-[r:DIRECTED {`file_path`: "...", `source_id`: "...", 
`created_at`: 1790828551, `description`: "...", `entity_type`: "..."}]->(target)
     RETURN r
   $$, $1) AS (r agtype);
   ```
   
   Conditions when it occurs:
   
   - **3 concurrent writers** on separate pooled connections, each running the 
statement above in its own transaction
   - Graph at roughly **500,000 vertices and 1,400,000 edges**, and growing
   - `search_path` includes `ag_catalog`, set per connection
   - Transaction-scoped advisory locks serialise writers on the same endpoint 
pair, so two sessions do not write the same edge concurrently. Different edges 
do proceed concurrently.
   - Always on **high-degree endpoints**: the vertices involved are the most 
connected in the graph, appearing in a large share of all edges. Low-degree 
endpoints have not produced it.
   - Some property values are large: `source_id` can hold a few hundred 
separator-joined identifiers.
   
   ### Frequency
   
   Roughly **3 to 5 failures per 100 documents ingested**, where each document 
performs many edge upserts. It is not deterministic and not tied to any 
particular endpoint pair.
   
   It appeared only **late in a large ingest**, once the graph had grown to the 
size above. Early in the same ingest, with the same code and concurrency, it 
did not occur.
   
   ### What we ruled out
   
   - **Catalog corruption.** `ag_label` is correct before, during and after.
   - **A DDL storm.** Our client library was calling `create_graph()` on every 
pool checkout, around 31,000 failed calls an hour against an existing graph. We 
suspected the repeated failed DDL was invalidating cached label information and 
removed it entirely. The failure rate did **not** change, so this is not the 
trigger.
   - **A missing or renamed label.** The labels in the errors never existed.
   - **Our own string building.** The endpoint ids are bound parameters, and 
the property literal is JSON-escaped. The strings in the error are not a 
quoting artefact of ours; they are the contents of the properties we are 
legitimately writing, appearing where a label name should be.
   
   ### Impact
   
   Each occurrence aborts the transaction and fails the unit of work. In a 
long-running ingest this means permanent, silent data loss unless the caller 
retries, and the error is not distinguishable from a genuine schema problem by 
its type.
   
   ### Reproduction
   
   We do not have a minimal standalone reproduction. It needs a graph of 
substantial size and sustained concurrency before it appears, and we have only 
reproduced it incidentally on a live ingest. We are happy to gather more 
diagnostics on request, for example a stack trace with debug symbols or 
`pg_stat_activity` at the point of failure. Tell us what would be most useful 
and we will capture it.
   
   If it helps narrow things, the detail we find most suggestive is that the 
bad name contains the **property keys of the edge being created**, in order, 
which points at the label-name pointer landing inside the agtype payload buffer 
for that same statement.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to