vakaobr opened a new issue, #2588:
URL: https://github.com/apache/age/issues/2588
### Summary
Under sustained concurrent Cypher writes, AGE intermittently fails to
resolve a label name and raises an undefined-table error for a relation whose
name is **not a label at all**. The name is made of fragments of the data being
written: agtype property keys, property values, and occasionally a single
control byte equal to a label id.
The graph's label catalog is correct and unchanged throughout, so this does
not look like catalog corruption. It looks like the label name is being read
through a pointer that no longer refers to the label.
### Version
- Apache AGE **1.7.0** (`PG17/v1.7.0-rc0`)
- PostgreSQL **17.10** (Debian 17.10-1.pgdg12+1)
- Linux x86_64, AGE loaded via `shared_preload_libraries`
### What happens
The error surfaces from a `CREATE` of an edge:
```
ERROR: relation "mygraph.\x01" does not exist
ERROR: relation "mygraph.<entity name fragment>" does not exist
ERROR: relation
"mygraph.dfile_pathsource_idcreated_atdescriptionentity_type<entity
name><filename><chunk id>" does not exist
ERROR: relation "mygraph.<text fragment><SEP><text fragment><binary bytes>"
does not exist
```
The third example is the clearest. `file_path`, `source_id`, `created_at`,
`description` and `entity_type` are **the property keys of the edge being
created**, concatenated, followed by a property value and an identifier. `\x01`
is `chr(1)`, which is the `id` of `_ag_label_vertex` in `ag_catalog.ag_label`
for this graph.
So the resolver appears to be producing a name either from the raw label id
or from memory adjacent to the agtype payload, instead of from the label
catalog.
### The catalog is fine
Queried during and after the failures, unchanged throughout:
```sql
SELECT name, kind, id FROM ag_catalog.ag_label l
JOIN ag_catalog.ag_graph g ON g.graphid = l.graph
WHERE g.name = 'mygraph';
name | kind | id
------------------+------+----
_ag_label_vertex | v | 1
_ag_label_edge | e | 2
base | v | 3
DIRECTED | e | 4
```
Four labels, exactly as expected. `information_schema.tables` for the graph
schema likewise shows four tables. No label named anything like the strings in
the errors has ever existed.
### Workload shape
The write is an idempotent edge upsert. Edge properties are inlined in the
`CREATE` clause because `SET r += {...}` does not persist edge properties on
AGE (the endpoint ids are parameterised):
```sql
SELECT r FROM cypher('mygraph', $$
MATCH (source:base {entity_id: $src_id})
WITH source
MATCH (target:base {entity_id: $tgt_id})
WITH source, target
OPTIONAL MATCH (source)-[old:DIRECTED]-(target)
DELETE old
WITH source, target
CREATE (source)-[r:DIRECTED {`file_path`: "...", `source_id`: "...",
`created_at`: 1790828551, `description`: "...", `entity_type`: "..."}]->(target)
RETURN r
$$, $1) AS (r agtype);
```
Conditions when it occurs:
- **3 concurrent writers** on separate pooled connections, each running the
statement above in its own transaction
- Graph at roughly **500,000 vertices and 1,400,000 edges**, and growing
- `search_path` includes `ag_catalog`, set per connection
- Transaction-scoped advisory locks serialise writers on the same endpoint
pair, so two sessions do not write the same edge concurrently. Different edges
do proceed concurrently.
- Always on **high-degree endpoints**: the vertices involved are the most
connected in the graph, appearing in a large share of all edges. Low-degree
endpoints have not produced it.
- Some property values are large: `source_id` can hold a few hundred
separator-joined identifiers.
### Frequency
Roughly **3 to 5 failures per 100 documents ingested**, where each document
performs many edge upserts. It is not deterministic and not tied to any
particular endpoint pair.
It appeared only **late in a large ingest**, once the graph had grown to the
size above. Early in the same ingest, with the same code and concurrency, it
did not occur.
### What we ruled out
- **Catalog corruption.** `ag_label` is correct before, during and after.
- **A DDL storm.** Our client library was calling `create_graph()` on every
pool checkout, around 31,000 failed calls an hour against an existing graph. We
suspected the repeated failed DDL was invalidating cached label information and
removed it entirely. The failure rate did **not** change, so this is not the
trigger.
- **A missing or renamed label.** The labels in the errors never existed.
- **Our own string building.** The endpoint ids are bound parameters, and
the property literal is JSON-escaped. The strings in the error are not a
quoting artefact of ours; they are the contents of the properties we are
legitimately writing, appearing where a label name should be.
### Impact
Each occurrence aborts the transaction and fails the unit of work. In a
long-running ingest this means permanent, silent data loss unless the caller
retries, and the error is not distinguishable from a genuine schema problem by
its type.
### Reproduction
We do not have a minimal standalone reproduction. It needs a graph of
substantial size and sustained concurrency before it appears, and we have only
reproduced it incidentally on a live ingest. We are happy to gather more
diagnostics on request, for example a stack trace with debug symbols or
`pg_stat_activity` at the point of failure. Tell us what would be most useful
and we will capture it.
If it helps narrow things, the detail we find most suggestive is that the
bad name contains the **property keys of the edge being created**, in order,
which points at the label-name pointer landing inside the agtype payload buffer
for that same statement.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]