Wondering if anyone is intrested in looking further into this.

Once we push this feature it unblocks a lot of work gated for a gRPC server.  I 
really think this can be the biggest win for a 3.0 version.

I followed rzo1's advice too.  He steered me to some industry standards I was 
unaware of. That motivated me to check out a few other standards.  I've also 
since battle tested this against 1000s of documents and it's still fast.  

But would love any thoughts or suggestions on this - 

On 2026/07/21 12:06:57 Martin Wiesner wrote:
> Got no capacity (for this) atm.
> 
> Best
> Martin
> 
> > Am 21.07.2026 um 11:18 schrieb Richard Zowalla <[email protected]>:
> > 
> > Any additional input?
> > 
> > On 2026/07/17 13:55:56 Kristian Rickert wrote:
> >> Everyone:
> >> 
> >> The draft PR (#1182) now reflects all 4 points, one commit per point so
> >> each is reviewable on its own.
> >> 
> >> tl;dr - Everything promised in the last mail is now spec text plus
> >> enforcing code plus contract tests, nothing is prose-only.
> >> 
> >> 1. Namespaced ids (86f4bee). Every key the toolkit defines carries the
> >> "opennlp:" prefix: opennlp:sentences, opennlp:tokens, opennlp:pos,
> >> opennlp:entities, and opennlp:lemmas / opennlp:stems on the adapters.
> >> Extensions use their own prefix, bare ids stay legal for application-local
> >> layers. The rule is stated on Layers, on LayerKey, and in the manual
> >> chapter.
> >> 2. Per-key scope (6c38c23). LayerKey now declares POSITIONAL or DOCUMENT.
> >> Positional is the default and guarantees a span on every annotation;
> >> LayerKey.document(id, type) creates a key for whole-document values
> >> (language id, category distributions, provenance), and Annotation.of(value)
> >> builds the span-less entry. The container polices it in both directions at
> >> insertion, naming the layer: a span-less annotation under a positional key
> >> is rejected, a spanned one under a document key likewise. One consequence
> >> you should know before commenting: key equality is now (id, type, scope),
> >> so a positional and a document key with the same id and type are distinct
> >> layers, same as the existing id-with-different-type rule.
> >> 3.  Invariants (a657c00). Transcribed onto the Document interface and into
> >> the manual: insertion order is preserved and never reordered, layers are
> >> immutable once added and detached from the caller's input list, add is
> >> once-only operation that rejects duplicates by naming the key, therefore
> >> index references stay valid for the document's lifetime. The wording
> >> matches what the contract tests enforce, nothing new was invented for the
> >> prose.
> >> 4. Gold versus predicted (48321ff).  Documented as the convention on Layers
> >> and in the manual: gold:opennlp:tokens beside opennlp:tokens, and a
> >> contract test pins that both live independently and neither can replace the
> >> other.
> >> 
> >> Thoughts are encouraged for the draft "opennlp:" prefix itself (versus
> >> something dotted) and the gold:opennlp:tokens double-prefix shape.  I
> >> prefer the ":" as the "." convention feels less of a separation.
> >> 
> >> This is a great moment to discuss, as it'll establish a new convention.
> >> 
> >> - Kristian
> >> 
> >> On Thu, Jul 16, 2026 at 1:58 PM Kristian Rickert <[email protected]> 
> >> wrote:
> >> 
> >>> Hi Richard,
> >>> 
> >>> Thanks for doing the survey before forming an opinion; the convergence you
> >>> found is the best argument the design could get. I am adopting all four
> >>> points. Short version first:
> >>> 
> >>> tl;dr - Spans are baked in by construction. The only span-less things are
> >>> whole-document facts, and those are declared at the key level, not skipped
> >>> at return time.
> >>> 
> >>> 1. Key identity. Equality is the (id, type) pair, pinned by contract
> >>> tests: same id and type address one layer, same id with a different type 
> >>> is
> >>> a distinct coexisting layer. Standard keys get namespaced ids before the
> >>> freeze ("opennlp:" prefix, extensions use their own prefix, bare ids stay
> >>> legal for application-local layers).
> >>> 
> >>> 2. Annotations without spans. The span is structurally mandatory today:
> >>> the annotation record rejects a null span and the container rejects
> >>> out-of-bounds spans, so no code path returns a span-less result. The 
> >>> census
> >>> of every annotator built so far:
> >>> 
> >>> 
> >>> +---------------------------------+--------------------------------------+
> >>>  | Annotator                       | Span anchors on
> >>> |
> >>> 
> >>> +---------------------------------+--------------------------------------+
> >>>  | Sentence / tokenizer adapters   | detected spans; tokens shifted to
> >>> |
> >>>  |                                 | document coordinates
> >>> |
> >>>  | POS / lemmatizer / stemmer      | each token's span
> >>> |
> >>>  | Name finder                     | the entity span
> >>> |
> >>>  | Glossary                        | the match span
> >>> |
> >>>  | PII                             | the mention span
> >>> |
> >>>  | Coreference (chains)            | each member mention's span
> >>> |
> >>>  | Money / quantity / temporal /   | the mention span
> >>> |
> >>>  |   currency conversion           |
> >>> |
> >>>  | Document date                   | the electing mention's span
> >>> |
> >>>  | Geocoding (locations)           | the resolved mention's span
> >>> |
> >>>  | Containment (hierarchy)         | the location mention's span
> >>> |
> >>>  | Dependencies                    | the dependent token's span
> >>> |
> >>>  | Relations                       | the covering span of both arguments
> >>> |
> >>>  | Embeddings                      | the source annotation's span
> >>> |
> >>> 
> >>> +---------------------------------+--------------------------------------+
> >>> 
> >>> Only one layer spans the document as a whole (the region ballot, currently
> >>> on a whole-document span). Proposal: scope is declared per key, never per
> >>> annotation. Positional keys guarantee a span on every entry;
> >>> document-scoped keys carry values without spans (language id, categories,
> >>> provenance, the ballot). That is the UIMA offset-bearing versus 
> >>> offset-free
> >>> split you cited, at library weight, and document-scoped layers staying
> >>> lists fits the real payloads (a categorizer returns a scored 
> >>> distribution).
> >>> 
> >>> 3. Invariants. All three are implemented and contract-tested, and the spec
> >>> will state them: insertion order is preserved and never sorted; layers are
> >>> immutable once added and detached from the caller's input; providing an
> >>> existing layer is rejected loudly, add is once-only. Index references
> >>> (dependencies point at token positions) therefore stay valid for the
> >>> document's lifetime.
> >>> 
> >>> 4. Gold versus predicted. Documented as a convention on the same namespace
> >>> scheme (gold:opennlp:tokens beside opennlp:tokens); with the once-only
> >>> rule, competing versions always live under distinct keys and never replace
> >>> each other.
> >>> 
> >>> Acceptance criterion adopted, and it is already an observation rather than
> >>> a promise: the feature branches on the PoC fork add all of the layers 
> >>> above
> >>> without a single edit to the container package.
> >>> 
> >>> Given this feedback, mind if I fold points 1 through 4 into the spec text?
> >>> Not finalizing, just updating.
> >>> 
> >>> Rungs,
> >>> Kristian
> >>> 
> >>> On Thu, Jul 16, 2026 at 1:22 PM Richard Zowalla <[email protected]> wrote:
> >>> 
> >>>> Thanks for the summary. I did some homework on how other systems solve
> >>>> this before forming an opinion. Short version: the proposed design is not
> >>>> novel architecture. That is its strength. The same pattern has emerged
> >>>> independently in at least four mature systems.
> >>>> 
> >>>> INCEpTION defines no container of its own. It sits on the UIMA CAS
> >>>> [1][2]. The CAS holds the original text and typed feature structures with
> >>>> begin and end offsets into it. That is classic standoff annotation, the
> >>>> same rule the proposal states for spans. What INCEpTION calls a layer is 
> >>>> a
> >>>> UIMA type plus configurable behaviors. It distinguishes span layers,
> >>>> relation layers, chain layers for coreference, and metadata layers 
> >>>> without
> >>>> offsets. The difference is weight. UIMA declares its type system in XML
> >>>> descriptors and carries reflection and serialization machinery with it. 
> >>>> The
> >>>> proposal is deliberately the lightweight end of that spectrum. For a
> >>>> library, as opposed to an annotation platform, that is the right 
> >>>> decision,
> >>>> IMHO.
> >>>> 
> >>>> Stanford CoreNLP is the closest match (to us), almost line for line. Its
> >>>> Annotation is a type safe heterogeneous map in the sense of Bloch
> >>>> [3][4][5]. Keys are class objects parameterized by the value type. The
> >>>> container never enumerates keys. Any module adds a new key without 
> >>>> touching
> >>>> the core. The proposed LayerKey is the same pattern with string identity
> >>>> instead of class identity. That design has held up in Standford's CoreNLP
> >>>> for over fifteen years.
> >>>> 
> >>>> spaCy has one owning Doc object, spans as views into it, and doc.spans as
> >>>> a dict of named span groups under open string keys [6]. Pipeline 
> >>>> components
> >>>> declare assigns and requires metadata, and the pipeline analyzer 
> >>>> validates
> >>>> ordering. That is a direct precedent for the proposed requires and 
> >>>> provides
> >>>> methods. GATE uses named annotation sets with untyped feature maps [7]. 
> >>>> It
> >>>> is open ended like the proposal but stringly typed in its values, which 
> >>>> the
> >>>> Class parameter in LayerKey correctly avoids.
> >>>> 
> >>>> So the field has converged four times on the same shape: standoff spans
> >>>> over immutable original text, open keyed typed layers, a container 
> >>>> ignorant
> >>>> of payloads, and pipelines ordered by declared dependencies.
> >>>> The guardrail note in the ticket, that the doc container must never learn
> >>>> about specific layers, is the single most important sentence in it.
> >>>> 
> >>>> Four points where the precedents diverge and where I would tighten the
> >>>> spec before the freeze:
> >>>> 
> >>>> 1. Key identity. CoreNLP gets collision free identity from class objects.
> >>>> String ids do not. Two extensions can both mint "entities". The spec 
> >>>> should
> >>>> define key equality and adopt namespaced ids for the standard keys, IMHO.
> >>>> 2. Annotations without spans. Language id, document categories,
> >>>> provenance. Every precedent grew a home for these. UIMA has types without
> >>>> offsets, spaCy has doc level attributes, INCEpTION has metadata layers.
> >>>> Retrofitting this later creates exactly the parallel surface the proposal
> >>>> wants to prevent, so we should think about that upfront.
> >>>> 3. Index based references need stated invariants. UIMA and CoreNLP
> >>>> reference annotations by object identity. Indices are better for
> >>>> immutability and serialization, but only if layer order is defined and
> >>>> existing layers are never mutated or reordered. The spec should -
> >>>> therefore- also say what happens when an annotator provides a layer that
> >>>> already exists. This needs to be defined.
> >>>> 4. Gold versus predicted layers: GATE annotation sets and UIMA views
> >>>> exist because corpora might carry competing versions of the same layer.
> >>>> Open string keys already handle this. The convention should simply be
> >>>> documented.
> >>>> 
> >>>> At a first glance, the validatio plan from the Jira looks right.
> >>>> Dependency parsing is precisely the consumer that stresses cross layer
> >>>> references, and coreference will inherit the same rules.
> >>>> I would also make the acceptance criterion, that a new layer can be added
> >>>> without touching the container package.
> >>>> 
> >>>> Gruß
> >>>> Richard
> >>>> 
> >>>> [1] INCEpTION Developer Guide.
> >>>> https://inception-project.github.io/releases/26.8/docs/developer-guide.html
> >>>> [2] Klie JC, Bugert M, Boullosa B, Eckart de Castilho R, Gurevych I. The
> >>>> INCEpTION Platform: Machine Assisted and Knowledge Oriented Interactive
> >>>> Annotation. COLING 2018. https://aclanthology.org/C18-2002.pdf
> >>>> [3] TypesafeMap. Stanford CoreNLP API.
> >>>> https://nlp.stanford.edu/nlp/javadoc/javanlp-3.5.0/edu/stanford/nlp/util/TypesafeMap.html
> >>>> [4] Annotation. Stanford CoreNLP API.
> >>>> https://nlp.stanford.edu/nlp/javadoc/javanlp-3.5.0/edu/stanford/nlp/pipeline/Annotation.html
> >>>> [5] CoreAnnotation. Stanford CoreNLP API.
> >>>> https://nlp.stanford.edu/nlp/javadoc/javanlp-3.5.0/edu/stanford/nlp/ling/CoreAnnotation.html
> >>>> [6] spaCy Doc API. https://spacy.io/api/doc
> >>>> [7] GATE Documentation, Language Resources: Corpora, Documents and
> >>>> Annotations. https://gate.ac.uk/sale/tao/splitch5.html
> >>>> 
> >>>>> Am 15.07.2026 um 21:33 schrieb Kristian Rickert <[email protected]>:
> >>>>> 
> >>>>> I accidentally sent a draft. To clear things up, I’ve opened
> >>>> OPENNLP-1888
> >>>>> to spark a design discussion around introducing a generic Document type
> >>>>> into the 3.0 API.
> >>>>> 
> >>>>> Currently, our Java API is tool-by-tool; as we look toward adding more
> >>>>> complex, structured results (like dependency arcs or relations),
> >>>> continuing
> >>>>> this pattern means we'll end up building parallel surface areas for
> >>>> every
> >>>>> new capability.
> >>>>> 
> >>>>> The proposal in the JIRA ticket aims to solve this by introducing a
> >>>> small,
> >>>>> unified container in opennlp-api that uses typed annotation layers over
> >>>> the
> >>>>> original text. I want to emphasize right away that this deprecates and
> >>>>> replaces absolutely nothing in the existing task-level API; it simply
> >>>>> provides a clean, additive way to handle combined-document views moving
> >>>>> forward. I'd really love it if we could take a look at the full
> >>>>> architecture, constraints, and motivations in the ticket and share your
> >>>>> thoughts.
> >>>>> 
> >>>>> I think this design covers what you'd see in other APIs as well as some
> >>>> new
> >>>>> frontier ideas.
> >>>>> 
> >>>>> Thoughts?
> >>>>> 
> >>>>> On Wed, Jul 15, 2026 at 3:17 PM Richard Zowalla <[email protected]>
> >>>> wrote:
> >>>>> 
> >>>>>> Did this mail get mangled?
> >>>>>> 
> >>>>>>> Am 15.07.2026 um 00:17 schrieb Kristian Rickert <[email protected]
> >>>>> :
> >>>>>>> 
> >>>>>>> Two main points:
> >>>>>>> 
> >>>>>>> - OPENNLP-1888 has not been started; it was created as a design
> >>>>>>> discussion regarding a generic document shape I'm recommending.
> >>>> Details
> >>>>>> in
> >>>>>>> the ticket
> >>>>>>> - Thank you everyone for the detailed feedback. I'll ensure every PR
> >>>> has
> >>>>>>> a section in the documentation.
> >>>>>>> 
> >>>>>>> OPENNLP-1888 was opened to discuss a generic Document type for the new
> >>>>>> API
> >>>>>>> interface. Details are in the ticket, but here is the initial shape
> >>>> and I
> >>>>>>> would like anyone's thoughts on it:
> >>>>>>> 
> >>>>>>> public interface Document {
> >>>>>>> CharSequence text();                            // the original text,
> >>>>>>> never a normalized form
> >>>>>>> <T> List<Annotation<T>> get(LayerKey<T> layer); // empty list when the
> >>>>>>> layer is absent
> >>>>>>> Set<LayerKey<?>> layers();
> >>>>>>> }
> >>>>>>> 
> >>>>>>> public record Annotation<T>(Span span, T value) { }
> >>>>>>> 
> >>>>>>> public final class LayerKey<T> {
> >>>>>>> public static <T> LayerKey<T> of(String id, Class<T> type) { ... }
> >>>>>>> }
> >>>>>>> 
> >>>>>>> public interface DocumentAnnotator {
> >>>>>>> Document annotate(Document document);           // returns a new
> >>>>>> document
> >>>>>>> with layers added
> >>>>>>> Set<LayerKey<?>> requires();
> >>>>>>> Set<LayerKey<?>> provides();
> >>>>>>> }
> >>>>>> 
> >>>> 
> >>>> 
> >> 
> 
> 

Reply via email to