Published September 16, 2026 | Version 0.2.0

annotations4all

Authors/Creators

  • 1. Humboldt-Universität zu Berlin, NFDI4Memory

Description

annotations4all is a schema-first Python library for LLM-assisted span annotation and named entity recognition (NER). You define a tag schema, the library turns it into chat prompts for any OpenAI-compatible endpoint, and maps the annotated model response back to offset-based spans (label, start, end, optional meta).

It is deliberately a library, not a pipeline: it owns neither a fixed tag ontology nor a corpus nor a metric set. Tag schemas, prompts and material stay with the research workflow that uses the library.

What the library provides

  • Schema-first prompt taggers. ConfigurableTagger renders a user-defined tag schema (("TAG", "description") pairs) into system/user messages and parses the model response. Tag names and descriptions belong to the user; common NER labels such as PER, LOC and ORG are examples, not a canonical ontology.
  • Response parsing with visible errors. <<TAG>>…</TAG>> annotations (optionally <<TAG:meta>>…</TAG>>) are decoded by a forgiving stack parser that reports structured warnings for malformed, unclosed or unanchorable tags instead of failing — decoding problems stay distinguishable from model mistakes.
  • OpenAI-compatible transport. OpenAICompatClient is a narrow /v1/chat/completions client for any compatible server (streaming and non-streaming) with typed usage metadata and a generic extra_body passthrough for provider-specific request fields (e.g. reasoning options). The semantics of such fields are explicitly not part of the compatibility promise; the passthrough is.
  • Per-label mode and merge (new in 0.2.0). For nested or hierarchical schemas a single nested model answer is unreliable. The fallback annotates one label (or label group) per model call and merges the flat runs with merge_span_sets into one nested annotation set — deterministic, order-independent, and derived from containment only. Conflicting spans are reported, never silently repaired: duplicate (same label and bounds twice), intra-label-overlap (same label, partial overlap) and crossing (different labels, partial overlap).
  • Anchor rule. A span's parent is always the smallest span that strictly encloses it; when children are searched inside an already known parent span, the parser can be bounded to that window (bounds=(start, end)) so no global similarity search can re-anchor them elsewhere. Span bounds are half-open ([start, end)), following Prodigy/spaCy/standoff/TEI convention: touching spans do not overlap.

Scope and limits

  • Annotation, not evaluation. The metrics for flat (SemEval-style) and nested evaluation live in the sibling package annotations-evaluate, which derives hierarchies from containment itself and does not import this library; the shared contract is the written semantics specification, not shared code.
  • Cost of the nested fallback. The per-label mode costs one model call per run — with one label per run that is approximately the number of labels. For large hierarchical schemas it is therefore only sensible for the outer level; recursive prompting per detected element is the alternative. It also loses cross-label context, so cross-label consistency is enforced (and reported) by the merge instead of being hidden.
  • No fixed ontology, no corpus, no provider lock-in. The library enforces no taxonomy and bundles no corpus; the supported transport is one explicit OpenAI-compatible contract.
  • Legacy domain taggers for historical corpora remain importable but are unstable examples, not part of the stable surface.

Release notes 0.2.0

  • New merge_span_sets: pure, syntax-agnostic merge of flat span sets into a containment-derived nesting plus a conflict list (duplicate, intra-label-overlap, crossing), with flatten_merged/iter_merged helpers and the intra-label-overlap policies flag, keep-first and union.
  • New PerLabelTagger: one model call per label or label group (groups), annotate() returns the merged structure plus the raw per-run responses, spans and parser warnings, so a decoding failure does not look like "nothing found". Packaged prompt templates for German and English.
  • Anchor rule exposed in the parser: parse_region_response/parse_region_response_detailed accept bounds=(start, end) and then anchor annotations only inside that window, without the global fallback search.
  • Tag-schema validation moved to taggers/schema.py and shared by both prompt taggers; the package API additionally exports the merge and per-label symbols.
  • Test suite grown to 89 tests at 89.5 % statement coverage, including order-independence tests for the merge and a live smoke test that runs the complete loop (prompt → endpoint → parse → merge) against a configured OpenAI-compatible endpoint.
  • Documentation: tutorial section for the per-label mode and merge (costs and limits), updated German and English READMEs.

Installation, documentation, tests

  • pip install annotations4all — Python 3.11 or newer; the only runtime dependencies are an OpenAI-compatible client, a fuzzy matcher and typing extensions (an optional extra covers Ollama).
  • docs/tutorial.md documents the API with runnable examples against local, OpenAI-compatible servers (llama.cpp, Ollama).
  • Gates: ruff check, ruff format --check and pytest (coverage threshold 75 %).

Releases and versioning

Every release is deposited separately and carries its own version DOI; the concept DOI 10.5281/zenodo.22011370 always resolves to the latest version. This deposit is version 0.2.0 and supersedes 0.1.1 (0.1.0 was withdrawn from PyPI and is not a valid release).

License and citation

MIT. Machine-readable citation metadata (CITATION.cff) ships with the source archive. Source code, issues and release notes: scm.cms.hu-berlin.de/annotations4all/annotations4all.

Files

annotations4all-0.2.0.zip

Files (1.2 MB)

Name Size Download all
md5:4f6f8fd589d66f62be7197a1de9727e2
1.2 MB Preview Download

Additional details

Related works

Continues
Publication: arXiv:2502.04351 (arXiv)
Has part
Software: 10.5281/zenodo.22015763 (DOI)
Software: 10.5281/zenodo.22795382 (DOI)
Is identical to
Software: https://scm.cms.hu-berlin.de/annotations4all/annotations4all (URL)
Is referenced by
Poster: 10.5281/zenodo.18834190 (DOI)

Funding

Deutsche Forschungsgemeinschaft
NFDI consortium 4Memory 501609550

Software

Repository URL
https://scm.cms.hu-berlin.de/annotations4all/annotations4all
Programming language
Python
Development Status
Active