maiven gateway · online··
manifest · 2d ago
Sign in Back to catalog
documents
source 2 PII columns
Docling-extracted documents — the index over our unstructured corpora.
Definition
One row per source binary (PDF / image / scan). Holds the extracted text, the extractor + pipeline that produced it (VLM vs EasyOCR), document language, page count, and a Presidio entity-count summary. The text is searchable via the chunks table downstream; this table is the catalog of what we have indexed.
11 of 11
Column
Description
Tests · PII
document_id
UUID5 derived from (source_uri, sha256). Stable across re-extracts of the same byte-identical file.
document +4 PII3 tests
source_uri
gs:// URI of the source binary. Used by the eraser to delete the underlying object when the data subject requests erasure.
1 testPII
sha256
SHA-256 of the source binary. Idempotency key for the extractor.
1 test
mime
MIME type guessed from blob content-type or filename suffix.
doc_lang
ISO-2 lowercased language code. Currently always `en` (heuristic in extractor).
pages
Page count when available from Docling, else derived from chunk page_end.
extracted_text
Full markdown export of the DoclingDocument. Carries layout cues for downstream embedding.
PII
extracted_at
ISO-8601 UTC timestamp the extractor ran. Parsed to DateTime64(3) in staging.
1 test
extractor
Extractor library + version (e.g. `docling@2.14.0`).
pipeline
Which Docling pipeline ran — `vlm` (GPU) or `ocr-easyocr` (CPU fallback).
pii_scan
JSON object — `{engine, entities}`. `entities` is a map of entity-type → count; the actual values are never stored.