One row per (extraction-day × corpus × pipeline × doc_lang) summarising the unstructured zone. Use this metric to answer "how many documents do we have per corpus?", "is the VLM path producing as many chunks as EasyOCR?", "which corpus is surfacing the most distinct PII entity types?" — without touching chunk text. For the actual text you want, call `search_documents` instead.
- Reading `chunk_text_bytes_total` as a quality signal
`chunk_text_bytes_total` sums `length(chunk_text)` across the included chunks — it's a corpus-SIZE proxy. A long invoice line-items table inflates it without saying anything about extraction quality. Use `avg_chunks_per_document` and the `documents_with_chunks` / `documents_count` fraction for quality signals.
- Treating `pii_entity_kinds_total` as a PII occurrence count
`pii_entity_kinds_total` SUMs distinct-entity-type counts per document, not occurrences. Two documents each tagged `{PERSON, LOCATION}` contribute 4, not 2 — the *kinds* are counted, not the entities. For an occurrence-level view you would need an entity-level mart, which is deliberately not exposed here.
- Confusing `documents_with_chunks` with data-loss
`documents_with_chunks` < `documents_count` means the extractor wrote a documents-row but no chunks for some files — typically an empty-template PDF or an extraction regression. It is an extraction-quality signal, not a data-loss signal. The raw zone still has the file.
- “How many documents do we have in each corpus today?”
- “Which corpus produced the most chunks per document on average?”
- “Are we extracting more pages via vlm or ocr-easyocr?”
- “Which corpus is surfacing the most distinct PII entity types?”