maiven gateway · online··
manifest · 1d ago
Sign in
Back to catalog
MAIVENmodelmetrics

metric_documents_by_corpus

beta
metric table

How big is each unstructured corpus, and what did Docling find?

Definition

One row per (extraction-day × corpus × pipeline × doc_lang) summarising the unstructured zone. Use this metric to answer "how many documents do we have per corpus?", "is the VLM path producing as many chunks as EasyOCR?", "which corpus is surfacing the most distinct PII entity types?" — without touching chunk text. For the actual text you want, call `search_documents` instead.

Watch out for
  • Reading `chunk_text_bytes_total` as a quality signal

    `chunk_text_bytes_total` sums `length(chunk_text)` across the included chunks — it's a corpus-SIZE proxy. A long invoice line-items table inflates it without saying anything about extraction quality. Use `avg_chunks_per_document` and the `documents_with_chunks` / `documents_count` fraction for quality signals.

  • Treating `pii_entity_kinds_total` as a PII occurrence count

    `pii_entity_kinds_total` SUMs distinct-entity-type counts per document, not occurrences. Two documents each tagged `{PERSON, LOCATION}` contribute 4, not 2 — the *kinds* are counted, not the entities. For an occurrence-level view you would need an entity-level mart, which is deliberately not exposed here.

  • Confusing `documents_with_chunks` with data-loss

    `documents_with_chunks` < `documents_count` means the extractor wrote a documents-row but no chunks for some files — typically an empty-template PDF or an extraction regression. It is an extraction-quality signal, not a data-loss signal. The raw zone still has the file.

Questions this answers
  • “How many documents do we have in each corpus today?”
  • “Which corpus produced the most chunks per document on average?”
  • “Are we extracting more pages via vlm or ocr-easyocr?”
  • “Which corpus is surfacing the most distinct PII entity types?”
10 of 10
Column
Type
Description
Tests · PII
extracted_day
Date
Calendar date the extractor ran. Sort + partition prefix.
1 test
corpus
LowCardinality(String)
Slug derived from the GCS path — `unstructured/<corpus>/...`. `unknown` when the slug can't be resolved.
1 test
pipeline
LowCardinality(String)
Docling pipeline that produced the doc — `vlm` (GPU) or `ocr-easyocr` (CPU fallback).
1 test
doc_lang
LowCardinality(String)
ISO-2 language code, lowercased. `unknown` when not detected.
1 test
documents_count
UInt64
Count of documents at this grain.
1 test
documents_with_chunks
UInt64
Subset of `documents_count` that produced at least one chunk. Drives the empty-chunk fraction formula measure.
chunks_count
UInt64
Total chunks across all documents at this grain.
pages_total
UInt64
Sum of Docling-reported `pages` (or chunk page-end fallback) across the grain.
chunk_text_bytes_total
UInt64
Sum of `length(chunk_text)` over the included chunks — a corpus-size proxy. The actual text never appears here.
pii_entity_kinds_total
UInt64
Sum of `pii_entity_kinds` (distinct-entity-type count from Presidio) across documents. Distinct kinds per doc, summed — not an occurrence count.