mart_document_chunks
Embedded chunks — the semantic-search surface.
One row per chunk with an `embedding Array(Float32)` column and a `vector_similarity('hnsw', 'cosineDistance', N)` index. AI reaches this mart only via `search_documents` — never via raw SQL. The `embedding` column is initially `[]`; the embedding worker fills it idempotently.
- Querying chunks via raw SQL
There is no raw-SQL surface over this mart for AI. The only path is the gateway's `search_documents` MCP tool, which embeds the query and runs persona-scoped cosineDistance. Direct SELECT on `chunk_text` or `embedding` requires operator credentials.
- Swapping the embedding model without re-embedding
The HNSW index is built for `cosineDistance` with a fixed dimension. Swap requires updating both the `EMBEDDING_DIM` env var on the worker AND the `embedding_dim` dbt var, then rebuilding the mart (which drops the old index) and re-running the embedder over every row. Mixed-dim or mixed-model rows break the index silently.
- “Find the 10 most relevant chunks about <topic>.”
- “How many chunks did we produce per document this month?”