lakehouse/PHASES.md at 97a376482cefac2cfb99dc8ddea338d97a500733

root 97a376482c Phase C: Decoupled embedding refresh

Implements the llms3.com-inspired pattern: embeddings refresh
asynchronously, decoupled from transactional row writes. New rows arrive,
ingest marks the vector index stale, a later refresh embeds only the
delta (doc_ids not already in the index).

Schema additions (DatasetManifest):
- last_embedded_at: Option<DateTime> - when the index was last refreshed
- embedding_stale_since: Option<DateTime> - set when data written, cleared on refresh
- embedding_refresh_policy: Option<RefreshPolicy> - Manual | OnAppend | Scheduled

Ingest paths (pipeline::ingest_file + pg_stream) call
registry.mark_embeddings_stale after writing. No-op if the dataset has
never been embedded — stale semantics only kick in once last_embedded_at
is set.

Refresh pipeline (vectord::refresh::refresh_index):
- Reads the dataset Parquet, extracts (doc_id, text) pairs
- Accepts Utf8 / Int32 / Int64 id columns (covers both CSV and pg schemas)
- Loads existing embeddings via EmbeddingCache (empty on first-time build)
- Filters to rows whose doc_id is NOT in the existing set
- Chunks (chunker::chunk_column), embeds via Ollama (batches of 32),
  writes combined index, clears stale flag

Endpoints:
- POST /vectors/refresh/{dataset_name} - body {index_name, id_column,
  text_column, chunk_size?, overlap?}
- GET /vectors/stale - lists datasets whose embedding_stale_since is set

End-to-end verified on threat_intel (knowledge_base.threat_intel):
- Initial refresh: 20 rows -> 20 chunks -> embedded in 2.1s,
  last_embedded_at set
- Idempotent second refresh: 0 new docs -> 1.8ms (pure delta check)
- Re-ingest to 54 rows: mark_embeddings_stale fires -> stale_since set
- /vectors/stale surfaces threat_intel with timestamps + policy
- Delta refresh: 34 new docs embedded in 970ms (6x faster than full
  re-embed); stale_cleared = true

Not in MVP scope:
- UPDATE semantics (same doc_id, different content) - would need
  per-row content hashing
- OnAppend policy auto-trigger - just declares intent; actual scheduler
  deferred
- Scheduler runtime - the Scheduled(cron) variant declares the intent so
  operators can see which datasets expect what, but the cron itself is
  separate

Per ADR-019: when a profile switches to vector_backend=Lance, this
refresh path benefits — Lance's native append replaces our "read all +
rewrite" Parquet rebuild pattern. Current MVP works well enough at
~500-5K rows to validate the architecture; Lance unblocks the 5M+ case.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

2026-04-16 03:00:43 -05:00

9.6 KiB

Raw Blame History

Phase Tracker

Phase 0: Bootstrap ✅

Phase 1: Storage + Catalog ✅

Phase 2: Query Engine ✅

Phase 3: AI Integration ✅

Phase 4: Frontend ✅

Phase 5: Hardening ✅

Phase 6: Ingest Pipeline ✅

Phase 7: Vector Index + RAG ✅

Phase 8: Hot Cache + Incremental Updates ✅

Phase 8.5: Agent Workspaces ✅

Phase 9: Event Journal ✅

Phase 10: Rich Catalog v2 ✅

Phase 11: Embedding Versioning ✅

Phase 12: Tool Registry ✅

Phase 13: Security & Access Control ✅

Phase 14: Schema Evolution ✅

Phase 15+: Horizon

9.6 KiB Raw Blame History Unescape Escape

Phase Tracker

Phase 0: Bootstrap ✅

Phase 1: Storage + Catalog ✅

Phase 2: Query Engine ✅

Phase 3: AI Integration ✅

Phase 4: Frontend ✅

Phase 5: Hardening ✅

Phase 6: Ingest Pipeline ✅

Phase 7: Vector Index + RAG ✅

Phase 8: Hot Cache + Incremental Updates ✅

Phase 8.5: Agent Workspaces ✅

Phase 9: Event Journal ✅

Phase 10: Rich Catalog v2 ✅

Phase 11: Embedding Versioning ✅

Phase 12: Tool Registry ✅

Phase 13: Security & Access Control ✅

Phase 14: Schema Evolution ✅

Phase 15+: Horizon

9.6 KiB

Raw Blame History