lakehouse

Author	SHA1	Message	Date
root	04770c97eb	HNSW vector index: 100K search in 27ms (58x faster than brute-force) - instant-distance HNSW implementation for approximate nearest neighbors - HnswStore: build from stored embeddings, in-memory index, thread-safe - POST /vectors/hnsw/build — build index from Parquet (100K in 35s release) - POST /vectors/hnsw/search — fast ANN search - GET /vectors/hnsw/list — list loaded indexes Benchmark (100K × 768d, release build): Brute-force: 1,567ms HNSW: 31ms (50x) HNSW warm: 27ms (58x) Build cost: 35s one-time for 100K vectors (release mode) ef_construction=40, ef_search=50 — good recall/speed balance Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-03-27 20:00:50 -05:00
root	6cd1daeb51	Phase 11: Embedding versioning — model-proof vector layer - IndexRegistry: tracks all vector indexes with model metadata (model_name, model_version, dimensions, build stats) - Index metadata persisted as JSON in vectors/meta/ - Rebuilt on startup for crash recovery - GET /vectors/indexes — list all indexes (filter by source/model) - GET /vectors/indexes/{name} — get index metadata - Background jobs auto-register metadata on completion - Multi-version support: same data, different models, coexist - Per ADR-014: enables incremental re-embed on model upgrade Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-03-27 09:27:10 -05:00
root	3b695cd592	Dual-pipeline supervisor for embedding ingestion - 4 parallel pipelines (tuned for i9 + A4000) - Range-based work splitting (2500 chunks per range) - Round-robin retry on failure (3 attempts before dead-letter) - Checkpointing to disk every 1000 chunks (crash recovery) - On restart, loads checkpoint and skips completed ranges - Dead-letter queue for permanently failed ranges - Vectors assembled in order after all pipelines finish - Batch size 64 for GPU throughput Architecture: Supervisor → splits 100K chunks into 40 ranges ├── Pipeline 0: grabs range, embeds, reports progress ├── Pipeline 1: grabs range, embeds, reports progress ├── Pipeline 2: grabs range, embeds, reports progress └── Pipeline 3: grabs range, embeds, reports progress Failed range → back to queue → next available pipeline retries Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-03-27 09:06:28 -05:00
root	6a532cb248	Background job system for embedding — fixes 100K timeout - JobTracker: create/update/complete/fail jobs with progress tracking - POST /vectors/index now returns immediately with job_id (HTTP 202) - Embedding runs in tokio::spawn background task - GET /vectors/jobs/{id} returns live progress (chunks embedded, rate, ETA) - GET /vectors/jobs lists all jobs - Progress logged every 100 batches with chunks/sec and ETA - 100K embedding job running successfully at 44 chunks/sec - System stays responsive during embedding (queries in 23ms) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-03-27 09:03:07 -05:00
root	26fc98c885	Phase 7: Vector index + RAG pipeline - vectord crate: chunk → embed → store → search → RAG - chunker: configurable chunk size + overlap, sentence-boundary aware splitting - store: embeddings as Parquet (binary blob f32 vectors), portable format - search: brute-force cosine similarity (works up to ~100K vectors) - rag: full pipeline — embed question → search index → retrieve context → LLM answer - Endpoints: POST /vectors/index, /vectors/search, /vectors/rag - Gateway wired with vectord service - Tested: 200 candidate resumes indexed in 5.4s, semantic search + RAG working - 20 unit tests passing (chunker, search, ingestd, shared) - AI gives honest "no match found" when context doesn't support an answer Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-03-27 08:12:28 -05:00

5 Commits