End-to-end RAG pipeline: chunking + embedding ingestion, vector index with metadata filtering, hybrid retrieval, and streamed LLM responses with citation rendering. Tuned chunk strategy and re-ranking for measurable improvements in answer accuracy and reduced hallucinations on a domain-specific corpus.
