Building a Local RAG Pipeline Over Engineering Documents
Why Local
Most write-ups on RAG systems assume you can send your documents to a cloud API. For a personal project over publicly available standards, that's fine. The moment you include project-specific content — mooring analysis reports, load case summaries, client engineering files — you can't. Offshore engineering documents are almost always covered by NDAs. The data cannot leave the machine.
Running everything locally also removes a cost variable. Once the model is downloaded, inference is free. For a system you want to query hundreds of times while working through a document set, that matters. The tradeoff is quality: locally runnable models are smaller and less capable than GPT-4-class APIs. For this use case, I found the tradeoff acceptable.
The Stack
- llama.cpp — local LLM inference. Runs quantised GGUF models on CPU with optional GPU offloading. I used Mistral 7B Instruct (Q4_K_M quantisation) for generation.
- sentence-transformers — embedding layer. all-MiniLM-L6-v2 for speed, BAAI/bge-base-en-v1.5 for quality. Both run entirely locally.
- FAISS — vector index. Flat L2 for small corpora (under 10,000 chunks), IVF for larger ones.
- pdfminer.six — PDF extraction. Not the prettiest library but gives the most control over the extracted text structure.
- Python — everything glued together with a simple FastAPI server and a minimal browser frontend for the query interface.
The Chunking Problem
Chunking is where most RAG tutorials wave their hands and move on. For technical PDFs — especially engineering standards — it's the hardest part of the whole pipeline.
The problem is that engineering documents are structurally inconsistent. A clause in ISO 19901-7 might be three sentences in one section and two pages in another. Tables often span multiple pages. Figures have captions that aren't adjacent to the figure in the extracted text. Cross-references — "see Table C.4 in Annex C" — mean the relevant content is nowhere near the referencing clause.
Naive chunking by character count or sentence count produces chunks that cut across logical boundaries. A chunk might start mid-sentence in a normative requirement and end mid-way through a related note. These chunks retrieve poorly and produce confused responses.
Embeddings
I tested three embedding models on this corpus. all-MiniLM-L6-v2 is fast and small (80 MB), decent on general text, but noticeably worse on technical terminology. BAAI/bge-base-en-v1.5 is significantly better — handles domain-specific language well and produces cleaner top-5 results. all-mpnet-base-v2 offers marginal improvement over bge-base with a meaningful speed penalty on CPU, not worth it.
I settled on bge-base-en-v1.5. One thing that helped significantly: embedding a short document identifier and section heading alongside the chunk text. Instead of embedding just the chunk, embedding "[ISO 19901-7 §8.3] — [chunk text]" improved retrieval for queries that reference document structure.
What Didn't Work
Fixed-size character chunking. The obvious first attempt — 500 characters with 100-character overlap. Results were poor because technical clauses are not correlated with character count. A safety-critical requirement might be 80 characters; a worked example might be 2,000.
Sentence-level chunking. Better, but engineering standards are full of numbered clauses that are individually short but only meaningful in the context of the parent clause. Sentence chunking loses that hierarchy.
Ignoring tables. pdfminer extracts table content as a stream of text without the grid structure. A table with 6 columns and 12 rows becomes 72 text fragments with no reliable row/column relationship. I ended up excluding tables from the retrieval corpus and noting their existence in adjacent clause chunks.
No metadata filtering. Without knowing whether a retrieved chunk is normative or informative, from which standard, and from which edition, the LLM conflates them. Early versions regularly produced responses that mixed ISO normative requirements with API informative guidance as if they were equivalent.
What Worked
Section-aware chunking. Instead of splitting by size, I split by section boundary — detecting heading patterns in the extracted text (numbered clauses like "8.3.2" are consistent enough to parse reliably). Each chunk corresponds to one clause or sub-clause. Chunks that are too long get split further, with section context prepended to each piece.
Metadata injection. Every chunk stores: document name, edition/year, section number, section title, normative or informative flag. At query time, the LLM prompt includes this metadata alongside the chunk text. Responses are noticeably more accurate about what is required versus what is recommended.
HyDE — Hypothetical Document Embeddings. Instead of embedding the raw query, generate a short hypothetical answer first and embed that. A query like "what safety factor applies to chain breaking load?" produces a poor embedding. A hypothetical answer — "the safety factor for chain breaking load in intact conditions is..." — produces a much better embedding that retrieves the relevant normative clause. This was the single largest quality improvement in the pipeline.
Reciprocal Rank Fusion over vector + keyword search. Vector search alone misses exact-match lookups like specific clause numbers or defined terms. Combining vector retrieval with simple keyword search (BM25) and merging with RRF gives better coverage than either alone.
HyDE + RRF together account for the majority of the quality gain. Either one alone is worth implementing; both together is significantly better than either.
Current State
The pipeline currently indexes ISO 19901-7 (2013 edition), API RP 2SK (3rd edition), and a set of my own project notes and analysis summaries. Total corpus is around 4,200 chunks. FAISS flat index, queried in under 200 ms on a mid-range laptop CPU.
Response quality is good enough to be genuinely useful for navigation and orientation tasks. It is not good enough to be authoritative on specific numerical requirements — those always need to be verified against the source document.
A cleaned-up version of this pipeline is what powers the Ask Dimitris chatbot in the About section — ported to Cloudflare Workers for public use. That implementation is documented step by step in the RAG Guide project.