zaid a.

Technical Manual Query System

Our operations library runs to something like eighteen documents once you count the OM Parts A through D, the type-specific manuals, the Cabin Safety Manual, the OCC Manual, EFB procedures, the FCOMs and AFMs for each fleet, and the MCAR Air Operations and Aircrew regulations sitting underneath all of it, and none of that counts the revisions. A line pilot asking about duty time limitations, or an audit team asking for documented evidence that a procedure traces back to an approved source, was until recently a matter of opening several PDFs and reading until the answer turned up or the patience ran out. I spent about a week building a system that fixes this, and the constraint I set myself going in was that it had to cost nothing beyond time, because I wanted to prove the concept was worth anything before committing budget or a business case to it.

The reason this is a harder problem than it looks is worth stating plainly. A PDF does not store the fact that a passage is a chapter heading, or that a block of numbers is a table, or that a paragraph belongs under section 3.2.1 rather than 3.2. It stores coordinates and glyphs, which is to say that it records where ink goes on a page, not what the ink means. Every text-extraction library, including the one I used, is reverse-engineering semantic structure out of a format that was only ever designed to reproduce a printed page. That is a fundamentally lossy operation, and any system built on it inherits the loss somewhere. You cannot extract perfect structure from a PDF. The real decision is where in the pipeline the imperfection gets absorbed, and I chose to absorb it early, at the chunking stage, rather than let it surface later as a wrong answer with a confident citation attached to it.

The architecture is what's called Retrieval Augmented Generation, and the concept underneath the acronym is simple even if the acronym is not. You build a searchable database of manual content, you retrieve the sections most relevant to whatever is being asked, and you hand those sections to a language model as context so it generates an answer grounded in text that actually exists in the library rather than in whatever it happens to remember from training. It combines two things that are each old on their own, database retrieval and language generation, and the retrieval step is what keeps the generation step honest.

This is worth distinguishing from ordinary keyword search, because the distinction is the entire reason the system is useful. A keyword search for "fuel reserve requirements" returns every paragraph containing those words in that order and misses every paragraph that discusses the same requirement in different language. A vector database does something different. Each chunk of manual text gets converted into a high-dimensional vector, essentially a long list of numbers representing where that chunk sits in a conceptual space, such that chunks about similar things sit near each other regardless of the specific words used. Let's suppose the query is phrased as "how much fuel do we need to carry as reserve." A keyword search finds nothing useful. The vector search finds the relevant section anyway, because the vectors for "reserve fuel," "fuel reserve requirements," and "how much fuel to carry" all land in roughly the same region of that space. I used ChromaDB, an open-source embedding database, running in persistent mode so the processed library survives a restart of the hosting platform rather than needing to be rebuilt from scratch every time.

The document-processing pipeline was the hardest part to get right. The chunking approach walks through each page detecting chapter headers with regular expressions matched against patterns like "Chapter 3" or "CHAPTER 3," detecting section numbers like "3.2.1," and creating a new chunk at each section boundary while carrying forward metadata on which chapter, section, manual, and page number that chunk belongs to. Chunk size sits at roughly 800 words. That figure is empirical rather than derived, I tried smaller chunks and lost context, tried larger chunks and diluted relevance with material the query wasn't actually asking about, and 800 words is where the balance held up in testing. The metadata is what makes the citation system possible later, because every answer the model gives can be traced back to a specific manual, chapter, section, and page rather than a vague gesture at "the manual says."

Text extraction itself runs on PyPDF2, which is reliable for straightforward paragraphs and unreliable for tables, because a table on a PDF page is, as far as the file format is concerned, just more coordinates and glyphs with no marker saying this is a table. I tried PyMuPDF and pdfplumber, both of which handle tables better, and both introduced dependency conflicts in the Streamlit Cloud environment that weren't worth fighting for a first version. So the limitation stands for now. Narrative procedures and requirements retrieve well, tabular data does not, and that gap is documented rather than hidden.

On the model side, I initially tried Anthropic's API, which I already knew well, but the subscription and per-token pricing model doesn't sit comfortably against an operational tool that might see a burst of queries during an audit and then nothing for two weeks. Google's API gives effectively unlimited queries at normal professional usage. That matters less for the raw cost and more for what it does to behaviour, nobody hesitates to run a query because they're worried about what it costs. I settled on the Pro model rather than the newer experimental releases, because the experimental models trade API stability for features I didn't need, and an operational tool that line crew and auditors rely on has no business running on an API that might change its behaviour under me without warning.

The system prompt does real work here. It establishes the model's role as an operations assistant for the company specifically, instructs it to cite manual sections by chapter, section, and page, to note the relevant ICAO or MCAR reference where one applies, to say plainly when the retrieved excerpts don't fully answer the question rather than filling the gap with something plausible-sounding, to hold a higher bar of precision for anything safety-critical, and to keep aircraft types distinct rather than blending an ATR procedure into an A320 answer because both happened to come up in the same query. The query and the retrieved excerpts are formatted so the model can tell the difference between the source material and the question being asked, which sounds obvious until you've seen a model blur the two.

The interface is built in Streamlit, which turns a Python script into a web application without touching HTML, CSS, or JavaScript directly. A sidebar handles manual management, a button to process or update the library, a running count of chunks and manuals loaded. The main panel takes a query, a slider for how many source chunks to retrieve, and shows the answer alongside an expandable list of the sources actually used. Deployment to Streamlit Cloud is connecting a GitHub repository, pointing at the main application file, and clicking deploy, three to five minutes from push to live application, and every subsequent push to GitHub redeploys automatically. The API key lives in Streamlit's encrypted secrets storage rather than the repository, which matters because a key committed to Git history stays there even after it's deleted from the current version.

Two bugs accounted for most of the debugging time. An early model version worked in some contexts and silently failed in others until I moved to a later stable release, and a manuals directory referenced as "./manuals" worked on my machine and broke on Streamlit Cloud because the working directory there isn't the same as the application file's location, fixed by resolving the path relative to the file itself rather than the working directory. Everything else was ordinary Git friction, the kind that comes from not yet having the staging and branching model in your hands rather than from anything wrong with the code.

Initial processing of the full library, about eighteen documents and several thousand pages combined, took fifteen to twenty minutes, and that happens once. Adding or updating a single manual afterward takes one to three minutes depending on length. A query returns in three to five seconds end to end, roughly one second of vector search and two to four seconds of generation, well inside what's acceptable for a tool used during a briefing or an audit rather than in the cockpit.

The citation is not a nicety, it's the point. An answer that says "according to OM Part A, Section 3.2.1, page 45" followed by an expandable list of every source consulted is an answer an auditor can independently verify against the controlled document, which is exactly the standard that matters in a regulatory environment. Ask how we determined a particular interpretation of a duty limitation and I can show the manual, the section, and the page, the same trail I'd have if someone had done the search by hand, just faster. And because the system searches across every loaded manual at once, a question like "compare ATR and A320 fuel reserve requirements" pulls relevant chunks from both OM-B documents and the model synthesises the comparison directly, which is the kind of cross-referencing that's tedious to do by hand and easy to get wrong under time pressure.

There's no conversation memory. It should be noted that this is deliberate rather than an oversight, self-contained answers with full citations serve regulatory purposes better than answers that lean on implicit context from three questions back. Adding memory would be straightforward, session state and prior exchanges folded into the prompt, and I'm not convinced yet that it's worth the added token cost against what it buys operationally.

On data handling, the manuals sit in a private GitHub repository, the application requires authentication, and the API sees only the excerpts sent with a given query rather than retaining anything. For content genuinely sensitive enough to need it, the whole stack could move on-premises, but operations manuals already distributed to every line crew don't carry a risk profile that justifies that complexity.

None of this is locked in. Adding a manual means dropping a PDF into the folder, committing, and letting the file-hash comparison decide that only the new file needs reprocessing. Changing the model is one line in utils.py. The vector database doesn't care which model generates the answer, so there's no dependency on any single provider surviving the next few years. The next real piece of work is the table problem, and the proper fix isn't a better text-extraction library, it's routing queries that ask about tables to send the actual PDF page as an image so the model reads the layout visually instead of through mangled extracted text, at the cost of more API calls and a slightly longer response for that subset of queries. After that I'd like a batch mode that takes a list of standard questions and produces a cited FAQ document for crew briefings or audit preparation, which is the same underlying machinery pointed at a different output format.

One could reasonably ask why not buy something off the shelf. I looked. What exists is either an enterprise document-search product priced and scoped for a company several times our size, or a consumer product that can't handle the semantic density of aviation regulatory text, or something in between that doesn't give citation accuracy good enough to survive an audit. Building it myself cost a week and gave me something those products don't, I understand exactly how every answer is produced, which matters considerably when a regulator asks how the tool works, and I can extend it into our other systems without asking a vendor's permission first.

The distance between having an idea for an operational tool and having one running in production has closed considerably in the last few years, to the point that the limiting factor is understanding the problem well enough to design the solution and having the patience to sit through the debugging, not access to a team of specialists or a budget most flight operations departments don't have. Every airline, every maintenance organisation, every training department, and every regulator sitting over them is carrying the same volume of documentation and the same shortage of time to search it, and the technical barrier to doing something about it has largely gone. What this system does, and all it does, is help a human find the right paragraph faster. It doesn't decide anything, and the citation trail behind every answer is the same documentary evidence we'd have produced doing the search by hand. Where the regulatory frameworks for tools like this eventually land is still being written, by EASA, by the FAA, by everyone watching this space closely. Being early to that conversation, with a working system and a documented audit trail already in place, is a better position than arriving at it after the frameworks exist and having to retrofit compliance into a tool that was never built with citation in mind.