Self-initiated demo by Flow Lab. Tallowby Freight is invented.
Staff ask in their own words. The assistant searches the company’s PDF policies, quotes the exact passage, links to the page, skips versions that were replaced, and says so plainly when the documents don’t cover a question.
Five steps, and each one can be checked on its own. In a client project the same steps can be connected to SharePoint, Google Drive or a shared folder.
Text is taken page by page. Code, version, effective date and status come from the document’s own control box.
Numbered headings (4., 4.2) become passages. Each one keeps its page number, so a link opens the right page.
A hand-made list of everyday words (petrol → mileage, nicked → stolen), and about a thousand sample questions a model wrote for the passages at import. Typos are matched to the nearest known word.
The answer is the passage itself. If the best match is weak, or shares too few of the question’s words, the assistant says it can’t find it.
Superseded versions are used only when a question is about the past. When a rule changed, the old wording is shown with the date it stopped applying.
40 questions written separately, in another assistant session that saw the documents but not the code, and checked to be unlike the questions used for tuning. The search was frozen before they were first run. They include traps: topics no document covers, and rules that changed between versions. The left column is the search in your browser. The right column is a recorded run of the full pipeline: for each question the same search picked five passages, and a language model answered from those five only. Every quote it gave is checked here, word for word. Every miss on both sides is listed below.
The demo shows the part that decides whether staff trust the answers. A client project adds the rest.
SharePoint, Google Drive, Notion or a shared folder. A new version replaces the old one on upload, with no re-typing.
A language model writes two or three sentences from the quoted passages. If it can’t cite a passage, it doesn’t answer.
A page on the intranet, Slack or Teams, or Telegram. Access by role, so drivers don’t see finance-only documents.
Past a few hundred pages, search adds embeddings next to keyword ranking and re-ranks the best passages. The test set works the same way, so the gain is measured, not assumed.
Scanned PDFs go through text recognition first. Tables are kept row by row, the way the London hotel limit is kept here.
A log of what people asked and what the documents didn’t answer. The test set grows each month, and gaps go back to the document owner.