Field notes
Ask your business anything: private AI search over your own documents
Every business has The Folder. Or seven folders, two email accounts, and a scanner tray. Somewhere in there is the answer to "what did we actually agree to with that distributor in 2023," and finding it costs forty minutes of somebody's evening.
Private document search is the first AI capability we install for most clients, because it attacks that exact forty minutes. Here's what it is and what living with it looks like, without the jargon.
What it is, without the jargon
A model running on a server you own reads and indexes your documents: contracts, SOPs, invoices, quotes, leases, correspondence. Afterward, anyone you authorize can ask questions in plain English and get an answer with the source document attached. The technical term is retrieval-augmented generation. The experience is closer to "the filing cabinet learned to talk."
The part that matters for a business: the model, the index, and every question asked all live on your hardware, behind your firewall. Covered in full in our on-prem AI overview.
A worked example
The question: "What's the termination clause in our agreement with the produce distributor?"
The old way: remember which year the contract was signed, guess at the filename, open four near-identical PDFs, skim to section 11, hope it's the signed version.
With private search: type the question, get the clause quoted back with a link to the exact page of the exact PDF, including the note that a 2024 amendment changed the notice period. Thirty seconds, and you read the source yourself before acting on it.
A few more that come up constantly:
- Owner: "Summarize this lease renewal against the current one and flag what changed."
- Owner: "What did we quote the bakery for cameras last spring?"
- Office manager: "Which employees are due for the safety training refresher per the handbook schedule?"
- Bookkeeper: "Find every invoice from that vendor over two thousand dollars this year."
The honest boundary: it finds and synthesizes what's written in your documents. It can't answer things nobody wrote down, and a messy document set produces messier answers. Garbage in still applies.
Why on-prem instead of a chatbot subscription
Look at the list above. Contracts, quotes, HR records, payroll invoices: these are precisely the documents you can't paste into a consumer chatbot, and shouldn't ship to any third-party AI service whose terms you haven't lawyered. On your own hardware the calculus flips: nothing leaves the building, nothing trains on your data, and there's no per-seat toll on asking questions. For practices with patient records, this is the difference between "interesting but forbidden" and "a normal IT project"; our practice guide covers that side.
What it's honestly bad at
- Documents that were never digitized well. Clean scans read fine through OCR. Coffee-stained handwriting from 2011 does not.
- Heavy math across spreadsheets. "Total these 400 invoices by quarter" is a reporting job, not a search job. Right tool.
- Anything unwritten. Tribal knowledge stays tribal until someone writes it down; the system then makes it findable forever, which is a good reason to finally write it down.
- Judgment. It quotes and cites; deciding is still your job.
What setup actually involves
- An assessment of where documents live. Server shares, cloud drives, the scanner's output folder, exported email. Consolidation is often half the value.
- Access scoping, treated as non-negotiable. The search must respect permissions. Payroll and HR answers go only to the people cleared for payroll and HR. This is configured before anyone types a question, not after.
- Hardware sized to the job. For search alone, a modest GPU server does it, and the same box typically carries other duties.
- The index build, then quiet upkeep. First indexing takes a while; after that, new documents become searchable on a schedule. From there it's a managed system like any other on your rack: backed up, monitored, patched.
Common questions
When staff ask it questions, does anything go to an AI company?
No. Questions, documents, and answers stay on your server. That's the entire point of running it on-premises.
Can it read scanned paper?
Typed documents scanned cleanly, yes, through OCR during indexing. Faded thermal receipts and rough handwriting are best-effort, and we'll tell you honestly what your archive will and won't yield during the assessment.
Can we limit who can search what?
Yes, and we won't deploy it any other way. Permissions mirror your file access: if someone can't open the folder, they can't search its contents either.
How current is the index?
New and changed files get picked up on a schedule, typically same-day. The invoice scanned this morning is findable this afternoon.
Got a folder nobody can find anything in?
Bring one real question you couldn't answer last month and we'll show you what private search would have done with it. Call (415) 555-0134 or send us a note. →