How Much Does a Custom AI Document Assistant Cost?
What a custom internal document Q&A assistant really costs to build and run, what drives the price up, and when a local LLM beats an API for compliance-heavy teams.
TL;DR
A working internal document assistant over a few thousand files is a 3 to 6 week build for a small team; the price is driven far more by permissions, document messiness and hosting constraints than by the model itself. Running costs are usually small next to the build, unless you need everything to stay on your own hardware.

A custom AI assistant that answers questions over your own documents typically takes a small team three to six weeks to build to a usable first version, and the cost is set mostly by how messy your documents are, how strict your access rules are, and whether the data can leave your servers. The language model is the cheapest and least risky line item. Everything around it is where budgets slip.
This guide breaks the cost into the parts you will actually be quoted for, so you can compare proposals and decide whether to build at all.
What you are really buying
"Chat with our documents" sounds like one feature. In practice it is four systems bolted together:
- Ingestion: pulling files from SharePoint, Google Drive, a CRM, email, or a shared folder, and keeping them in sync.
- Preparation: turning PDFs, scans, spreadsheets and slide decks into clean text with structure preserved.
- Retrieval: indexing that text so the right passages come back for a question. This is the RAG layer we covered in RAG explained for founders.
- Generation and UI: the model that writes the answer, plus the interface, citations, feedback buttons and logging.
Vendors who quote a single low number are usually pricing only the last part.
The cost drivers, in order of impact
1. Document quality and variety
A clean folder of Word files with sensible titles is a one week ingestion job. Ten years of scanned contracts, policies in six revisions, and tables that only make sense visually is a month. OCR, table extraction, deduplication and version handling are unglamorous engineering, and they are where most of the effort lands.
Ask any vendor how they handle a scanned PDF with a table in it. If the answer is vague, expect the assistant to confidently misquote your own pricing sheet.
2. Permissions
If everyone in the company may see everything, retrieval is simple. If the assistant must never show HR documents to engineers, or a client's file to another account manager, every query has to be filtered by who is asking. That means mirroring your source system's access control into the index and keeping it in sync when permissions change.
This is the single most underestimated item in document assistant projects. We wrote up the patterns in permission-aware RAG. Budget for it explicitly, because retrofitting it later means rebuilding the index.
3. Where it runs
Three realistic options, roughly in ascending build cost and descending running cost:
- Cloud API for the model, your own database for the index. Fastest to ship. Documents are chunked and stored in something like PostgreSQL with pgvector, and the model is called per question. Per-query cost is small but never zero, and your text goes to the provider.
- Cloud infrastructure you control, open-weight model. A GPU instance running a model through a server such as Ollama or vLLM. Data stays inside your account. You pay for the instance whether or not anyone is asking questions.
- Fully on-premises. A box in your own rack or a private data centre. Highest setup cost, best answer to a regulator, and the only option for some healthcare and government buyers in the Gulf and the EU.
The trade-offs are laid out in detail in self-hosting an LLM vs API. The short version: if compliance forces the decision, the decision is cheap to make, and the cost is just the price of doing business in your sector.
4. Answer quality requirements
A tool for the sales team to find the latest case study can tolerate an occasional miss. A tool that nurses use to check a clinical protocol cannot. The stricter the bar, the more you spend on evaluation sets, citation enforcement, guardrails that refuse to answer when retrieval is weak, and human review loops.
The cost difference is not abstract. For the sales tool, a wrong answer means someone opens the wrong PDF and closes it. For the clinical tool, every answer needs a citation to the exact protocol revision, a refusal path when the retrieved passages do not actually contain the answer, and a clinician signing off on the evaluation set before launch. That review loop alone can add more engineering weeks than the entire sales pilot.
Insist on an evaluation set from day one: real questions from real staff, with the correct answer and the correct source document. Start with the ten questions your team asks most often, and grow it from the feedback buttons in the interface. It is the only honest way to know whether a change made the assistant better or worse, and it is what you should hold a vendor to.
Rough effort bands
Treat these as ranges for a scoped proposal, not a price list. Every serious quote should be preceded by a look at your actual documents.
- Pilot, single source, no permissions, cloud API. One to two weeks of engineering. Good for proving the idea to a sceptical board.
- Production for one department, permissions mirrored from one system, citations, feedback capture. Three to six weeks.
- Company-wide, multiple sources, strict access control, self-hosted model, audit logs. Two to four months, plus ongoing maintenance.
Ongoing costs come in three flavours: model usage or GPU hosting, re-indexing as documents change, and someone owning the evaluation set and quality. The last one is the one companies forget, and it is why so many pilots quietly rot after launch.
Where the model itself fits
Founders often assume the choice of model is the expensive decision. It rarely is. Model costs keep falling, and switching providers behind a retrieval layer is a small job if the system was designed for it.
What matters more is how the model is used. In our own outreach engine, which scrapes each prospect's website with a self-hosted Firecrawl instance and a local model, we tried multi-step chains that first extracted facts, then planned, then drafted. A single call that extracts the facts and writes the email in one pass turned out to be both cheaper and better. The same lesson holds for document assistants: a well-prompted single call over well-retrieved passages beats an elaborate agent loop most of the time. We covered that in single call vs agent chains.
The practical consequence for your budget: spend on retrieval and document preparation, not on orchestration frameworks.
Build, buy, or wait
Buy if your documents already live in Microsoft 365 or Google Workspace, your access rules match those platforms, and you have no data residency constraint. The bundled assistants are good enough for many teams and the price is predictable. Run the same test on them you would run on a vendor: hand one your worst scanned PDF with a table and see whether it quotes the numbers correctly.
Build if any of these apply:
- Documents are spread across systems that no single vendor indexes, such as a CRM, a shared drive and ten years of email.
- Access rules are more complex than folder permissions, like the account-manager-per-client case above.
- Data cannot leave your infrastructure, or your regulator wants to see exactly where it goes. This is the usual position for healthcare and government buyers in the Gulf and the EU.
- You want the assistant embedded in your own product or workflow rather than in a separate chat window.
Wait if you cannot name ten questions your staff ask repeatedly that a document would answer. Without that list, you do not have a use case, you have a demo. Those same ten questions become the seed of your evaluation set, so writing them down costs nothing and tells you immediately whether the project is real.
Questions to ask any vendor
- How will you handle our scanned PDFs and tables, and can you show me output from a sample?
- How are permissions enforced at query time, and what happens when someone's access changes?
- What does the evaluation set look like, and who maintains it after launch?
- Which parts run on our infrastructure, and which parts call an external API?
- What does re-indexing cost when we add a thousand documents a month?
- If we switch models next year, what has to change?
A vendor who answers all six crisply is worth paying. One who redirects the conversation to which model they use is selling you the cheap part.
If you are weighing a document assistant for a compliance-heavy team and want a scoped estimate against your actual files rather than a slide, let's talk.
Frequently asked questions
Can we just use ChatGPT Enterprise or Copilot instead of building one?
If your documents live in one vendor's ecosystem and you have no hard data residency rules, yes, and you should try that first. Custom builds win when documents are spread across systems, access rules are complex, or the data cannot leave your infrastructure.
What is the biggest hidden cost?
Document preparation. Scanned PDFs, tables, versions of the same policy and inconsistent naming eat more engineering time than the AI layer, and they decide whether answers are trustworthy.
Does a local model make it more expensive?
Upfront, yes: you pay for a GPU box or a dedicated server and for someone to tune throughput. Per query it is close to free, so it pays back on high volume or when compliance rules out cloud APIs anyway.
How do we know it actually works before rolling it out?
Insist on an evaluation set: 50 to 100 real questions with known correct answers, scored before launch and after every change. If a vendor cannot show you that, they are guessing.
How Pykero can help
Services related to this article.
Building something like this?
Pykero Agency designs and ships production web, mobile, SaaS, and AI products.
Talk to us →

