Consider whether the same questions keep going to the same few people in your firm. Which delivery terms apply to this wholesale client? What is our policy on late returns? Which price list is the current one? Together they take a large part of someone's week.

The answers are in a few people's heads, in shared drives and in old email threads. When one of those people is on leave or busy with a customer, the rest of the team waits or guesses.

This guide explains AI knowledge base development in plain words for owners and managers of small firms. It covers what the system does, which documents work, how the build runs, what drives the effort, how privacy is handled, where it can go wrong and who checks the answers.

What an internal AI knowledge base actually does

Picture a new hire in their second week. A wholesale client asks about delivery, and the new hire does not know which terms apply. Today they ask a colleague. With an internal knowledge base, they ask the system and get an answer drawn from your own files.

The method behind this is called retrieval-augmented generation, or RAG: before writing anything, the model checks a separate, trusted collection of documents outside its original training.¹ In practice, a document retriever first selects the documents most relevant to the question.² It then writes the answer from those passages, so the answer rests on your company's own material.³ The answer can include citations or references, so the person can open the source documents and check.¹,²

This matters because a general chat model's training data is static and has a knowledge cut-off date, which can lead to out-of-date answers.¹ It has never seen your price list.

It is the right fit for knowledge that lives in a few people's heads, and it is also why an AI system is more than a chatbot.

This guide is about questions from your own staff. Customer questions that arrive by WhatsApp or email are a related task: the same approved documents can feed a drafted reply, and a person approves it before it is sent. A WhatsApp message that is really an order is a third task, order entry: the system prepares the order line, and a person checks it before it is saved.

Workflow · Staff question to checked, sourced answer
  1. Trigger

    A staff member asks a question

    For example, a staff member might ask: which delivery terms apply to this wholesale client?

  2. Prepared by the system

    Relevant passages found

    The system searches the approved documents and picks the sections that match the meaning of the question.

  3. Prepared by the system

    Answer drafted with sources

    A short answer is written from those sections only, with links to the documents it used.

  4. Checked by a person

    Staff member checks the source

    The person opens the cited document before using the answer with a customer. Gaps or conflicts go to the named document owner.

  5. Result

    Answer used, gaps fixed

    The checked answer is used, and the document owner updates any missing or outdated file.

Which documents work, and which cause trouble

Start with materials such as manuals, policies and product specifications.⁴ A typical use is an internal assistant that answers staff questions from company manuals and documentation.⁴ If a document is clear enough for a careful new hire, it is usually clear enough for the system.

Consider using a maintained sheet of client-specific prices and delivery days as a source. If your data changes often, ask how the system picks up changes before using them in answers. Rules for order changes and substitutions work best as a short written policy, so there is something clear to answer from. Turning incoming order emails into sheet rows is a separate task: a knowledge base answers questions, it does not enter orders.

Trouble starts with files that confuse people too. Multiple versions of the same price list, with no sign of which one is current. Scanned pages where the text cannot be read reliably. Folders that nobody owns, so nobody knows whether the content is still true.

During setup, large documents can be split into smaller sections so each part can be matched on its own.³

The simple rule is that outdated files give outdated answers. Without clear rules and ownership, such a system can show information that is old, one-sided or no longer allowed, which damages trust.⁵ Sorting the files before the build is the most useful work an owner can do.

How AI knowledge base development works, step by step

A RAG build has three main stages: preparing the documents, finding the right passages and writing the answer.⁶ Your team supplies the documents and the test questions. Whoever does the technical setup, an outside partner or a capable member of your own staff, handles the search settings, so you do not need to understand them; you decide which documents go in, who sees what and who checks.

Gather and clean the documents

Collect the files people actually use when they answer questions. Remove old versions, mark the current price list and note which scans need retyping.

Cut the documents into sections

Each document is split into short sections that can be searched one by one.

Set up search that matches meaning

Plain word search fails when staff phrase a question differently from the document, so the search has to understand what the person means.³ In plain words: consider whether a search finds useful passages when a question uses different wording from the documents.

Decide who sees what

The system can be set up so that each user only reaches the information their role allows.⁶ A warehouse worker may see delivery rules but not salary policies.

Test with real staff questions

Collect questions people really asked last month, in the words they used. Run them through the system and check every answer against the source. Wrong or missing answers usually point to a document that needs fixing.

A sensible first trial covers one folder and one team. Agree beforehand how you will judge it: are the answers right, do they cite the right file, and do the usual go-to people get fewer interruptions? If not, stop or fix the documents before going further.

Name a document owner

Each document or folder gets one named person who keeps it current and receives questions the system could not answer. Without this step, quality slowly drops after launch.

What drives the time, effort and running cost

You do not retrain a model to add your company's information. Retraining is expensive in both money and computing power, and RAG is a cheaper way to give a model new information.¹ For most organizations, updating the retrieval store can replace fine-tuning the model.⁴ The retrieval store is simply the searchable copy of your documents, so keeping answers current means keeping files current.

RAG can send only the most relevant sections rather than an entire document, reducing query size and improving efficiency.⁶ For your budget, the practical step is to ask for a written monthly running estimate based on how many questions your staff ask in a normal week.

Consider document volume, organization and staff time when planning the build. A firm with a well-kept shared drive starts much faster than one where the rules live in email threads. Also plan regular hours for the document owners after launch, because updating files and handling unanswered questions is ongoing work. Ask for a build quote after discussing the documents and number of users. Ask for a price range based on how your files are organized.

If your staff ask questions in Latvian, test Latvian questions during setup, not after launch.

Privacy, GDPR and who sees which answers

With RAG, sensitive data can stay on your own servers while still being used by a local model or a trusted outside one.⁶ Some open models can even be downloaded and run on a local computer.⁷

Keeping everything on your own servers gives the most control, but someone has to run and update that setup. If you have no IT staff, consider working with a trusted outside provider. An outside provider is easier to run, so compare providers on written answers: where the data is stored, how long questions and files are kept, how you delete them, whether they are used to train the provider's models, and whether you get a log of who accessed what.

A few practical rules can help when handling sensitive files. Choose on purpose which files go in, rather than connecting a whole drive. Keep personal data out unless the answers truly need it. Agree data processing terms in writing with any outside provider. This is general guidance, not legal advice, so check specific cases with your own advisor.

Where it goes wrong, and who checks the answers

A correct source does not guarantee a correct answer. A model can produce wrong information from accurate documents when it misreads the context.² It may also give an answer when it should say that it does not know.²

Most enterprise AI initiatives fail because architecture and governance are fragmented rather than because models are weak.⁵ In a small firm, that means clear owners and clear rules matter more than which model you pick.

Answers based on verifiable sources are more reliable and contain fewer invented details.⁴ That only helps if someone looks at the source. So staff check the cited document before using an answer with a customer, and a named person handles gaps and conflicts. More on this in our guide to building AI workflows with human review.

Responsibility stays with people. The staff member who sends an answer is responsible for it, just as with an answer written today. The document owner fixes a wrong or missing source, and the owner or manager decides which kinds of question the system may answer at all.

Warning

A correct document can still produce a wrong answer. Before any answer reaches a customer, the staff member opens the cited source and checks that it says what the answer says.

When to bring in a specialist

If your answers sit in one tidy folder and a curious staff member wants to experiment, an off-the-shelf or free AI tool may be enough to learn what works. Look for one that answers only from files you upload, shows which file it used and needs no IT specialist to set up. Products change often, so judge any tool by these tests rather than by name. Start small and check every answer.

If staff already find answers quickly in one shared document, keep using it. This effort pays off when answers sit across many files and the same few people keep being interrupted.

Get help when files are spread across several tools, when access rules matter, or when nobody in the team has time to maintain the system. First learn how the team works and who answers what, then plan the software and who will keep it current after launch.

When choosing an outside partner, look for one who asks about your documents and your team before talking about software. They should work in your staff's language, set access rules, show sources in every answer and agree in writing who maintains the system after launch.

Frequently Asked Questions

What is an AI-based knowledge base?

It is software that answers staff questions from your company's own approved documents rather than from general training data.¹,³ It first selects the most relevant documents to augment the question.² Answers can include sources that a person can open and check.¹,²

How do I build an AI knowledge base for my company?

Start by gathering company documents, such as manuals, policies and specifications.⁴ The build then covers three stages: preparing the documents, finding the right passages and writing the answer.⁶ Test it with real staff questions and name a person who keeps each document current.

Does ChatGPT have a knowledge base?

A general chat model's training data has a cut-off date, which can lead to out-of-date answers.¹ RAG connects a model to your own documents so its answers draw on them.⁴,⁶ That connection, together with access rules and visible sources, is what an internal knowledge base adds.

Does an internal AI knowledge base work in Latvian?

It can, but test it before launch with real questions in Latvian, Lithuanian, Estonian, Russian and English, whichever your staff use. Search that matches meaning rather than exact words helps when staff phrase questions differently from the document.³ Documents that mix languages need checking during setup, so treat language as something to test, not something to assume.

Sources (7)

Have a process you'd like to talk through?

Book a discovery call (opens in a new tab)