ROSEWOOD SYSTEMS
BLOG TALK TO US
ARTICLE

What to Do Before You Train an AI on Your Company's Documents

August 5, 2026 · by Nia Rowe

There is a predictable moment in most organizations right now. Someone realises the company has years of accumulated knowledge sitting in folders, and that AI could theoretically make all of it searchable and useful. The idea is good. The first attempt usually disappoints.

It rarely disappoints for technical reasons. It disappoints because of decisions nobody made before the documents were uploaded.

Here is what is worth settling first.

1. Decide what "correct" means before you start

Ask three people in your organization the same client question and you will often get three different answers. All defensible, none identical.

An AI trained on all your material will faithfully reproduce that inconsistency, and it will do so confidently. That is worse than the original problem, because a human hedges when unsure and a language model frequently does not.

So the first task is not technical. Pick the questions that matter most — the twenty or thirty that actually recur — and decide what the right answer is. If your organization cannot agree, the AI cannot resolve it for you.

2. Separate what is current from what merely exists

Most document repositories are archaeology. The 2019 onboarding guide sits beside the 2024 revision. A policy that was superseded two years ago is still the top search result internally because it has the clearest filename.

A retrieval system does not know which is authoritative. It knows which is textually similar to the question.

Before anything is ingested, someone has to walk the material and mark what is current, what is historical, and what should never have been kept. This is unglamorous and it is the single highest-leverage hour you will spend on the project.

3. Work out what must never go in

Some material is genuinely useful to the AI and genuinely should not be in it.

Client records containing personal information. Anything under professional privilege. Credentials. Health or financial details. Employment files. Material you are contractually forbidden to share with a third party — and most AI processing does involve a third party.

Draw that line explicitly and in writing before ingestion, not after someone notices. If you operate in Canada, remember that consent obtained for one purpose does not automatically extend to a new one.

4. Decide who it answers to

An internal tool used by trained staff who understand its limits is a different product from a client-facing assistant, even if the underlying technology is identical.

Client-facing raises the bar considerably: it needs guardrails for questions outside its knowledge, clear disclosure that it is AI, escalation paths to a person, and a much lower tolerance for confident-but-wrong. Decide which you are building before you build it, because retrofitting the stricter version onto the looser one is expensive.

5. Plan for what happens when it is wrong

Not if. AI-generated output is probabilistic — it can be incorrect, incomplete, outdated, or unsupported by your own source material, and it can vary between two runs of the same question.

Any system that people will genuinely rely on needs an answer to: how does someone report a bad answer, who reviews it, and how does the fix get made? Without that loop, errors persist and trust drains quietly until people stop using the tool.

The pattern underneath all five

Every one of these is an organizational decision wearing a technical costume. Which is why the projects that succeed usually start with a week of unglamorous conversation rather than a platform evaluation.

The technology is genuinely capable now. The bottleneck has moved.

If you are weighing this for your own organization and want to think it through with someone who builds these systems, that is exactly the conversation I enjoy.

Nia Rowe
AI Architect, Rosewood Systems
Nia builds intelligent software, custom AI partners trained on an organization’s own knowledge, and practical AI education. Based in Toronto. More about Nia →

Ready? Let’s build something.

Purpose-built software, an AI partner trained on your business, or practical AI education for your team. We’d love to help.

inquiries@rosewoodsystems.io