What Material Does a Knowledge Base Need at Cold Start? An Actionable Checklist
Cold start is not a migration of the entire shared drive. It is the disciplined selection of authoritative rules, procedures, terminology, real user questions, exceptions and evidence around a bounded business problem.
Key takeaway
Start with one bounded domain and a set of real questions. Collect authoritative rules, procedures, terminology, user wording, exceptions and checkable evidence. Record version, owner, scope, permissions and status for every source. Judge the first release by reliable question coverage, not file count.

An enterprise knowledge base does not need the whole shared drive on day one. A workable start is a bounded business problem supported by authoritative rules, standard procedures, product facts, real user wording and a small set of exceptions — with a version, owner and scope attached to every source.
Cold-start projects are often run backwards. A team scans the shared drive from the top, gathers hundreds of contracts, manuals and meeting notes, yet cannot say what employees will ask in week one. The right order is questions first, evidence second. A knowledge base delivers checkable answers, not a larger folder.
Cold start ends when the first questions are reliable, not when the archive is complete
Begin with a one-sentence scope statement: “For the after-sales team, answer installation, warranty and returns questions for product lines A and B; do not answer live stock, case-by-case compensation or unreleased-product questions.” That sentence fixes the users, topics, time boundary and refusal boundary. A file that cannot support an in-scope answer is not essential to the first release.
A smaller scope makes missing evidence visible and version decisions easier for business owners. Smaller organisations can follow the method in A Knowledge Base for Small Teams: Start Usable, Not Perfect: make one frequent scenario work, prove that people will use and maintain it, then open a second domain. Cold start does not shrink the long-term ambition; it turns the first delivery into a promise that can be tested.
Work backwards from questions people actually ask, not forwards from the folder tree
Useful question samples come from support tickets, onboarding notes, sales enquiries, rejected approval requests and repeated questions in business chat. Preserve the asker's wording while removing names, contact details, contract values and other unnecessary sensitive data. Do not rely on managers to imagine what staff ask. A manager says “after-sales policy”; a frontline colleague says “can the customer return it after opening the box?” The second phrasing is what retrieval must recognise.
Add three pieces of context to each question: who asks it during which task, what happens if the answer is wrong, and which class of source ought to prove the answer. This separates frequent but low-risk questions from less frequent questions that cannot safely be wrong. Priority is not frequency alone; it combines business impact, the availability of an authoritative basis, and whether the knowledge base can answer independently.
Six source roles reveal exactly what a first corpus is missing
File extension is not a useful content class. “PDF, Word and Excel” tells you how to open an item, not what duty it performs in an answer. Collect the first corpus by six evidence roles:
- Authoritative rules and facts: current specifications, pricing positions, service terms, policies and internal rules that answer “what is it?” and “is it allowed?”
- Standard procedures: application, approval, configuration, troubleshooting and delivery steps, including prerequisites, inputs, outputs and a clear end state.
- Terminology and name mappings: aliases, former product names, departmental jargon, abbreviations and bilingual terms that connect user language to formal material.
- Real questions and standard phrasings: sanitised questions from tickets, training and frontline support, retained for retrieval tests rather than importing whole conversations.
- Exceptions, limits and escalation conditions: when the general rule does not apply, when a named owner must decide, and which questions the system must refuse.
- Checkable evidence and examples: form templates, good and bad examples, process diagrams and calculation definitions that point from an explanation to the next action.
A short, closed and repeatedly asked question can begin as a Q&A pair. A subject involving conditions, steps, versions or attachments belongs in a structured knowledge article. These are not merely short and long versions of one container; the boundary between an FAQ and an enterprise knowledge base explains why.
One source-inventory sheet should settle version, accountability and access
At minimum, record domain, title, source location, summary, authority level, current version or effective date, applicable products and departments, business owner, audience, original format, conversion status, conflicts and the trigger for the next review. ISO 15489-1:2016 places records, metadata, assigned responsibilities and controls in one management frame. The immediate cold-start lesson is that the content and the metadata explaining its origin must travel together.
The business owner is not whoever uploaded the file; it is the person authorised to confirm that the content remains valid. An administrator may transform formats and a project manager may chase deadlines, but only the after-sales owner can rule which returns policy is current. The same inventory can feed the later work on channels, permissions and operating loops described in the complete enterprise AI knowledge base implementation path.
Deduplication means deciding what represents current truth, not deleting files
When several files cover one topic, classify their relationship: exact duplicate, old and new versions, summary and source, partial conflict, or different applicability. An exact duplicate needs only one retrieval copy. An old version should leave current retrieval but retain an archive location. A summary must not impersonate the authority when the source is absent. Versions with different scope must put those conditions into their title and metadata.
Do not ask AI to “reconcile” a conflict. Open a conflict record containing the disputed statements, sources, affected questions, decision owner and resolution state. Until it is ruled on, the material can stay out of retrieval and the question can route to a named human. An explicit unknown is safer than a confident composite error. A conflict often exposes a business rule the company itself has not settled; the knowledge base project has no mandate to settle it silently.
Make material parseable before ingestion; neat pages are not the same as clear structure
Headings, tables and page numbers visible to a person may be invisible to a parser. Prefer first-release sources with selectable text, explicit heading levels, complete table headers and a natural reading order. Run optical character recognition on scans and sample-check the result; remove repeated headers and footers; restate critical conditions from process diagrams in text; and split a file that mixes products or years into sections with unambiguous scope.
This is not document beautification. It protects the causal chain from parsing to chunking to retrieval. Lose the heading and a passage no longer knows which product it describes. Lose a table header and a column of numbers loses meaning. Drop the retired-version marker from metadata and retrieval may present it as current policy. See how document format affects AI retrieval quality for the full treatment.
Assemble the minimum viable corpus by question coverage, never by file count
Put candidate questions on rows and authoritative sources on columns. Mark each intersection “direct answer”, “needs combination”, “missing condition” or “no evidence”. Admit sources that cover several high-value questions, then add the exceptions and terminology that bound those answers. A hundred-page manual supporting one low-value question may rank below a one-page current returns rule. A popular spreadsheet nobody will validate does not become first-release material merely because people cite it often.
The minimum viable corpus should give every selected question a conclusion, its conditions, a source and an owner. Where evidence is missing, log the gap instead of writing a plausible answer. The priority map should make the logic visible: business questions create evidence needs; authority determines admissibility; frequency and consequence set order; format quality determines whether retrieval can use the source consistently.

Accept with real questions and close cold start with auditable deliverables
Hold back part of the question sample for acceptance so the people preparing sources cannot rewrite every test around the expected answer. Review substance rather than prose style:
- Did the system retrieve the correct domain, rather than assemble a similar answer from a neighbouring topic?
- Does the answer state its conclusion, applicable conditions and original source, with a path back to the document?
- When evidence conflicts, expires or runs out, does it refuse to conclude and identify the human route?
- For different roles, does the answer respect source visibility without citing restricted material?
- Can a wrong or missed answer return to a named source owner as a concrete correction task?
Cold start should finish with a scope statement, real-question set, source inventory, authority list, conflict and gap log, minimum viable corpus, acceptance record and owner roster — not a note saying that a number of files were uploaded. ISO 30401:2018 treats a knowledge management system as something established, implemented, maintained, reviewed and improved. These deliverables therefore need to support later review and change.
When acceptance fails, trace the chain. No source means add evidence; a present source that cannot be found means inspect parsing and chunking; stale retrieval means inspect version control; an answer nobody trusts means inspect citations and ownership. Do not collapse every failure into “the model is too weak”. Most common causes of an unusable enterprise knowledge base leave evidence during cold start. A good cold start is not the biggest collection; it knows why each source is present, who confirms it, what it can answer and when it must say it does not know.