From research pool to knowledge base — Turning sources into a working system
A folder full of research is not yet project knowledge. A usable knowledge base emerges only when sources, relationships, contradictions and project results are structured so that people and AI can retrieve, inspect and use them for decisions.

After a strong research chain, something very valuable — and very inconvenient — appears: many documents. Different research runs illuminate different perspectives, overlap, sometimes contradict one another and often operate at different levels of detail. This is where a second stage begins, one that is easy to underestimate in AI project management.
The question is no longer: What else can we research? It becomes: How do we turn the collected material into a knowledge system the project can actually work with?
This transition can be observed very concretely: dozens of Deep Research documents get brought together, semantically linked and transferred into a shared knowledge base. The important principle is independent of the tool: a research pool is raw material. A knowledge base is a working architecture.
A research pool is valuable, but not yet controllable
Imagine 40, 60 or 100 research files. Every file may be useful. Yet practical problems appear immediately:
• Which file answers which question?
• Which claims are merely repeated?
• Where do sources disagree?
• Which finding is established and which is only plausible?
• Which information belongs to market, technology, risk or requirements?
• Which finding has already been translated into a project decision?
• Which source is outdated or now only historically relevant?
A person can still browse a small collection manually. As the project grows, that becomes increasingly expensive. AI can read larger volumes faster, but it also needs structure if provenance and context are not to disappear.
The core problem is therefore not the number of files. It is the lack of orientation between files.
A knowledge base has at least three layers
For AI projects, it helps to think of a knowledge base not as one large repository but as three connected layers.
1. The source layer
This is where originals live: research documents, studies, meeting notes, internal analyses, customer feedback, technical notes, results and later project artifacts.
This layer should be edited as little as possible. An original loses value when nobody can tell what the source actually said and what was later summarized or interpreted.
At minimum, a source should carry: a unique name or stable ID; provenance; date or version; topic or validity range; status: current, unreviewed, superseded, historical; sensitivity or access boundary where needed.
The source layer is the project's evidence memory.
2. The semantic layer
This is where relationships are created. A data-quality research note is linked to a technical feasibility analysis. A user observation points to a requirement. Two sources are marked as conflicting. Several research documents are grouped under one shared concept.
The semantic layer answers questions such as: Which documents discuss the same concept? Which findings support the same requirement? Which claims contradict one another? Which information belongs to one common risk? Which source is particularly relevant to a specific decision?
This layer turns storage into orientation.
3. The operational layer
This is where knowledge becomes actionable. Findings are translated into decisions, requirements, tasks, risks, checks and project artifacts.
The team no longer asks only, "What do our sources say?" It asks: Which findings change our scope? Which open questions block the next decision? Which requirements are sufficiently supported? Which assumption needs a test? Which project results must flow back into the knowledge base?
Only this layer turns a collection of knowledge into a project instrument.
Preserve the original; add the compression above it
A common consolidation mistake is to merge many sources into one huge "master summary". It looks tidy, but can destroy exactly the distinctions that may matter later.
A stronger principle is two-layered: Keep the original. Add a distilled layer.
A short source note can sit above an original document:
# SOURCE NOTE
Source: R-017
Topic: Data requirements for a local AI system
## Key claims
- ...
- ...
## Relevance to the project
- ...
## Boundaries / open questions
- ...
## Links
- [[Data quality]]
- [[Local architecture]]
- [[Risk: outdated information]]The short note is not the source. It is a navigation layer to the source.
That matters for humans and AI agents alike. A summary is fast to read. When a claim becomes decision-relevant, the path back to the original must remain open.
Semantic linking is more than folder structure
Folders mainly answer: Where is something stored? A knowledge architecture must also answer: What is it related to?
One file may simultaneously be relevant to:
• technical feasibility;
• privacy;
• cost;
• user experience;
• procurement;
• project risk.
A classic folder hierarchy forces one location or duplicate copies. Semantic links allow the same source to remain visible in several contexts.
This matters especially in AI projects because many decisions run across disciplines. An architecture choice may affect cost, data handling, team capability and later governance at the same time.
Knowledge is therefore closer to a network than a shelf.
Contradictions belong in the knowledge base — not under the carpet
When two research documents reach different conclusions, it is tempting to select the "better" statement and let the other disappear. For project work, that is dangerous.
Contradictions can arise from: different dates; different definitions; different target groups; different assumptions; unequal data quality; different measurement methods.
A knowledge system should therefore keep conflict visible.
# CONFLICT NOTE
Topic: Expected operating cost
Source A:
- Claim ...
Source B:
- Claim ...
Possible cause of the difference:
- ...
What must be checked?
- ...
Impact on decision:
- ...The contradiction itself becomes a work item. The system prevents uncertainty from disappearing inside a polished summary.
A knowledge base needs maps, not just files
As a collection grows, entry points become essential. Nobody should have to scan 80 filenames to understand how project knowledge is organized.
Maps of Content, overview pages or thematic maps are useful. They do not duplicate the entire content; they reveal the most important relationships.
A "Technical Architecture" map might contain:
# MAP — TECHNICAL ARCHITECTURE
## Foundations
- [[R-004 System requirements]]
- [[R-011 Infrastructure comparison]]
## Decisions
- [[D-003 Hybrid architecture]]
## Risks
- [[RK-007 Vendor dependency]]
- [[RK-012 Data leakage]]
## Open questions
- [[Q-018 Operating cost]]
- [[Q-024 Backup strategy]]Such a map does not replace the documents. It is a reading lens for the knowledge space.
Structure should grow from the content
A key point stands out here: an existing database structure should not be copied blindly just because it already exists. The architecture should fit the material and the project.
That means: do not invent 30 folders first and then force every source into them. Instead:
1. read the sources;
2. identify recurring themes;
3. identify meaningful relationships;
4. derive categories, maps and links from them;
5. simplify the structure later if it becomes too complex.
This is closer to real knowledge work. It avoids a common failure: a technically neat archive that explains nothing.
Compress without losing provenance
Real projects have platform limits. A browser workspace may accept only a limited number of files. A model has a finite context window. A colleague may be working on a mobile device and cannot load the entire knowledge space.
Compression is therefore useful — but it must be controlled.
A robust chain is: Source → extraction → synthesis → working context. Not: Source → summary → original forgotten.
For a smaller working context, 60 sources might be compressed into ten thematic syntheses. Those syntheses must continue to point back to the original sources. That preserves traceability.
The compressed layer is a transport format for knowledge, not its only store.
The database does not end when research ends
One of the strongest ideas from the course is that the knowledge base does not merely support project initiation. It accompanies the project throughout its lifecycle.
As work proceeds, new artifacts are added: feasibility analyses; decisions and rationales; tests and test results; customer feedback; change requests; errors and incidents; new research; status reports; retrospectives; final results.
The research base gradually becomes a project memory.
The knowledge base then connects planning, execution, review and future learning. The course used the image of a project's "spinal cord": information does not merely flow in; it connects the different parts of the work.
AI becomes an archivist — but only with explicit rules
A well-structured knowledge base lets AI recognise relationships more quickly. It can answer questions such as: Which decision depends on this source? Which open risks relate to this requirement? Which findings were added since the previous planning version? Which sources contradict the current assumption? Which documents must be read for a review?
But this does not mean AI automatically becomes a reliable knowledge manager. It needs rules:
• original sources are not silently overwritten;
• syntheses remain distinguishable from originals;
• contradictions are marked rather than smoothed away;
• unclear provenance is treated as uncertainty;
• outdated information is not treated as current;
• sensitive information has explicit access boundaries.
The quality of AI access depends directly on the quality of the knowledge architecture.
Think tool-neutral: method before Obsidian or Drive
Obsidian is very useful for linked Markdown knowledge spaces. A normal file system may be enough for a small project. Specialized knowledge-graph systems can also make sense.
The method comes first.
A minimal architecture could look like this:
01 Sources
02 Extractions
03 Concepts
04 Conflicts
05 Decisions
06 Risks
07 Project outputs
08 OverviewsThe folder names matter less than four properties:
1. Provenance remains visible.
2. Relationships are traceable.
3. Status and currency are recognisable.
4. Knowledge can be translated into a concrete project decision.
If those four properties exist, the structure can later be moved into Obsidian or another system without reinventing the logic.
A practical eight-step workflow
Step 1 — Inventory the sources
Give each file a stable identifier, provenance, date and broad topic.
Step 2 — Detect duplicates and variants
Mark identical or near-identical documents. Keep older versions discoverable but make their status clear.
Step 3 — Extract decision-relevant claims
Do not summarize everything. Extract only claims, boundaries, examples and open questions that can change project decisions.
Step 4 — Build concepts and themes
Turn recurring ideas into concept cards. A source can be linked to more than one concept.
Step 5 — Mark contradictions
Do not immediately resolve conflicting claims. First investigate why they differ and whether the difference matters to a decision.
Step 6 — Build overview maps
Create maps or indexes for major project areas so humans and AI can enter the knowledge space quickly.
Step 7 — Link project decisions
Decisions point back to the sources and syntheses on which they are based. The team can then see why the current plan looks the way it does.
Step 8 — Maintain the knowledge during execution
Feed new results, tests, feedback and changes back into the system. The knowledge base remains working material rather than a static archive.
When is a knowledge base operational?
Not when it is large. Not when it looks impressive. And not when every document has been filed somewhere.
It is operational when a person or agent can reliably answer: Which source supports this claim? What other information is connected to it? Where is there uncertainty or contradiction? Which version is current? Which decision was derived from it? Which next question or task follows from it?
Research then changes role. It is no longer a preliminary phase that disappears after the project plan. It becomes the first part of a continuous knowledge loop.
Conclusion: not more files, but more relationships
AI makes it easy to create large volumes of research. That makes the second capability more important: organizing knowledge without destroying its provenance, differences and uncertainty.
A research pool gathers perspectives. A knowledge base makes them retrievable, connected and actionable. It separates originals from syntheses, preserves contradictions, links findings to decisions and continues to grow during the project.
The decisive step is therefore not to load as much material as possible into one tool. It is to build a knowledge space in which every important finding has a location, a provenance, a relationship and a next use.
The next article can build directly on this foundation: once different people or projects have contributed different approaches, experiences and solutions, the shared knowledge base can be used to create a collective method map.
Worksheet: Build a mini knowledge base from six research documents
Choose six research or project documents on one shared topic. The goal is not a perfect vault but a small, traceable knowledge architecture.
1. Inventory the sources
Assign each file an ID and record provenance, date, topic and status.
2. Extract the key claims
Write no more than five claims per source that could change a project decision.
3. Create three concept cards
Identify three recurring themes and link every source to at least one concept card.
4. Document one contradiction
Find two claims that do not fully agree. Write the likely cause, open question and possible project impact.
5. Build one overview page
Create a page from which sources, concepts, the contradiction and open questions can be reached directly.
6. Link one decision
Write one small project decision and name the sources on which it depends.
7. Define the feedback loop
Decide which future project results must be added back to the knowledge base.
Reflection
Which information would most likely have disappeared in a simple folder structure? _________________________________________________
Which relationship between two sources changes your project decision most strongly? _______________________________________________
All materials to download — the topic overview and the worksheet:
0 comments
● Loading comments…