← System handbook

Foundry, Intake & Knowledge Search

Implementedknowledge

The governed route from uncontrolled historical material into Primitive OS.

TL;DR

Foundry already owns estate discovery, universal source tagging, filing, chunking, hybrid search, candidate review, Google Drive canonical baselines and versioned OS intake bundles. Docs and Workspace should reuse that machinery rather than building another crawler/RAG pipeline.

ELI5

Foundry is the loading dock. It finds old boxes, labels them, checks what they are, lets you approve what should come inside, puts an official copy in the filing room (Google Drive), and only then gives Primitive OS a receipt and a structured package.

Synthesis

This is the anti-fragmentation boundary for the planned second brain. primitive.tags/v1 stays attached to the source through Foundry, GitHub, Drive and OS intake. primitive.filing/v1 is a changeable navigation decision, not a rewrite of source identity. Search already combines lexical, embedding and tag signals, so the Docs 'Ask Primitive' layer should consume this governed index after integration review rather than inventing a second ingestion stack.

Pipeline

One route from estate to OS

Foundry intake
DiscoverFiles, repos, Drive, chats, archives
Index + tagprimitive.tags/v1 + provenance
Review + sealFreeze exact reviewed estate state
Intake candidatesApprove / reject / route
Google Drive baselineVerified canonical bytes + manifests
OS bundleVersioned, idempotent intake contract
OS promotionKnowledge / Labs / Capabilities
Search

The existing hybrid knowledge engine

Foundry v0.6 chunks source text and combines SQLite FTS5 lexical retrieval, an embedding-provider score and universal-tag overlap. The bundled primitive-feature-hash-384-v1 embedding is a deterministic local baseline, not a neural semantic model.

SignalDefault weight
Lexical / BM250.48
Embedding0.42
Universal-tag overlap0.10
Drive invariant

Google Drive is the human-accessible baseline for promoted knowledge

Anything referenced by Primitive Knowledge/Labs/capability intake must first have a verified canonical Drive copy/receipt. The original source remains provenance; the OS need not duplicate source bytes when Drive is the baseline.

Primitive Knowledge in Drive
01_SOURCESCanonical source bytes
02_CATALOGUECategory link records; no duplicate bytes
03_INDEXShared taxonomy
_manifestsTags + provenance + filing receipts
Boundary

What Docs should do next — and what it should not

  • Docs may provide a polished Ask/Search surface over the governed Foundry/OS index.
  • Docs should not rescan Google Drive or local machines.
  • Docs should not create a competing tag schema.
  • Docs should not promote a retrieved source into accepted knowledge.
  • Neural embeddings can be added behind Foundry's EmbeddingProvider without changing source IDs or tags.
Keep reading

Related pages

Reference / evidence

Canonical sources

These pages are explanatory read models. Where wording conflicts with implementation or a Control decision, the linked implementation/decision is authoritative and this page should be corrected.