Labs / evidence status

Experiments

What Primitive is testing, what evidence actually exists, and what has not earned the right to be described as fact.

Evidence maturityNot every experiment should become a feature.
HypothesisA claim worth testing
ExperimentProtocol + falsifier
EvidenceObserved result + provenance
ReviewInterpret limits / contradictions
Promote / park / refuteExplicit lifecycle decision
Experiment

27-state E/P/A semantic cube

LOCKED / STABLE
Question
Can a compact ternary state representation remain explicit about reference vs unknown?
Evidence now
Implemented in Workspace main; tests enforce E/P/A ∈ {-1,0,+1}, Q00–Q26 and UNKNOWN outside the cube.
Next / falsifier
Keep stable. New chart semantics must be versioned rather than silently changing E/P/A.
Experiment

Holarchy reading experiment

COMPLETED — evidence retained
Question
Does hierarchical multi-reader synthesis preserve consensus and reveal useful disagreements/new combinations?
Evidence now
50-call experiment across Sonnet, Luna, DeepSeek, Opus and Sol. Consensus core survived; hierarchy added combinations/disagreement maps; controls were stronger on fidelity/fact-checking.
Next / falsifier
Treat order effects and ±3/7-dimension links as open hypotheses, not laws. Give higher-level synthesisers source spot-check access.
Experiment

Microproof Swarm

REVIEWED — kernel adopted, crypto deferred
Question
Which parts of a recursive proof/settlement swarm are useful for Primitive's real control plane?
Evidence now
Control review adopted deterministic kernel/records/lifecycle concepts, deferred ZK/settlement, and rejected stake/slashing/token governance for Primitive's principal-led authority model.
Next / falsifier
Continue with bounded permits, receipts and kernel admission; revisit ZK only when a real cross-organisation predicate-proof need appears.
Experiment

Primitive Reflex v1

FEATURE BRANCH — not canonical main
Question
Can a fast, bounded System-1 router select experts/skills/tools/context and abstain safely before governed execution?
Evidence now
Provider-neutral core, service, tests, trace artifacts and CLI exist on feature/primitive-reflex-v1. The contract explicitly grants no execution authority.
Next / falsifier
Review/merge through the canonical Workspace train; then evaluate routing quality and downstream task outcomes.
Experiment

Neural Reflex / specialist adapters

EXPERIMENTAL — no result claimed
Question
Can validated routing procedures be compressed into small learned specialist adapters with calibrated abstention?
Evidence now
Runnable training/benchmark scaffold exists, but its README explicitly claims no trained adapter, benchmark result or optimal topology.
Next / falsifier
Human-reviewed/consented dataset, leakage-safe splits, held-out comparison of anchor-only, 0±1, 0±2, 0±3 and conventional top-k, then downstream outcome measurement.
Experiment

Route benchmark / F5-BENCH

SPECIFIED PILOT
Question
Which model/provider routes are cheapest while meeting task-specific quality thresholds?
Evidence now
A preregistered deterministic four-class pilot is specified on the canonical Bench Runtime. The packet explicitly says the small pilot cannot certify a route; it only screens/parks candidates.
Next / falsifier
Run only after route pinning/ZDR support exists; collect evidence into the Router rather than manually promoting leaderboard winners.
Experiment

Foundry Knowledge/Search v0.6–v0.7

IMPLEMENTED
Question
Can one governed intake/search layer preserve source identity while supporting filing, hybrid retrieval and OS ingestion?
Evidence now
Foundry implements primitive.tags/v1, primitive.filing/v1, chunking, FTS5, local vector baseline, hybrid ranking, Drive baseline and OS bundles.
Next / falsifier
Integrate the existing Foundry boundary into the OS/Docs Ask surface rather than creating another ingestion pipeline.
Experiment

Environment / Relation / Context cubes

WORKING THEORY / EXPERIMENT
Question
Do three coordinated frame-relative charts improve cross-domain reasoning without overwriting the canonical E/P/A representation?
Evidence now
Defined as an experiment in the recursive-systems architecture. Exact axis semantics and usefulness are not canonical and require Bench evidence.
Next / falsifier
Define each chart/version/reference/validity domain and benchmark retrieval/reasoning benefit before adoption.
Future automation

This page should eventually be generated from Labs + Foundry evidence

The current list is curated from canonical sources so it is useful immediately. The long-term implementation should derive experiment state from Labs objects, Bench receipts and Foundry intake instead of maintaining a second manual status database in Docs.