Labs / evidence status
Experiments
What Primitive is testing, what evidence actually exists, and what has not earned the right to be described as fact.
HypothesisA claim worth testing
ExperimentProtocol + falsifier
EvidenceObserved result + provenance
ReviewInterpret limits / contradictions
Promote / park / refuteExplicit lifecycle decision
Experiment
27-state E/P/A semantic cube
- Question
- Can a compact ternary state representation remain explicit about reference vs unknown?
- Evidence now
- Implemented in Workspace main; tests enforce E/P/A ∈ {-1,0,+1}, Q00–Q26 and UNKNOWN outside the cube.
- Next / falsifier
- Keep stable. New chart semantics must be versioned rather than silently changing E/P/A.
Experiment
Holarchy reading experiment
- Question
- Does hierarchical multi-reader synthesis preserve consensus and reveal useful disagreements/new combinations?
- Evidence now
- 50-call experiment across Sonnet, Luna, DeepSeek, Opus and Sol. Consensus core survived; hierarchy added combinations/disagreement maps; controls were stronger on fidelity/fact-checking.
- Next / falsifier
- Treat order effects and ±3/7-dimension links as open hypotheses, not laws. Give higher-level synthesisers source spot-check access.
Experiment
Microproof Swarm
- Question
- Which parts of a recursive proof/settlement swarm are useful for Primitive's real control plane?
- Evidence now
- Control review adopted deterministic kernel/records/lifecycle concepts, deferred ZK/settlement, and rejected stake/slashing/token governance for Primitive's principal-led authority model.
- Next / falsifier
- Continue with bounded permits, receipts and kernel admission; revisit ZK only when a real cross-organisation predicate-proof need appears.
Experiment
Primitive Reflex v1
- Question
- Can a fast, bounded System-1 router select experts/skills/tools/context and abstain safely before governed execution?
- Evidence now
- Provider-neutral core, service, tests, trace artifacts and CLI exist on feature/primitive-reflex-v1. The contract explicitly grants no execution authority.
- Next / falsifier
- Review/merge through the canonical Workspace train; then evaluate routing quality and downstream task outcomes.
Experiment
Neural Reflex / specialist adapters
- Question
- Can validated routing procedures be compressed into small learned specialist adapters with calibrated abstention?
- Evidence now
- Runnable training/benchmark scaffold exists, but its README explicitly claims no trained adapter, benchmark result or optimal topology.
- Next / falsifier
- Human-reviewed/consented dataset, leakage-safe splits, held-out comparison of anchor-only, 0±1, 0±2, 0±3 and conventional top-k, then downstream outcome measurement.
Experiment
Route benchmark / F5-BENCH
- Question
- Which model/provider routes are cheapest while meeting task-specific quality thresholds?
- Evidence now
- A preregistered deterministic four-class pilot is specified on the canonical Bench Runtime. The packet explicitly says the small pilot cannot certify a route; it only screens/parks candidates.
- Next / falsifier
- Run only after route pinning/ZDR support exists; collect evidence into the Router rather than manually promoting leaderboard winners.
Experiment
Foundry Knowledge/Search v0.6–v0.7
- Question
- Can one governed intake/search layer preserve source identity while supporting filing, hybrid retrieval and OS ingestion?
- Evidence now
- Foundry implements primitive.tags/v1, primitive.filing/v1, chunking, FTS5, local vector baseline, hybrid ranking, Drive baseline and OS bundles.
- Next / falsifier
- Integrate the existing Foundry boundary into the OS/Docs Ask surface rather than creating another ingestion pipeline.
Experiment
Environment / Relation / Context cubes
- Question
- Do three coordinated frame-relative charts improve cross-domain reasoning without overwriting the canonical E/P/A representation?
- Evidence now
- Defined as an experiment in the recursive-systems architecture. Exact axis semantics and usefulness are not canonical and require Bench evidence.
- Next / falsifier
- Define each chart/version/reference/validity domain and benchmark retrieval/reasoning benefit before adoption.
Future automation
This page should eventually be generated from Labs + Foundry evidence
The current list is curated from canonical sources so it is useful immediately. The long-term implementation should derive experiment state from Labs objects, Bench receipts and Foundry intake instead of maintaining a second manual status database in Docs.