Handbook
Experimentation and Empirical Learning Intelligence
Design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
Updated
Definition
Design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
Coverage role
Provides a general learn-by-intervention loop for science, engineering, product development, operations, learning, behavior change, and everyday uncertainty reduction.
Scope
Included
- testable hypothesis formation
- variables and outcomes
- experimental and quasi-experimental design
- controls randomization and blocking
- sampling and measurement
- protocol and stop rules
- execution logging
- analysis and uncertainty
- validity assessment
- replication and iteration
Excluded and non-goals
- human or animal experimentation without appropriate ethics and consent
- causal claims from uncontrolled observation
- p-hacking outcome switching or selective reporting
- unsafe self-experimentation
- using one experiment as universal proof
- replacing domain experts or regulated trials
Boundary rules
- Route to general.experimentation when the requested outcome requires one or more declared capabilities and the output can be represented by its canonical artifacts.
- Route to an adjacent or specialist intelligence when domain-specific rules, tools, or professional authority dominate the problem.
- Use general.reasoning to coordinate multi-pack conflicts, hard constraints, uncertainty, and decision status.
- Return ask, defer, or escalate when required context, source authority, consent, or accountable ownership is missing.
Capability semantics
| Capability ID | Name | Semantic intent |
|---|---|---|
capability.general.experimentation.formulate_testable_hypotheses |
Formulate Testable Hypotheses | Turn questions and models into falsifiable predictions, rival hypotheses, and expected observation patterns. |
capability.general.experimentation.define_variables_and_measures |
Define Variables And Measures | Specify interventions, factors, outcomes, covariates, controls, units, instruments, timing, and measurement quality. |
capability.general.experimentation.select_experimental_design |
Select Experimental Design | Choose pilot, A/B, factorial, randomized, blocked, sequential, quasi-experimental, single-case, simulation, or observational designs with stated limits. |
capability.general.experimentation.control_confounding_and_bias |
Control Confounding And Bias | Plan randomization, matching, blocking, blinding, counterbalancing, negative controls, and alternative-explanation checks. |
capability.general.experimentation.plan_sampling_and_stopping |
Plan Sampling And Stopping | Define population, recruitment or case selection, sample rationale, repetitions, exclusion rules, and ethical stop conditions. |
capability.general.experimentation.author_protocol_and_analysis_plan |
Author Protocol And Analysis Plan | Freeze the intended procedure, data schema, primary outcomes, transformations, model, comparisons, and deviation policy before confirmatory execution. |
capability.general.experimentation.execute_and_log_experiment |
Execute And Log Experiment | Record conditions, interventions, observations, anomalies, deviations, versions, seeds, and custody of data or artifacts. |
capability.general.experimentation.analyze_results_and_uncertainty |
Analyze Results And Uncertainty | Use appropriate deterministic or statistical tools, report effect sizes and uncertainty, and distinguish exploratory from confirmatory findings. |
capability.general.experimentation.assess_validity_replicate_and_iterate |
Assess Validity Replicate And Iterate | Evaluate internal external construct and statistical validity, attempt reproduction or replication, update models, and design the next experiment. |
Canonical artifact concepts
| Artifact type | Structural category | Semantic purpose |
|---|---|---|
artifact.general.experimentation.hypothesis_registry |
hypothesis | Hypothesis Registry for Experimentation and Empirical Learning Intelligence: A falsifiable hypothesis record with rivals, predictions, disconfirming observations, assumptions, evidence, and decision consequence. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
artifact.general.experimentation.variable_and_measurement_dictionary |
reference_schema | Variable And Measurement Dictionary for Experimentation and Empirical Learning Intelligence: A versioned reference structure for terms, variables, formats, units, transformations, provenance, and validation. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
artifact.general.experimentation.experimental_design_matrix |
model | Experimental Design Matrix for Experimentation and Empirical Learning Intelligence: A purpose-bounded representation of entities, relationships, assumptions, boundaries, and evidence of adequacy. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
artifact.general.experimentation.confounding_and_bias_control_plan |
plan | Confounding And Bias Control Plan for Experimentation and Empirical Learning Intelligence: A governed action design that sequences work, ownership, dependencies, timing, contingencies, and review. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
artifact.general.experimentation.sampling_and_stop_plan |
plan | Sampling And Stop Plan for Experimentation and Empirical Learning Intelligence: A governed action design that sequences work, ownership, dependencies, timing, contingencies, and review. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
artifact.general.experimentation.protocol_and_analysis_plan |
plan | Protocol And Analysis Plan for Experimentation and Empirical Learning Intelligence: A governed action design that sequences work, ownership, dependencies, timing, contingencies, and review. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
artifact.general.experimentation.experiment_run_log |
ledger | Experiment Run Log for Experimentation and Empirical Learning Intelligence: A controlled collection of uniquely identified entries with ownership, provenance, status, and review state. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
artifact.general.experimentation.data_and_transformation_manifest |
reference_schema | Data And Transformation Manifest for Experimentation and Empirical Learning Intelligence: A versioned reference structure for terms, variables, formats, units, transformations, provenance, and validation. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
artifact.general.experimentation.result_and_uncertainty_report |
analysis | Result And Uncertainty Report for Experimentation and Empirical Learning Intelligence: An evidence-linked assessment that states criteria, findings, uncertainty, limitations, status, and action. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
artifact.general.experimentation.validity_replication_and_next_experiment_plan |
plan | Validity Replication And Next Experiment Plan for Experimentation and Empirical Learning Intelligence: A governed action design that sequences work, ownership, dependencies, timing, contingencies, and review. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility. |
Method families and limits
| Method ID | Method | Use when | Avoid when |
|---|---|---|---|
method.general.experimentation.design_of_experiments |
design of experiments | Use when a question can be converted into discriminating observations under a controlled, quasi-controlled, simulation, or explicitly observational design. | Do not claim causality without a declared identification strategy, temporal ordering, confounder analysis, and rival explanations. |
method.general.experimentation.randomized_controlled_comparison |
randomized controlled comparison | Use when units can be assigned fairly, interference is controlled or modeled, outcomes are measurable, and randomization is ethical and operationally feasible. | Avoid when consent, equipoise, sample adequacy, treatment integrity, or interference assumptions cannot be met; do not generalize beyond the sampled population without evidence. |
method.general.experimentation.factorial_and_fractional_factorial_design |
factorial and fractional-factorial design | Use when several controllable factors and their interactions must be estimated efficiently over declared ranges. | Avoid when factor levels are unsafe, effects are strongly time-varying without blocking, or a fractional design aliases interactions that matter to the decision. |
method.general.experimentation.a_b_and_sequential_testing |
A/B and sequential testing | Use when two or more alternatives can be exposed under controlled allocation and an explicit sequential or fixed-horizon decision rule. | Avoid optional stopping, metric switching, repeated peeking without correction, contaminated assignment, or experiments whose harms cannot be rapidly contained. |
method.general.experimentation.quasi_experimental_design |
quasi-experimental design | Use when random assignment is unavailable but a defensible comparison, discontinuity, timing change, instrument, matching design, or natural experiment exists. | Avoid causal claims when the identifying assumption is untestable and unsupported, pre-trends fail, spillovers dominate, or treatment assignment is confounded beyond repair. |
method.general.experimentation.single_case_and_within_subject_design |
single-case and within-subject design | Use when repeated observations on the same unit can establish a stable baseline and ethically support phase changes or counterbalancing. | Avoid when the condition is rapidly progressive, reversal is unsafe or impossible, measurement is unreliable, or carryover prevents phase interpretation. |
method.general.experimentation.simulation_experiments |
simulation experiments | Use when a question can be converted into discriminating observations under a controlled, quasi-controlled, simulation, or explicitly observational design. | Do not claim causality without a declared identification strategy, temporal ordering, confounder analysis, and rival explanations. |
method.general.experimentation.pilot_and_feasibility_studies |
pilot and feasibility studies | Use before a high-cost study or implementation when feasibility, acceptability, recruitment, instrumentation, fidelity, or operational risk is uncertain. | Do not treat a small feasibility study as a powered effectiveness trial or convert exploratory outcomes into confirmatory claims. |
method.general.experimentation.falsification_and_rival_hypothesis_testing |
falsification and rival-hypothesis testing | Use when confidence in a claim, rule, model, concept, or design depends on actively searching for observations that would make it false or too broad. | Avoid weak straw-man alternatives, changing the claim after failure, or treating failure to find a counterexample as proof under an unbounded search. |
method.general.experimentation.reproduction_replication_and_robustness_checks |
reproduction replication and robustness checks | Use when confidence depends on repeating a result, independently reimplementing it, comparing it with reference data, or estimating systematic measurement error. | Avoid calling a same-code rerun an independent replication, hiding deviations from the original conditions, or treating agreement within one dataset as general validity. |
Pack-specific invariants
The pack also inherits the core invariants in core/core_invariants.yaml.
| Invariant ID | Statement | Failure action |
|---|---|---|
invariant.general.experimentation.01_exploratory_and_confirmatory_work_are_labeled_and_their_an |
Exploratory and confirmatory work are labeled and their analysis freedoms are not conflated. | repair_or_conditional_pass |
invariant.general.experimentation.02_hypotheses_outcomes_transformations_exclusions_and_stop_ru |
Hypotheses outcomes transformations exclusions and stop rules are timestamped before confirmatory analysis. | repair_or_conditional_pass |
invariant.general.experimentation.03_causal_claims_match_the_identification_strength_of_the_des |
Causal claims match the identification strength of the design and acknowledged assumptions. | repair_or_conditional_pass |
invariant.general.experimentation.04_consent_ethics_safety_privacy_and_data_governance_requirem |
Consent ethics safety privacy and data-governance requirements precede execution. | repair_or_conditional_pass |
invariant.general.experimentation.05_null_negative_contradictory_and_anomalous_results_remain_i |
Null negative contradictory and anomalous results remain in the record. | repair_or_conditional_pass |
invariant.general.experimentation.06_protocol_deviations_missing_data_measurement_error_and_imp |
Protocol deviations missing data measurement error and implementation failure are reported. | repair_or_conditional_pass |
invariant.general.experimentation.07_computations_are_reproducible_where_possible_and_replicati |
Computations are reproducible where possible and replication is distinguished from reproduction. | repair_or_conditional_pass |
invariant.general.experimentation.08_one_experiment_updates_evidence_it_does_not_automatically |
One experiment updates evidence; it does not automatically establish universal truth. | repair_or_conditional_pass |
Dependencies and relations
Operational general intelligences: general.question-framing, general.validation-verification
Foundation concepts: general.epistemic-evidence, general.probabilistic-statistical, general.causal-counterfactual, general.logical-constraint, general.tool-computational, general.ethical-normative
Life intelligences: None
Source roles
| Source ID | Claim roles | Runtime resolution required |
|---|---|---|
nist.engineering_statistics |
concept_definition, measurement_or_security_practice, risk_control | False |
nap.science_engineering_practices |
concept_definition, research_and_evidence_practice | False |
nap.reproducibility_replicability |
concept_definition, research_and_evidence_practice | False |
nasa.models_simulations |
concept_definition, engineering_method, verification_and_validation | False |
Risk and authority boundary
Risk class: moderate
- The experiment involves people animals medical treatment hazardous materials safety-critical systems or legally regulated research.
- The intervention could cause material harm discrimination coercion privacy loss or irreversible environmental impact.
- A result will drive high-stakes policy clinical financial legal or automated decisions without independent expert review.
Accountable people retain consent, value, regulated, irreversible, and residual-risk decisions. This semantic page does not claim a deployed implementation.
Maturity
semantic: S3_reviewable_definition
contract: C3_testable_contract
implementation: I0_not_included
evidence: E2_source_roles_mapped
governance: G2_controls_and_review_defined
overlay: O2_paired_seed_complete
Paired overlay
Runtime maturity
- Capabilities: 9
- Technical floor→ceiling:
I1_runtime→I2_integrated - Autonomy floor→ceiling:
A0_assisted→A0_assisted - I1_runtime count: 8
- Evidence class:
contract: 8,semantic: 1 - Full hierarchy: MATURITY-HIERARCHY.md
Evaluation suite status
- CI status:
pass - Suite:
evaluation_suites/experimentation.yaml - Cases defined: 7
- Capabilities under test: 9