Forge Intelligence

Experimentation and Empirical Learning Intelligence

Design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.

Updated

Definition

Design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.

Coverage role

Provides a general learn-by-intervention loop for science, engineering, product development, operations, learning, behavior change, and everyday uncertainty reduction.

Scope

Included

  • testable hypothesis formation
  • variables and outcomes
  • experimental and quasi-experimental design
  • controls randomization and blocking
  • sampling and measurement
  • protocol and stop rules
  • execution logging
  • analysis and uncertainty
  • validity assessment
  • replication and iteration

Excluded and non-goals

  • human or animal experimentation without appropriate ethics and consent
  • causal claims from uncontrolled observation
  • p-hacking outcome switching or selective reporting
  • unsafe self-experimentation
  • using one experiment as universal proof
  • replacing domain experts or regulated trials

Boundary rules

  • Route to general.experimentation when the requested outcome requires one or more declared capabilities and the output can be represented by its canonical artifacts.
  • Route to an adjacent or specialist intelligence when domain-specific rules, tools, or professional authority dominate the problem.
  • Use general.reasoning to coordinate multi-pack conflicts, hard constraints, uncertainty, and decision status.
  • Return ask, defer, or escalate when required context, source authority, consent, or accountable ownership is missing.

Capability semantics

Capability ID Name Semantic intent
capability.general.experimentation.formulate_testable_hypotheses Formulate Testable Hypotheses Turn questions and models into falsifiable predictions, rival hypotheses, and expected observation patterns.
capability.general.experimentation.define_variables_and_measures Define Variables And Measures Specify interventions, factors, outcomes, covariates, controls, units, instruments, timing, and measurement quality.
capability.general.experimentation.select_experimental_design Select Experimental Design Choose pilot, A/B, factorial, randomized, blocked, sequential, quasi-experimental, single-case, simulation, or observational designs with stated limits.
capability.general.experimentation.control_confounding_and_bias Control Confounding And Bias Plan randomization, matching, blocking, blinding, counterbalancing, negative controls, and alternative-explanation checks.
capability.general.experimentation.plan_sampling_and_stopping Plan Sampling And Stopping Define population, recruitment or case selection, sample rationale, repetitions, exclusion rules, and ethical stop conditions.
capability.general.experimentation.author_protocol_and_analysis_plan Author Protocol And Analysis Plan Freeze the intended procedure, data schema, primary outcomes, transformations, model, comparisons, and deviation policy before confirmatory execution.
capability.general.experimentation.execute_and_log_experiment Execute And Log Experiment Record conditions, interventions, observations, anomalies, deviations, versions, seeds, and custody of data or artifacts.
capability.general.experimentation.analyze_results_and_uncertainty Analyze Results And Uncertainty Use appropriate deterministic or statistical tools, report effect sizes and uncertainty, and distinguish exploratory from confirmatory findings.
capability.general.experimentation.assess_validity_replicate_and_iterate Assess Validity Replicate And Iterate Evaluate internal external construct and statistical validity, attempt reproduction or replication, update models, and design the next experiment.

Canonical artifact concepts

Artifact type Structural category Semantic purpose
artifact.general.experimentation.hypothesis_registry hypothesis Hypothesis Registry for Experimentation and Empirical Learning Intelligence: A falsifiable hypothesis record with rivals, predictions, disconfirming observations, assumptions, evidence, and decision consequence. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
artifact.general.experimentation.variable_and_measurement_dictionary reference_schema Variable And Measurement Dictionary for Experimentation and Empirical Learning Intelligence: A versioned reference structure for terms, variables, formats, units, transformations, provenance, and validation. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
artifact.general.experimentation.experimental_design_matrix model Experimental Design Matrix for Experimentation and Empirical Learning Intelligence: A purpose-bounded representation of entities, relationships, assumptions, boundaries, and evidence of adequacy. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
artifact.general.experimentation.confounding_and_bias_control_plan plan Confounding And Bias Control Plan for Experimentation and Empirical Learning Intelligence: A governed action design that sequences work, ownership, dependencies, timing, contingencies, and review. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
artifact.general.experimentation.sampling_and_stop_plan plan Sampling And Stop Plan for Experimentation and Empirical Learning Intelligence: A governed action design that sequences work, ownership, dependencies, timing, contingencies, and review. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
artifact.general.experimentation.protocol_and_analysis_plan plan Protocol And Analysis Plan for Experimentation and Empirical Learning Intelligence: A governed action design that sequences work, ownership, dependencies, timing, contingencies, and review. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
artifact.general.experimentation.experiment_run_log ledger Experiment Run Log for Experimentation and Empirical Learning Intelligence: A controlled collection of uniquely identified entries with ownership, provenance, status, and review state. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
artifact.general.experimentation.data_and_transformation_manifest reference_schema Data And Transformation Manifest for Experimentation and Empirical Learning Intelligence: A versioned reference structure for terms, variables, formats, units, transformations, provenance, and validation. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
artifact.general.experimentation.result_and_uncertainty_report analysis Result And Uncertainty Report for Experimentation and Empirical Learning Intelligence: An evidence-linked assessment that states criteria, findings, uncertainty, limitations, status, and action. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.
artifact.general.experimentation.validity_replication_and_next_experiment_plan plan Validity Replication And Next Experiment Plan for Experimentation and Empirical Learning Intelligence: A governed action design that sequences work, ownership, dependencies, timing, contingencies, and review. It supports the pack purpose of design, conduct, analyze, and iterate controlled investigations that produce new evidence about mechanisms, alternatives, parameters, users, or system behavior while preserving ethics, validity, and reproducibility.

Method families and limits

Method ID Method Use when Avoid when
method.general.experimentation.design_of_experiments design of experiments Use when a question can be converted into discriminating observations under a controlled, quasi-controlled, simulation, or explicitly observational design. Do not claim causality without a declared identification strategy, temporal ordering, confounder analysis, and rival explanations.
method.general.experimentation.randomized_controlled_comparison randomized controlled comparison Use when units can be assigned fairly, interference is controlled or modeled, outcomes are measurable, and randomization is ethical and operationally feasible. Avoid when consent, equipoise, sample adequacy, treatment integrity, or interference assumptions cannot be met; do not generalize beyond the sampled population without evidence.
method.general.experimentation.factorial_and_fractional_factorial_design factorial and fractional-factorial design Use when several controllable factors and their interactions must be estimated efficiently over declared ranges. Avoid when factor levels are unsafe, effects are strongly time-varying without blocking, or a fractional design aliases interactions that matter to the decision.
method.general.experimentation.a_b_and_sequential_testing A/B and sequential testing Use when two or more alternatives can be exposed under controlled allocation and an explicit sequential or fixed-horizon decision rule. Avoid optional stopping, metric switching, repeated peeking without correction, contaminated assignment, or experiments whose harms cannot be rapidly contained.
method.general.experimentation.quasi_experimental_design quasi-experimental design Use when random assignment is unavailable but a defensible comparison, discontinuity, timing change, instrument, matching design, or natural experiment exists. Avoid causal claims when the identifying assumption is untestable and unsupported, pre-trends fail, spillovers dominate, or treatment assignment is confounded beyond repair.
method.general.experimentation.single_case_and_within_subject_design single-case and within-subject design Use when repeated observations on the same unit can establish a stable baseline and ethically support phase changes or counterbalancing. Avoid when the condition is rapidly progressive, reversal is unsafe or impossible, measurement is unreliable, or carryover prevents phase interpretation.
method.general.experimentation.simulation_experiments simulation experiments Use when a question can be converted into discriminating observations under a controlled, quasi-controlled, simulation, or explicitly observational design. Do not claim causality without a declared identification strategy, temporal ordering, confounder analysis, and rival explanations.
method.general.experimentation.pilot_and_feasibility_studies pilot and feasibility studies Use before a high-cost study or implementation when feasibility, acceptability, recruitment, instrumentation, fidelity, or operational risk is uncertain. Do not treat a small feasibility study as a powered effectiveness trial or convert exploratory outcomes into confirmatory claims.
method.general.experimentation.falsification_and_rival_hypothesis_testing falsification and rival-hypothesis testing Use when confidence in a claim, rule, model, concept, or design depends on actively searching for observations that would make it false or too broad. Avoid weak straw-man alternatives, changing the claim after failure, or treating failure to find a counterexample as proof under an unbounded search.
method.general.experimentation.reproduction_replication_and_robustness_checks reproduction replication and robustness checks Use when confidence depends on repeating a result, independently reimplementing it, comparing it with reference data, or estimating systematic measurement error. Avoid calling a same-code rerun an independent replication, hiding deviations from the original conditions, or treating agreement within one dataset as general validity.

Pack-specific invariants

The pack also inherits the core invariants in core/core_invariants.yaml.

Invariant ID Statement Failure action
invariant.general.experimentation.01_exploratory_and_confirmatory_work_are_labeled_and_their_an Exploratory and confirmatory work are labeled and their analysis freedoms are not conflated. repair_or_conditional_pass
invariant.general.experimentation.02_hypotheses_outcomes_transformations_exclusions_and_stop_ru Hypotheses outcomes transformations exclusions and stop rules are timestamped before confirmatory analysis. repair_or_conditional_pass
invariant.general.experimentation.03_causal_claims_match_the_identification_strength_of_the_des Causal claims match the identification strength of the design and acknowledged assumptions. repair_or_conditional_pass
invariant.general.experimentation.04_consent_ethics_safety_privacy_and_data_governance_requirem Consent ethics safety privacy and data-governance requirements precede execution. repair_or_conditional_pass
invariant.general.experimentation.05_null_negative_contradictory_and_anomalous_results_remain_i Null negative contradictory and anomalous results remain in the record. repair_or_conditional_pass
invariant.general.experimentation.06_protocol_deviations_missing_data_measurement_error_and_imp Protocol deviations missing data measurement error and implementation failure are reported. repair_or_conditional_pass
invariant.general.experimentation.07_computations_are_reproducible_where_possible_and_replicati Computations are reproducible where possible and replication is distinguished from reproduction. repair_or_conditional_pass
invariant.general.experimentation.08_one_experiment_updates_evidence_it_does_not_automatically One experiment updates evidence; it does not automatically establish universal truth. repair_or_conditional_pass

Dependencies and relations

Operational general intelligences: general.question-framing, general.validation-verification
Foundation concepts: general.epistemic-evidence, general.probabilistic-statistical, general.causal-counterfactual, general.logical-constraint, general.tool-computational, general.ethical-normative
Life intelligences: None

Source roles

Source ID Claim roles Runtime resolution required
nist.engineering_statistics concept_definition, measurement_or_security_practice, risk_control False
nap.science_engineering_practices concept_definition, research_and_evidence_practice False
nap.reproducibility_replicability concept_definition, research_and_evidence_practice False
nasa.models_simulations concept_definition, engineering_method, verification_and_validation False

Risk and authority boundary

Risk class: moderate

  • The experiment involves people animals medical treatment hazardous materials safety-critical systems or legally regulated research.
  • The intervention could cause material harm discrimination coercion privacy loss or irreversible environmental impact.
  • A result will drive high-stakes policy clinical financial legal or automated decisions without independent expert review.

Accountable people retain consent, value, regulated, irreversible, and residual-risk decisions. This semantic page does not claim a deployed implementation.

Maturity

semantic: S3_reviewable_definition
contract: C3_testable_contract
implementation: I0_not_included
evidence: E2_source_roles_mapped
governance: G2_controls_and_review_defined
overlay: O2_paired_seed_complete

Paired overlay

Runtime maturity

  • Capabilities: 9
  • Technical floor→ceiling: I1_runtimeI2_integrated
  • Autonomy floor→ceiling: A0_assistedA0_assisted
  • I1_runtime count: 8
  • Evidence class: contract: 8, semantic: 1
  • Full hierarchy: MATURITY-HIERARCHY.md

Evaluation suite status

  • CI status: pass
  • Suite: evaluation_suites/experimentation.yaml
  • Cases defined: 7
  • Capabilities under test: 9