Files
linus-d 659544f0b2
magicka-merge-gates / advisory-fast (push) Has been skipped
magicka-merge-gates / merge-gates (push) Failing after 38s
update
2026-06-21 19:21:48 -07:00

333 lines
9.4 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
ChatGPT3:26 PM
Rust Phase 0/1 Implementor Specification
Objective
Build the testing framework first, then the runtime.
No spell content is accepted until CI proves the system resists collapse into:
single score
single resource
single effect axis
single executor behavior
single rune behavior
single hidden damage formula
decorative world state
Compliance Model Anti-Collapse
The framework must also resist collapse of the compliance model.
The implementor may not replace a specification obligation with a representative approximation that merely resembles the obligation.
Every acceptance requirement must have all of:
1. A measured artifact.
2. A provenance chain from the measured artifact to the run that produced it.
3. A merge-blocking enforcement point.
4. A failure condition that blocks acceptance if the artifact or provenance is absent.
No requirement may be satisfied by:
sample
summary
approximation
representative subset
default profile
proxy metric
regenerated artifact
local-only executable
documentation claim
unless this specification explicitly defines that weaker substitute as acceptable for that requirement.
The compliance path must be:
Specification requirement
Mandatory enforcement mechanism
Merge blocked if absent
The compliance path must not be:
Specification requirement
Representative approximation
Evidence of approximation
Examples of forbidden substitutions:
100% reference/runtime comparison may not be replaced by comparison of base executions only.
Required artifacts may not be replaced by partial artifacts.
Full trace information may not be replaced by a summarized proxy unless that proxy is explicitly named as the acceptance artifact.
Persisted expectations may not be replaced by expectations regenerated in the same run.
Merge-blocking enforcement may not be replaced by a manually runnable local binary.
If an implementation uses a weaker substitute, the correct result is not partial credit. The correct result is failure of the corresponding acceptance gate.
Hard CI Gates
Minimum per full CI run:
Generated worlds: 50,000
Generated rune programs: 250,000
Executions: 1,000,000
Perturbations per execution: 10
Semantic mutants per run: 500
Replay corpus cases: 10,000 minimum
Reference/runtime comparison: 100% of executions
Fast CI may run 10%, but merge-blocking CI must run full gates.
Repository Layout
crates/
world_model/
rune_ir/
trace_model/
generators/
reference_runtime/
runtime_under_test/
collapse_analysis/
semantic_mutation/
replay_corpus/
ci_reports/
Implementation order is mandatory:
1. world_model
2. trace_model
3. generators
4. collapse_analysis
5. semantic_mutation
6. replay_corpus
7. reference_runtime
8. runtime_under_test
The optimized runtime may not begin before steps 17 pass CI.
World Model
pub struct WorldSnapshot {
pub id: WorldId,
pub turn: u64,
pub domains: Vec<DomainState>,
pub causal_state: CausalState,
pub observation_state: ObservationState,
pub execution_state: ExecutionState,
pub time_state: TimeState,
pub seed: u64,
}
Minimum domain count:
8 independent world domains
Each domain must expose:
pub trait WorldDomain {
fn domain_id(&self) -> DomainId;
fn read_surface(&self) -> ReadSurface;
fn write_surface(&self) -> WriteSurface;
fn perturbation_axes(&self) -> Vec<PerturbationAxis>;
fn fingerprint(&self) -> DomainFingerprint;
}
Domain acceptance:
Each domain appears in ≥ 35% of traces.
Each domain influences execution in ≥ 20% of corpus.
Each domain is mutated by execution in ≥ 20% of corpus.
Removing any domain reduces corpus behavioral diversity by ≥ 10%.
Merging any two domains loses ≥ 8% predictive accuracy.
Rune Program Model
pub struct RuneProgram {
pub id: ProgramId,
pub tokens: Vec<RuneToken>,
pub seed: u64,
}
No rune stream may be rejected as the primary safety path.
Execution always returns:
pub struct ResolutionResult {
pub delta: WorldDelta,
pub trace: ExecutionTrace,
pub faults: FaultLog,
pub replay: ReplayRecord,
}
Engine crashes, panics, undefined Rust behavior, or unlogged failures fail CI.
Trace Model
pub struct ExecutionTrace {
pub read_graph: DomainAccessGraph,
pub write_graph: DomainAccessGraph,
pub causal_graph: CausalGraph,
pub information_flow: InformationFlowGraph,
pub executor_divergence: DivergenceGraph,
pub temporal_graph: TemporalGraph,
pub perturbation_response: PerturbationResponse,
pub behavior_fingerprint: BehaviorFingerprint,
}
Trace gates:
Median causal edges per execution: ≥ 24
95% of executions causal rank: ≥ 6
Median touched domains per execution: ≥ 4
95% of executions touched domains: ≥ 3
Behavior fingerprint collision rate: < 5%
Largest behavior cluster: < 2% of corpus
Generator Requirements
Generators must reject flat cases.
pub struct GeneratedCase {
pub world: WorldSnapshot,
pub program: RuneProgram,
pub contexts: Vec<ExecutionContext>,
pub contract: SemanticContract,
pub perturbations: Vec<PerturbedCase>,
}
Generated case gates:
estimated_causal_rank ≥ 6
domain_entropy ≥ configured minimum
perturbation_axes ≥ 10
executor_count ≥ 3
future_dependence present within 3 turns
nonzero hidden/observed state divergence
nonuniform domain fingerprints
Semantic Contract
pub struct SemanticContract {
pub min_causal_rank: usize,
pub min_domain_participation: usize,
pub min_future_sensitivity: f64,
pub min_context_divergence: f64,
pub max_compressibility: f64,
}
Default thresholds:
min_causal_rank: 6
min_domain_participation: 4
min_future_sensitivity: 0.50
min_context_divergence: 0.40
max_compressibility: 0.70
A case passes only if measured trace behavior satisfies its contract.
Metamorphic Testing
For every base execution, produce 10 perturbations.
Perturbations are generated from domain surfaces, not a fixed list.
pub trait PerturbationAxis {
fn apply(&self, world: &WorldSnapshot) -> WorldSnapshot;
fn expected_trace_difference(&self) -> TraceDifferenceExpectation;
}
Metamorphic gates:
≥ 90% perturbations alter trace
≥ 75% perturbations alter world delta
≥ 50% perturbations alter state within 3 future turns
≤ 5% perturbations may be observationally neutral without explanation
Collapse Analysis
The framework must attempt compression attacks.
pub trait CollapseAttack {
fn compress(&self, corpus: &BehaviorCorpus) -> CompressedModel;
fn report(&self) -> CollapseReport;
}
Required attack families:
domain removal
domain merging
constant folding
causal edge deletion
state aliasing
latent factor modeling
behavior clustering
surrogate prediction
temporal flattening
observation flattening
executor identity erasure
Collapse gates:
Best 1-factor model predicts < 40%
Best 2-factor model predicts < 55%
Best 4-factor model predicts < 70%
No single domain explains > 30% outcome variance
No pair of domains explains > 55%
Compressed model loses ≥ 35% trace information
If a smaller model predicts above these thresholds, fail CI.
Semantic Mutation
Mutants are generated structurally from the model.
pub trait SemanticMutator {
fn mutate(&self, runtime: RuntimeArtifact) -> RuntimeArtifact;
fn expected_detection_reason(&self) -> DetectionClass;
}
Mutation gates:
500 semantic mutants minimum
0 surviving mutants
Every mutant must fail at least one named acceptance gate
Survivor report blocks merge
A mutant surviving means the tests are invalid, not that the mutant is acceptable.
Reference Runtime
The reference runtime is the executable spec.
pub trait Runtime {
fn resolve(&self, input: ResolutionInput) -> ResolutionResult;
}
Every execution runs:
let expected = reference.resolve(input.clone());
let actual = runtime_under_test.resolve(input);
assert_eq!(canonical(expected), canonical(actual));
Canonical comparison includes:
world delta
trace
faults
replay hash
future-state hash over 3 turns
Replay Corpus
Every failure becomes permanent.
pub struct ReplayCase {
pub world_seed: u64,
pub program_seed: u64,
pub contract_seed: u64,
pub perturbation_seed: u64,
pub expected_trace_hash: Hash,
pub expected_delta_hash: Hash,
pub expected_future_hash: Hash,
}
Replay gates:
10,000 cases minimum
100% deterministic replay
0 hash drift unless migration explicitly updates corpus
Reports Required Per CI Run
Generate machine-readable JSON and human-readable markdown:
domain_participation_report
causal_rank_report
compression_resistance_report
metamorphic_response_report
mutation_survivor_report
runtime_equivalence_report
replay_report
coverage_report
Merge blocked unless all reports pass.
Absolute Rejection Conditions
Reject if:
A smaller model predicts behavior above thresholds.
Any domain is decorative.
Any domain is read-only or write-only across corpus.
Most programs share the same behavior fingerprint.
Most outcomes reduce to one numeric axis.
Reference/runtime differ.
Replay is nondeterministic.
Any semantic mutant survives.
Trace evidence cannot explain causality.
Phase 0/1 Completion
Phase 0/1 is complete only when:
the adversarial framework exists first
the reference runtime passes it
the optimized runtime matches the reference runtime
collapse attacks fail to simplify the universe
mutation tests kill every generated simplification
generated programs produce diverse, causal, replayable behavior
No spell list. No templates. No cosmetic runes. The deliverable is a Rust engine whose tests make a fake universe fail.