Files
magicka-vm/plan.md
T
linus-d 659544f0b2
magicka-merge-gates / advisory-fast (push) Has been skipped
magicka-merge-gates / merge-gates (push) Failing after 38s
update
2026-06-21 19:21:48 -07:00

9.4 KiB
Raw Blame History

ChatGPT3:26 PM Rust Phase 0/1 Implementor Specification Objective Build the testing framework first, then the runtime.

No spell content is accepted until CI proves the system resists collapse into:

single score single resource single effect axis single executor behavior single rune behavior single hidden damage formula decorative world state Compliance Model Anti-Collapse The framework must also resist collapse of the compliance model.

The implementor may not replace a specification obligation with a representative approximation that merely resembles the obligation.

Every acceptance requirement must have all of:

  1. A measured artifact.
  2. A provenance chain from the measured artifact to the run that produced it.
  3. A merge-blocking enforcement point.
  4. A failure condition that blocks acceptance if the artifact or provenance is absent.

No requirement may be satisfied by:

sample summary approximation representative subset default profile proxy metric regenerated artifact local-only executable documentation claim

unless this specification explicitly defines that weaker substitute as acceptable for that requirement.

The compliance path must be:

Specification requirement Mandatory enforcement mechanism Merge blocked if absent

The compliance path must not be:

Specification requirement Representative approximation Evidence of approximation

Examples of forbidden substitutions:

100% reference/runtime comparison may not be replaced by comparison of base executions only. Required artifacts may not be replaced by partial artifacts. Full trace information may not be replaced by a summarized proxy unless that proxy is explicitly named as the acceptance artifact. Persisted expectations may not be replaced by expectations regenerated in the same run. Merge-blocking enforcement may not be replaced by a manually runnable local binary.

If an implementation uses a weaker substitute, the correct result is not partial credit. The correct result is failure of the corresponding acceptance gate.

Hard CI Gates Minimum per full CI run:

Generated worlds: 50,000 Generated rune programs: 250,000 Executions: 1,000,000 Perturbations per execution: 10 Semantic mutants per run: 500 Replay corpus cases: 10,000 minimum Reference/runtime comparison: 100% of executions Fast CI may run 10%, but merge-blocking CI must run full gates.

Repository Layout crates/ world_model/ rune_ir/ trace_model/ generators/ reference_runtime/ runtime_under_test/ collapse_analysis/ semantic_mutation/ replay_corpus/ ci_reports/ Implementation order is mandatory:

  1. world_model
  2. trace_model
  3. generators
  4. collapse_analysis
  5. semantic_mutation
  6. replay_corpus
  7. reference_runtime
  8. runtime_under_test The optimized runtime may not begin before steps 17 pass CI.

World Model pub struct WorldSnapshot { pub id: WorldId, pub turn: u64, pub domains: Vec, pub causal_state: CausalState, pub observation_state: ObservationState, pub execution_state: ExecutionState, pub time_state: TimeState, pub seed: u64, } Minimum domain count:

8 independent world domains Each domain must expose:

pub trait WorldDomain { fn domain_id(&self) -> DomainId; fn read_surface(&self) -> ReadSurface; fn write_surface(&self) -> WriteSurface; fn perturbation_axes(&self) -> Vec; fn fingerprint(&self) -> DomainFingerprint; } Domain acceptance:

Each domain appears in ≥ 35% of traces. Each domain influences execution in ≥ 20% of corpus. Each domain is mutated by execution in ≥ 20% of corpus. Removing any domain reduces corpus behavioral diversity by ≥ 10%. Merging any two domains loses ≥ 8% predictive accuracy. Rune Program Model pub struct RuneProgram { pub id: ProgramId, pub tokens: Vec, pub seed: u64, } No rune stream may be rejected as the primary safety path.

Execution always returns:

pub struct ResolutionResult { pub delta: WorldDelta, pub trace: ExecutionTrace, pub faults: FaultLog, pub replay: ReplayRecord, } Engine crashes, panics, undefined Rust behavior, or unlogged failures fail CI.

Trace Model pub struct ExecutionTrace { pub read_graph: DomainAccessGraph, pub write_graph: DomainAccessGraph, pub causal_graph: CausalGraph, pub information_flow: InformationFlowGraph, pub executor_divergence: DivergenceGraph, pub temporal_graph: TemporalGraph, pub perturbation_response: PerturbationResponse, pub behavior_fingerprint: BehaviorFingerprint, } Trace gates:

Median causal edges per execution: ≥ 24 95% of executions causal rank: ≥ 6 Median touched domains per execution: ≥ 4 95% of executions touched domains: ≥ 3 Behavior fingerprint collision rate: < 5% Largest behavior cluster: < 2% of corpus Generator Requirements Generators must reject flat cases.

pub struct GeneratedCase { pub world: WorldSnapshot, pub program: RuneProgram, pub contexts: Vec, pub contract: SemanticContract, pub perturbations: Vec, } Generated case gates:

estimated_causal_rank ≥ 6 domain_entropy ≥ configured minimum perturbation_axes ≥ 10 executor_count ≥ 3 future_dependence present within 3 turns nonzero hidden/observed state divergence nonuniform domain fingerprints Semantic Contract pub struct SemanticContract { pub min_causal_rank: usize, pub min_domain_participation: usize, pub min_future_sensitivity: f64, pub min_context_divergence: f64, pub max_compressibility: f64, } Default thresholds:

min_causal_rank: 6 min_domain_participation: 4 min_future_sensitivity: 0.50 min_context_divergence: 0.40 max_compressibility: 0.70 A case passes only if measured trace behavior satisfies its contract.

Metamorphic Testing For every base execution, produce 10 perturbations.

Perturbations are generated from domain surfaces, not a fixed list.

pub trait PerturbationAxis { fn apply(&self, world: &WorldSnapshot) -> WorldSnapshot; fn expected_trace_difference(&self) -> TraceDifferenceExpectation; } Metamorphic gates:

≥ 90% perturbations alter trace ≥ 75% perturbations alter world delta ≥ 50% perturbations alter state within 3 future turns ≤ 5% perturbations may be observationally neutral without explanation Collapse Analysis The framework must attempt compression attacks.

pub trait CollapseAttack { fn compress(&self, corpus: &BehaviorCorpus) -> CompressedModel; fn report(&self) -> CollapseReport; } Required attack families:

domain removal domain merging constant folding causal edge deletion state aliasing latent factor modeling behavior clustering surrogate prediction temporal flattening observation flattening executor identity erasure Collapse gates:

Best 1-factor model predicts < 40% Best 2-factor model predicts < 55% Best 4-factor model predicts < 70% No single domain explains > 30% outcome variance No pair of domains explains > 55% Compressed model loses ≥ 35% trace information If a smaller model predicts above these thresholds, fail CI.

Semantic Mutation Mutants are generated structurally from the model.

pub trait SemanticMutator { fn mutate(&self, runtime: RuntimeArtifact) -> RuntimeArtifact; fn expected_detection_reason(&self) -> DetectionClass; } Mutation gates:

500 semantic mutants minimum 0 surviving mutants Every mutant must fail at least one named acceptance gate Survivor report blocks merge A mutant surviving means the tests are invalid, not that the mutant is acceptable.

Reference Runtime The reference runtime is the executable spec.

pub trait Runtime { fn resolve(&self, input: ResolutionInput) -> ResolutionResult; } Every execution runs:

let expected = reference.resolve(input.clone()); let actual = runtime_under_test.resolve(input);

assert_eq!(canonical(expected), canonical(actual)); Canonical comparison includes:

world delta trace faults replay hash future-state hash over 3 turns Replay Corpus Every failure becomes permanent.

pub struct ReplayCase { pub world_seed: u64, pub program_seed: u64, pub contract_seed: u64, pub perturbation_seed: u64, pub expected_trace_hash: Hash, pub expected_delta_hash: Hash, pub expected_future_hash: Hash, } Replay gates:

10,000 cases minimum 100% deterministic replay 0 hash drift unless migration explicitly updates corpus Reports Required Per CI Run Generate machine-readable JSON and human-readable markdown:

domain_participation_report causal_rank_report compression_resistance_report metamorphic_response_report mutation_survivor_report runtime_equivalence_report replay_report coverage_report Merge blocked unless all reports pass.

Absolute Rejection Conditions Reject if:

A smaller model predicts behavior above thresholds. Any domain is decorative. Any domain is read-only or write-only across corpus. Most programs share the same behavior fingerprint. Most outcomes reduce to one numeric axis. Reference/runtime differ. Replay is nondeterministic. Any semantic mutant survives. Trace evidence cannot explain causality. Phase 0/1 Completion Phase 0/1 is complete only when:

the adversarial framework exists first the reference runtime passes it the optimized runtime matches the reference runtime collapse attacks fail to simplify the universe mutation tests kill every generated simplification generated programs produce diverse, causal, replayable behavior No spell list. No templates. No cosmetic runes. The deliverable is a Rust engine whose tests make a fake universe fail.