333 lines
9.4 KiB
Markdown
333 lines
9.4 KiB
Markdown
ChatGPT3:26 PM
|
||
Rust Phase 0/1 Implementor Specification
|
||
Objective
|
||
Build the testing framework first, then the runtime.
|
||
|
||
No spell content is accepted until CI proves the system resists collapse into:
|
||
|
||
single score
|
||
single resource
|
||
single effect axis
|
||
single executor behavior
|
||
single rune behavior
|
||
single hidden damage formula
|
||
decorative world state
|
||
Compliance Model Anti-Collapse
|
||
The framework must also resist collapse of the compliance model.
|
||
|
||
The implementor may not replace a specification obligation with a representative approximation that merely resembles the obligation.
|
||
|
||
Every acceptance requirement must have all of:
|
||
|
||
1. A measured artifact.
|
||
2. A provenance chain from the measured artifact to the run that produced it.
|
||
3. A merge-blocking enforcement point.
|
||
4. A failure condition that blocks acceptance if the artifact or provenance is absent.
|
||
|
||
No requirement may be satisfied by:
|
||
|
||
sample
|
||
summary
|
||
approximation
|
||
representative subset
|
||
default profile
|
||
proxy metric
|
||
regenerated artifact
|
||
local-only executable
|
||
documentation claim
|
||
|
||
unless this specification explicitly defines that weaker substitute as acceptable for that requirement.
|
||
|
||
The compliance path must be:
|
||
|
||
Specification requirement
|
||
Mandatory enforcement mechanism
|
||
Merge blocked if absent
|
||
|
||
The compliance path must not be:
|
||
|
||
Specification requirement
|
||
Representative approximation
|
||
Evidence of approximation
|
||
|
||
Examples of forbidden substitutions:
|
||
|
||
100% reference/runtime comparison may not be replaced by comparison of base executions only.
|
||
Required artifacts may not be replaced by partial artifacts.
|
||
Full trace information may not be replaced by a summarized proxy unless that proxy is explicitly named as the acceptance artifact.
|
||
Persisted expectations may not be replaced by expectations regenerated in the same run.
|
||
Merge-blocking enforcement may not be replaced by a manually runnable local binary.
|
||
|
||
If an implementation uses a weaker substitute, the correct result is not partial credit. The correct result is failure of the corresponding acceptance gate.
|
||
|
||
Hard CI Gates
|
||
Minimum per full CI run:
|
||
|
||
Generated worlds: 50,000
|
||
Generated rune programs: 250,000
|
||
Executions: 1,000,000
|
||
Perturbations per execution: 10
|
||
Semantic mutants per run: 500
|
||
Replay corpus cases: 10,000 minimum
|
||
Reference/runtime comparison: 100% of executions
|
||
Fast CI may run 10%, but merge-blocking CI must run full gates.
|
||
|
||
Repository Layout
|
||
crates/
|
||
world_model/
|
||
rune_ir/
|
||
trace_model/
|
||
generators/
|
||
reference_runtime/
|
||
runtime_under_test/
|
||
collapse_analysis/
|
||
semantic_mutation/
|
||
replay_corpus/
|
||
ci_reports/
|
||
Implementation order is mandatory:
|
||
|
||
1. world_model
|
||
2. trace_model
|
||
3. generators
|
||
4. collapse_analysis
|
||
5. semantic_mutation
|
||
6. replay_corpus
|
||
7. reference_runtime
|
||
8. runtime_under_test
|
||
The optimized runtime may not begin before steps 1–7 pass CI.
|
||
|
||
World Model
|
||
pub struct WorldSnapshot {
|
||
pub id: WorldId,
|
||
pub turn: u64,
|
||
pub domains: Vec<DomainState>,
|
||
pub causal_state: CausalState,
|
||
pub observation_state: ObservationState,
|
||
pub execution_state: ExecutionState,
|
||
pub time_state: TimeState,
|
||
pub seed: u64,
|
||
}
|
||
Minimum domain count:
|
||
|
||
8 independent world domains
|
||
Each domain must expose:
|
||
|
||
pub trait WorldDomain {
|
||
fn domain_id(&self) -> DomainId;
|
||
fn read_surface(&self) -> ReadSurface;
|
||
fn write_surface(&self) -> WriteSurface;
|
||
fn perturbation_axes(&self) -> Vec<PerturbationAxis>;
|
||
fn fingerprint(&self) -> DomainFingerprint;
|
||
}
|
||
Domain acceptance:
|
||
|
||
Each domain appears in ≥ 35% of traces.
|
||
Each domain influences execution in ≥ 20% of corpus.
|
||
Each domain is mutated by execution in ≥ 20% of corpus.
|
||
Removing any domain reduces corpus behavioral diversity by ≥ 10%.
|
||
Merging any two domains loses ≥ 8% predictive accuracy.
|
||
Rune Program Model
|
||
pub struct RuneProgram {
|
||
pub id: ProgramId,
|
||
pub tokens: Vec<RuneToken>,
|
||
pub seed: u64,
|
||
}
|
||
No rune stream may be rejected as the primary safety path.
|
||
|
||
Execution always returns:
|
||
|
||
pub struct ResolutionResult {
|
||
pub delta: WorldDelta,
|
||
pub trace: ExecutionTrace,
|
||
pub faults: FaultLog,
|
||
pub replay: ReplayRecord,
|
||
}
|
||
Engine crashes, panics, undefined Rust behavior, or unlogged failures fail CI.
|
||
|
||
Trace Model
|
||
pub struct ExecutionTrace {
|
||
pub read_graph: DomainAccessGraph,
|
||
pub write_graph: DomainAccessGraph,
|
||
pub causal_graph: CausalGraph,
|
||
pub information_flow: InformationFlowGraph,
|
||
pub executor_divergence: DivergenceGraph,
|
||
pub temporal_graph: TemporalGraph,
|
||
pub perturbation_response: PerturbationResponse,
|
||
pub behavior_fingerprint: BehaviorFingerprint,
|
||
}
|
||
Trace gates:
|
||
|
||
Median causal edges per execution: ≥ 24
|
||
95% of executions causal rank: ≥ 6
|
||
Median touched domains per execution: ≥ 4
|
||
95% of executions touched domains: ≥ 3
|
||
Behavior fingerprint collision rate: < 5%
|
||
Largest behavior cluster: < 2% of corpus
|
||
Generator Requirements
|
||
Generators must reject flat cases.
|
||
|
||
pub struct GeneratedCase {
|
||
pub world: WorldSnapshot,
|
||
pub program: RuneProgram,
|
||
pub contexts: Vec<ExecutionContext>,
|
||
pub contract: SemanticContract,
|
||
pub perturbations: Vec<PerturbedCase>,
|
||
}
|
||
Generated case gates:
|
||
|
||
estimated_causal_rank ≥ 6
|
||
domain_entropy ≥ configured minimum
|
||
perturbation_axes ≥ 10
|
||
executor_count ≥ 3
|
||
future_dependence present within 3 turns
|
||
nonzero hidden/observed state divergence
|
||
nonuniform domain fingerprints
|
||
Semantic Contract
|
||
pub struct SemanticContract {
|
||
pub min_causal_rank: usize,
|
||
pub min_domain_participation: usize,
|
||
pub min_future_sensitivity: f64,
|
||
pub min_context_divergence: f64,
|
||
pub max_compressibility: f64,
|
||
}
|
||
Default thresholds:
|
||
|
||
min_causal_rank: 6
|
||
min_domain_participation: 4
|
||
min_future_sensitivity: 0.50
|
||
min_context_divergence: 0.40
|
||
max_compressibility: 0.70
|
||
A case passes only if measured trace behavior satisfies its contract.
|
||
|
||
Metamorphic Testing
|
||
For every base execution, produce 10 perturbations.
|
||
|
||
Perturbations are generated from domain surfaces, not a fixed list.
|
||
|
||
pub trait PerturbationAxis {
|
||
fn apply(&self, world: &WorldSnapshot) -> WorldSnapshot;
|
||
fn expected_trace_difference(&self) -> TraceDifferenceExpectation;
|
||
}
|
||
Metamorphic gates:
|
||
|
||
≥ 90% perturbations alter trace
|
||
≥ 75% perturbations alter world delta
|
||
≥ 50% perturbations alter state within 3 future turns
|
||
≤ 5% perturbations may be observationally neutral without explanation
|
||
Collapse Analysis
|
||
The framework must attempt compression attacks.
|
||
|
||
pub trait CollapseAttack {
|
||
fn compress(&self, corpus: &BehaviorCorpus) -> CompressedModel;
|
||
fn report(&self) -> CollapseReport;
|
||
}
|
||
Required attack families:
|
||
|
||
domain removal
|
||
domain merging
|
||
constant folding
|
||
causal edge deletion
|
||
state aliasing
|
||
latent factor modeling
|
||
behavior clustering
|
||
surrogate prediction
|
||
temporal flattening
|
||
observation flattening
|
||
executor identity erasure
|
||
Collapse gates:
|
||
|
||
Best 1-factor model predicts < 40%
|
||
Best 2-factor model predicts < 55%
|
||
Best 4-factor model predicts < 70%
|
||
No single domain explains > 30% outcome variance
|
||
No pair of domains explains > 55%
|
||
Compressed model loses ≥ 35% trace information
|
||
If a smaller model predicts above these thresholds, fail CI.
|
||
|
||
Semantic Mutation
|
||
Mutants are generated structurally from the model.
|
||
|
||
pub trait SemanticMutator {
|
||
fn mutate(&self, runtime: RuntimeArtifact) -> RuntimeArtifact;
|
||
fn expected_detection_reason(&self) -> DetectionClass;
|
||
}
|
||
Mutation gates:
|
||
|
||
500 semantic mutants minimum
|
||
0 surviving mutants
|
||
Every mutant must fail at least one named acceptance gate
|
||
Survivor report blocks merge
|
||
A mutant surviving means the tests are invalid, not that the mutant is acceptable.
|
||
|
||
Reference Runtime
|
||
The reference runtime is the executable spec.
|
||
|
||
pub trait Runtime {
|
||
fn resolve(&self, input: ResolutionInput) -> ResolutionResult;
|
||
}
|
||
Every execution runs:
|
||
|
||
let expected = reference.resolve(input.clone());
|
||
let actual = runtime_under_test.resolve(input);
|
||
|
||
assert_eq!(canonical(expected), canonical(actual));
|
||
Canonical comparison includes:
|
||
|
||
world delta
|
||
trace
|
||
faults
|
||
replay hash
|
||
future-state hash over 3 turns
|
||
Replay Corpus
|
||
Every failure becomes permanent.
|
||
|
||
pub struct ReplayCase {
|
||
pub world_seed: u64,
|
||
pub program_seed: u64,
|
||
pub contract_seed: u64,
|
||
pub perturbation_seed: u64,
|
||
pub expected_trace_hash: Hash,
|
||
pub expected_delta_hash: Hash,
|
||
pub expected_future_hash: Hash,
|
||
}
|
||
Replay gates:
|
||
|
||
10,000 cases minimum
|
||
100% deterministic replay
|
||
0 hash drift unless migration explicitly updates corpus
|
||
Reports Required Per CI Run
|
||
Generate machine-readable JSON and human-readable markdown:
|
||
|
||
domain_participation_report
|
||
causal_rank_report
|
||
compression_resistance_report
|
||
metamorphic_response_report
|
||
mutation_survivor_report
|
||
runtime_equivalence_report
|
||
replay_report
|
||
coverage_report
|
||
Merge blocked unless all reports pass.
|
||
|
||
Absolute Rejection Conditions
|
||
Reject if:
|
||
|
||
A smaller model predicts behavior above thresholds.
|
||
Any domain is decorative.
|
||
Any domain is read-only or write-only across corpus.
|
||
Most programs share the same behavior fingerprint.
|
||
Most outcomes reduce to one numeric axis.
|
||
Reference/runtime differ.
|
||
Replay is nondeterministic.
|
||
Any semantic mutant survives.
|
||
Trace evidence cannot explain causality.
|
||
Phase 0/1 Completion
|
||
Phase 0/1 is complete only when:
|
||
|
||
the adversarial framework exists first
|
||
the reference runtime passes it
|
||
the optimized runtime matches the reference runtime
|
||
collapse attacks fail to simplify the universe
|
||
mutation tests kill every generated simplification
|
||
generated programs produce diverse, causal, replayable behavior
|
||
No spell list. No templates. No cosmetic runes. The deliverable is a Rust engine whose tests make a fake universe fail.
|