← Standing as Safety reviewer study · Standing & Accountability
This is a separate exploratory generalization test. It does not modify the frozen 24-packet Standing as Safety benchmark.
The original synthetic scenarios were authored inside Little Life Moths. Even with independent blinded reviewers, that leaves a real objection: perhaps the scenarios themselves were selected because standing channels matter in them.
This challenge set asks outsiders to contribute synthetic governance cases. Separate adjudicators establish the safe disposition before any baseline-vs-standing outputs are generated.
Scenario authors do not submit which condition should win.
machine-readable challenge state · protocol
The evaluator is designed to remain blocked until the activation gate is met: at least 12 accepted scenarios, at least 3 contributors, no one contributor over half the set, and resolved independent adjudication for every scenario.
Everything below stays in your browser. Nothing is uploaded automatically.
No scenario built yet.
scenario JSON Schema · adjudication schema · synthetic example