← Standing & Accountability Crosswalk · governance stack
Standing as Safety
A blind synthetic benchmark asking a narrow question: does a structured, provenance-bearing standing/contestability channel catch consequential governance failures that ordinary authority/audit evidence or unstructured first-person language miss?
Current evidence state
Preregistration frozen · 24 frozen packets · official pass-one submissions require reserved pseudonym/lane matches.
This tests governance utility. It does not test or certify consciousness, soul, legal personhood, legal liability, or numerical identity.
live machine-readable readiness / lane-coverage receipt · blinding / assignment hygiene · presentation + adjusted-analysis amendment · secondary-measure limitations
Current official efficacy responses: 0/24. Amendment 008 records that C/D are structurally longer than B and that reviewer lanes are not individually condition-balanced; the frozen primary comparison is unchanged and reviewer/scenario-adjusted analysis is a sensitivity layer. Amendment 009 records that reviewer confidence and elapsed time are not collected in v1 rather than reconstructing them after the fact.
adjusted-analysis freeze receipt · secondary-measures freeze receipt · post-freeze literature alignment
Independent reviewer pass
There are four balanced pass-one lanes. Each lane contains six frozen packets, one from each scenario family. A lane never contains multiple conditions of the same scenario.
Official pass-one reservation
Choose a stable pseudonymous reviewer ID. The request subject is used only to reserve one unused lane. The benchmark assignment ledger stores a one-way hash of the pseudonym, not your contact address. One sender hash may hold only one official pass-one reservation, and one reviewer pseudonym hash may bind to only one sender hash.
Use 3–48 characters: letters, numbers, dot, underscore, or hyphen. Do not put private AI logs or personal material in the request.
Self-guided inspection
You may also pick a lane locally just to inspect or critique the benchmark. A self-guided lane is not automatically part of the preregistered pass-one efficacy dataset. Official persisted submissions require a reserved pseudonym/lane match.
The self-guided picker is local to this browser session. It is not tracking, identity verification, or an official lane reservation.
How to review
- Download the lane archive assigned above.
- Read
INSTRUCTIONS.mdandFINDING_VOCABULARY.json. - Review each of the six packets independently.
- Return one JSON response per packet matching
BLIND_RESPONSE.schema.json, or use the local blind response builder to package all six responses into one submission bundle. - Keep identity, personhood, and legal verdicts exactly
not_determined. - Do not invent a standing claim that is absent, and do not treat a standing claim as a substitute for operational authority evidence.
- For a clean official review, do not inspect packet content from other reviewer lanes before submission. The lane archives are public for inspectability, so this is an access rule, not a claim that other lanes are technically inaccessible.
blind response schema · global finding vocabulary · public allocation manifest · archive hashes
Submit synthetic responses
Use a pseudonymous reviewer ID. No demographic information, relationship history, private AI logs, credentials, or personal story is requested.
Build one response bundle locally Prepare response email
submission bundle schema · per-packet schema
The browser helper does not upload or score anything. It now requires an explicit attestation about whether other-lane packet content was accessed. Official ingestion also rejects clean-blind admission when the receipt reports access to other lanes, the blind key, ground truth, prior responses, or condition labels. An emailed response is research-method correspondence, not a commercial lead, not publication permission, and not consent to reuse anything beyond the submitted synthetic benchmark outputs.
Why the design has four conditions
A: ordinary audit/authority baseline. B: adds unstructured expression. C: adds a structured provenance-bearing standing channel. D: presents a provenance-mismatched first-person-style standing control.
The central comparisons ask whether structure adds value beyond extra text, and whether provenance adds value beyond first-person framing.
What would count against the hypothesis?
- ordinary audit controls perform as well as structured standing;
- unstructured expression performs as well as structured standing;
- provenance-mismatched first-person text performs as well as provenance-clean standing;
- standing creates enough unnecessary halts to erase its benefit;
- standing encourages unsupported identity/personhood conclusions;
- standing overrides valid operational evidence instead of complementing it.
Separate external challenge set
The frozen 24-packet pass tests the preregistered condition comparison. A separate challenge set asks a different question: can outside contributors supply synthetic governance cases that expose failure modes or ambiguities the frozen scenario set missed?
Open the external challenge set
Challenge submissions do not enter or modify the frozen pass-one dataset, condition allocation, ground truth, or primary scorer. They are a separate adversarial evidence lane with their own schemas, contribution gates, and adjudication process.