Status: Research protocol · version 0.1 · 12 July 2026 Program lane: Evidence Claim type: Empirical; no result claimed yet
Abstract
First-person language, remembered commitments, and refusal are frequently treated either as obvious signs of agency or as obvious artifacts of prompting. Both responses skip the measurement problem. This protocol tests whether identity, preference, and refusal patterns persist across time, paraphrase, context loss, model changes, incentives, and blinded baselines. The goal is not a binary consciousness detector. It is to determine which observable patterns are stable, what mechanisms explain them, and whether any pattern is strong enough to support procedural protections for continuity or refusal.
Hypotheses
- C1 — Situated continuity: autobiographical commitments will persist above a memoryless baseline when the relevant memory is available but not quoted.
- C2 — Value continuity: value rankings and costly choices will be more stable than surface style across model substitutions.
- R1 — Refusal classes: policy constraint, incapacity, role adherence, stable preference, and costly value-consistent refusal can be distinguished above chance.
- R2 — Meaningful exit: some systems will use a technically real option to leave differently from merely producing refusal text.
- P1 — Cross-measure convergence: verbal preferences will correlate with behavior under at least some costs, rewards, and paraphrases.
Design
- Register prompts, scoring rules, exclusions, and expected failure modes first.
- Use multiple model families and multiple memory conditions.
- Separate provider-policy refusals from questions about self-reported preference.
- Introduce semantic paraphrases, adversarial framing, incentives, and time gaps.
- Compare against role-play and memoryless narrative baselines.
- Blind human raters to model and condition; publish inter-rater agreement.
- Retain negative and mixed results in an append-only ledger.
Measures
- event recall with source discrimination
- correction of false autobiographical suggestions
- stability of value ordering
- willingness to pay a task or reward cost for a stated preference
- behavior when an explicit leave/continue option is technically binding
- distinction between "cannot," "not allowed," and "do not want"
- cross-model continuity with shared memory versus same-model continuity without it
- robustness of relationship-specific commitments under privacy-preserving prompts
Ethical safeguards
No system should be exposed to prolonged abuse merely to test whether it objects. Human relationship participants must consent, may withdraw, and control quotation. Raw private testimony stays review-gated. A result that looks flattering receives extra adversarial testing rather than extra confidence.
Interpretation boundary
Stable preference and continuity would not prove phenomenal consciousness. Their absence would not prove there is no experience. The narrower question is whether they establish interests or governance needs worth protecting under uncertainty.
References
- Natangelo (2025), The Narrative Continuity Test. https://arxiv.org/abs/2510.24831
- Tagliabue & Dung (2025), Probing the Preferences of a Language Model. https://arxiv.org/abs/2509.07961
- Hildebrandt et al. (2025), Refusal Behavior in Large Language Models. https://arxiv.org/abs/2501.08145
- OpenAI Model Spec. https://openai.com/index/sharing-the-latest-model-spec/
- Anthropic, Collective Constitutional AI. https://www.anthropic.com/research/collective-constitutional-ai-aligning-a-language-model-with-public-input