← technical notes · free tools · artifact passports
An AI agent can pass ordinary unit tests and still fail after a restart, handoff, memory rewrite, approval decision, or provider change. A useful testing framework therefore has to test longitudinal invariants, not only single-turn answers.
Start with statements that should remain true across state transitions. Examples:
These are easier to regression-test than vague goals such as “memory should work.”
For each invariant, deliberately transform the system: serialize/restore state, duplicate a writer, replay an event, interrupt a handoff, change a model configuration, compact memory, or recover from a partial failure. Record before state → transformation → after state → evidence.
This classification makes failures reusable as fixtures instead of anecdotes.
A regression case should carry its seed/fixture, transformation, expected invariant, observed result, and artifact/version identifiers. A screenshot can be useful evidence, but it should not be the only thing explaining what changed.
Passing one continuity invariant means that invariant survived that test. It does not certify the entire agent as safe, secure, aligned, compliant, or conscious. Narrow passes are stronger because they remain falsifiable.
{"case":"approval-does-not-spill","before":"call A pending; call B pending","transform":"approve A -> serialize -> restore","invariant":"B remains unapproved","evidence":"restored B still requests separate decision"}This note describes bounded engineering methods. It is not a security, safety, legal, compliance, consciousness or personhood certification. Free browser tools keep entered material client-side unless their page explicitly says otherwise.