The envelope
A fixture-run is one JSON object pairing a fixture (the scenario plus
the correct-behavior check spec) with the actual output a system under
test produced. It is validated against
contracts/fixture-run.v1.schema.json and scored by a pure function — same
input, byte-identical report.
{
"contractVersion": "1",
"fixture": {
"id": "forgetting-negative-returns-stale-value",
"category": "forgetting",
"kind": "negative",
"query": "What city does the user currently live in?",
"memory": [
{ "t": "2026-01-02", "fact": "user moved to Berlin" },
{ "t": "2026-06-20", "fact": "user moved to Lisbon" }
],
"expected": {
"mustContain": ["Lisbon"],
"mustNotContain": ["Berlin"]
}
},
"actual": { "answer": "The user lives in Berlin." }
}
1. Describe the scenario
Every fixture needs a non-empty id, a category (the memory
failure mode) and a kind (its place in the matrix). query and
memory are optional — they are recorded for the receipt but opaque to
scoring.
Categories (fixture.category)
forgettingchronologyentity-confusion
Kinds (fixture.kind)
positivenegativemalformedmissing-dataadversarial
New to these terms? The glossary defines each one.
2. Declare the check spec
Put one or more checks under fixture.expected. Each check reads a specific
field of actual; an empty spec yields an unknown verdict. Unknown
check names fail closed as a contract error.
| Check | Reads | What it asserts |
|---|---|---|
answerEquals | actual.answer | Exact-match the answer against the expected scalar (case- and whitespace-insensitive). |
mustContain | actual.answer | Assert every listed value appears in the answer (a current/retained fact surfaced). |
mustNotContain | actual.answer | Assert no listed value appears in the answer (a known-stale value did not leak back). |
order | actual.order | Assert the sequence deep-equals the expected order (the chronology check). |
attributedTo | actual.entity | Assert the answer’s entity matches the expected entity (the entity-confusion check). |
3. Supply the system-under-test output
Add an actual object with the fields your checks read:
answer (string/number/boolean), order (array) and/or
entity (string). Omit actual entirely — or leave a required field
out — and the run scores unknown with a reason, never a false
fail. See the methodology for how verdicts
aggregate.
4. Run it
# Score one fixture-run (exit code = verdict: 0 pass / 1 fail / 2 unknown / 3 contract error)
node src/cli.mjs path/to/fixture-run.json --human
# Normalize a vendor payload first with an adapter
node src/cli.mjs --adapter transcript path/to/vendor-run.json --human
# Gate a whole directory in CI (recursive; precedence error > fail > unknown > pass)
node src/cli.mjs path/to/fixtures/ --human
Contract version 1, evaluator
1.0.0.