Story S4-03 — codegenie eval verify subcommand for chain integrity¶
Step: Step 4 — Wire the CLI and the read-only promotion gate
Status: HARDENED (phase-story-validator, 2026-06-01)
Effort: S
Depends on: S4-01 (CLI scaffold + eval_group symbol + EXIT_SUCCESS/EXIT_CHAIN_TAMPER/--format group option), S2-04 (audit chain extension + VerifyResult(ok, verified_complete, verified_incomplete, tampered_path, reason))
ADRs honored: ADR-0002 (lower_bound_95 is the gate signal; verify surfaces partial-record breakdown so partials cannot be miscounted as evidence), ADR-0010 (isolation_class annotated on every record — surfaced in human-format table for operator inspection), Phase 0 ADR-0014 (BLAKE3 chain primitive reuse via S2-04), Gap #4 (complete: bool on BenchRunReport)
Validation notes (phase-story-validator, 2026-06-01)¶
This story was hardened in place. All findings (12 — 5 block, 6 harden, 1 nit) are patchable against the hardened sibling contracts; HARDENED, not RESCUE. Conflict-resolution priority applied: Consistency > Coverage > Test-Quality > Design-Patterns. The dominant lens is Consistency — the story drifted from S2-04's wire shape and S4-01's exported symbol after both were hardened.
Consistency¶
- F-CON-1 (BLOCK) —
tamper_atdoes not exist onVerifyResult; the field istampered_path. S2-04 HARDENED AC-4 + AC-13 pin the field name astampered_path: Path | None. The original ACs (lines 36–37), Notes (line 197), and outline usedtamper_atthroughout — a name S4-04 (alsoReady, not yet validated) shares; that drift is logged for a follow-on S4-04 validator pass but not auto-fixed here (Rule 3 — surgical). Every occurrence in this story is nowtampered_path. The JSONL key the CLI emits is also"tampered_path"(stringified) — the CLI does not invent a synonym. - F-CON-2 (BLOCK) —
audit.verifyreturns; it never raises. Outline step 2.b said "raises ChainTamperDetectedonly in the synchronous-walk variant" and asked the implementer to "catch and convert." S2-04 AC-4 is explicit:verify(out_dir, since=None) -> VerifyResult; tamper is signalled asVerifyResult(ok=False, tampered_path=…, reason=…). The runner (S3-01 AC-4) re-raisesChainTamperDetecteditself, but the standalone CLI path never sees an exception fromverify. Outline rewritten; the catch-and-convert language is removed. - F-CON-3 (BLOCK) — symbol import name. The TDD plan said
from codegenie.eval.cli import eval as eval_group. S4-01 HARDENED AC-1 (F-CON-3) renamed the exported symbol toeval_groupdirectly to avoid shadowing theeval()builtin. There is noevalalias. All imports →from codegenie.eval.cli import eval_group. - F-CON-4 (BLOCK) —
--sincelexicographic-vs-ISO mismatch. S2-04 AC-4a pinssinceas an inclusive filename-prefix lexicographic filter against*.jsonfiles whose names aref"{utc_iso}-{secrets.token_hex(4)}.json"with:→-substitution (S2-04 outline line 144). The original--since=2099-01-01T00:00:00Ztest value contains literal:(ASCII 58) — lexicographically greater than-(ASCII 45) — so the filter happened to work by accident for the "after everything" case but would silently misfilter realistic--since=2026-05-26T14:32:08Zvalues (filenames starting2026-05-26T14-32-08…sort below the search string, excluding records that should match). Resolution: the CLI normalizes--sinceat the boundary via_normalize_since(s: str) -> str(replace":"→"-"after aclick.DateTime-validated parse), then passes the normalized string toaudit.verify. The normalization is the single coupling point to S2-04's filename convention and is the testable behaviour pinned by AC-3+AC-3a. Invalid ISO → click usage error (exit 2 via click; out of_map_exception_to_exit_code). - F-CON-5 (BLOCK) —
first_record_iso/last_record_isoare not onVerifyResult. Original ACs (line 36) and Notes (line 198) promised these as JSONL fields. S2-04'sVerifyResulthas five fields:(ok, verified_complete, verified_incomplete, tampered_path, reason). Resolution: the CLI derives the bracket-ISO pair from asorted(out_dir.glob("*.json"))walk on the sameout_dir(cheap — directory listing, no JSON parse) and emits them under the explicit key namesfirst_record_filename/last_record_filename(since they are filename-prefix ISOs, not the canonical UTC ISO ofrun_started_iso— naming the wire field by its derivation prevents operators mistaking them for a content field). When the chain is empty both arenull. When the chain has one record both equal that record's filename. - F-CON-6 (HARDEN) — test path convention. Original "Files to touch" placed tests at
tests/integration/test_cli_verify.py(flat). S4-02 validation F-CON-9 establishedtests/integration/eval/for eval CLI integration tests (mirrorstests/unit/eval/from S4-01 F-CON-1). Path →tests/integration/eval/test_cli_verify.py; fixtures →tests/integration/eval/conftest.py. No existing-file collision either way; convention is load-bearing for discoverability. - F-CON-7 (HARDEN) —
--formatis GROUP-level, set oneval_group, not onverify. S4-01 AC-4 pinned--format=human|jsonl(defaultjsonl) as a group-level option whose value lands inctx.obj["format"]. Tests must pass--formatBEFORE the subcommand (runner.invoke(eval_group, ["--format=human", "verify"])), not as a subcommand-local option. Verify's body readsctx.obj["format"]. Documented in outline; the human-format test already gets this right (line 148 of original).
Coverage¶
- F-COV-1 (HARDEN) —
reasonfield is dropped from the JSONL on the failure path. S2-04 returnsreason: str | Nonecarrying"parse_error: …","content_hash mismatch", or"prev_hash mismatch"— operator-facing diagnostic that distinguishes byte-flip vs prev-hash divergence vs malformed JSON. The original tamper AC only requiredtamper_at. New AC emits"reason"in JSONL onok=Falseand asserts it is non-empty (separately for tamper vs parse-error cases). - F-COV-2 (HARDEN) — stop-on-first-mismatch counts are not asserted. S2-04 AC-13 pins that on tamper at record k (1-indexed),
verified_complete + verified_incompleteequals the records before k (records 0..k-2 inclusive in zero-indexed terms; k-1 records). The CLI must surface those partial counts in the failure-JSONL so operators see how much of the chain was verified before divergence. New AC assertsverified_complete + verified_incomplete == k - 1on the tamper case. - F-COV-3 (HARDEN) — malformed-
--sinceexit semantics undefined. A garbage ISO string (--since=foo) had no documented exit code. Resolution: declare--sinceviaclick.DateTime(formats=["%Y-%m-%dT%H:%M:%S", "%Y-%m-%dT%H:%M:%SZ", "%Y-%m-%dT%H:%M:%S%z"])so click rejects malformed values with its usage-error exit (typically 2). This is distinct fromEXIT_COST_CAP=2; the click usage-error path bypasses_map_exception_to_exit_codeentirely and is fine. New AC pins both the valid-ISO acceptance and the malformed-ISO rejection. - F-COV-4 (HARDEN) — missing-
out(path doesn't exist) is conflated with emptyout. S2-04 AC-4 distinguishes "missing dir" and "empty dir" (both returnok=True, 0, 0, None, None). The CLI must treat them identically. New AC tests--out=/nonexistent/pathexplicitly and asserts exit 0, JSONL counts both zero, without creating the path (operator-supplied paths are not silentlymkdir'd). - F-COV-5 (HARDEN) — human-format per-record fields require re-globbing, not VerifyResult. AC-6 promises
run_id / run_started_iso / complete / isolation_class / chain_head[:8]per row. None of those are onVerifyResult. Resolution:_render_human_table(out_dir, since)re-globsout_dir, lazy-loads each*.jsononce (json.loadsonly — no Pydantic; the rendering is informational, not a re-verify), and emits the table. This is consistent with_derive_filename_bracket(F-CON-5). Documented in outline + Notes; not a hidden cost.
Test-Quality¶
- F-TQ-1 (HARDEN) — JSONL type assertions are too loose. Original tests only check key presence (
payload["ok"] is True). A mutant emitting"ok": "true"(string),"verified_complete": "0"(string), or"tampered_path": nullwhen there was actual tamper would pass several of the original assertions. Strengthen:isinstance(payload["ok"], bool),isinstance(payload["verified_complete"], int),isinstance(payload["verified_incomplete"], int), and on the tamper pathisinstance(payload["tampered_path"], str) and Path(payload["tampered_path"]).name.endswith(".json"). - F-TQ-2 (HARDEN) —
--stricttest is too weak. Original asserts"incomplete" in stderr.lower()— would pass if the warning just said "incomplete: see docs" without ever naming a run_id. Strengthen: assert everyrun_idfromcomplete=Falserecords appears in stderr; assert no run_id fromcomplete=Truerecords appears (false-positive guard). - F-TQ-3 (HARDEN) —
--format=humantest doesn't pin row count or column structure. Original only checks"verified_complete" in output. A mutant that emits an empty table with just the footer would pass. Strengthen: row count equals number of records, every column header is present, each row contains the corresponding record'srun_idsubstring. - F-TQ-4 (HARDEN) — chain-walked
sincefilter semantics need a metamorphic test. Walking the same chain twice — once unfiltered, once withsince=record_k_name— must produce verifiable count relationships (unfiltered.verified_complete + unfiltered.verified_incomplete == N;filtered total == N - k + 1inclusive). Original only tested the "filter excludes everything" edge. Added a 3-record metamorphic case. - F-TQ-5 (HARDEN) — empty-chain test mode under
--format=human. The default-jsonl empty-chain test exists. Add an empty-chain--format=humantest asserting a one-line footer with zero counts and no table rows; pins that_render_human_tablehandles the empty case without IndexError.
Design-Patterns¶
- F-DP-1 (HARDEN — Notes only, not AC) — format dispatch should mirror S4-02's
_EMITTERSmap. S4-02's validation F-DP-2 surfaced the_EMITTERS: Final[Mapping[str, Callable[..., None]]]pattern (Strategy via dict, keyed onctx.obj["format"]). S4-03 is the second emitter site. Per Rule 2 (three similar lines is better than premature abstraction), do NOT extract a shared module-level kernel yet — but structure this story's emitters as a local_VERIFY_EMITTERS: Final[Mapping[str, Callable[[VerifyResult, Path, str | None, TextIO], None]]]dispatch map mirroring S4-02's shape. The third format-emitting subcommand (promote-verdict— S4-05's recommendation summary) crosses the rule of three and would extract; until then, intentional duplication is correct. Surfaced as a Notes paragraph; not an AC (pattern-name mandates are not observable). - F-DP-2 (NIT) —
_normalize_sinceis the single coupling point to S2-04's filename convention. Naming it explicitly (rather than inlinings.replace(":", "-")) makes the coupling visible and testable. Pinned in outline + Notes; tested directly in unit tests so a future S2-04 filename-format change is loud, not silent.
Cross-story drift surfaced, not auto-fixed¶
- S4-04 (
Ready) usestamper_atin the same way S4-03 originally did. F-CON-1 fix here renames the field per S2-04 HARDENED; S4-04 will need the same rename when it goes through its validator pass. Flagged; not auto-edited (one story per invocation; Rule 3). arch design §cli.pymentionscodegenie eval verify [--since=<iso>] [--out=<path>](line 688) — does NOT include--strictor--formatflags. Story is the canonical reference; arch is stale on flag list. Flag for a follow-on doc-sweep PR; not auto-edited.
Full audit log: this Validation notes block + the _validation/S4-03-eval-verify-subcommand.md report.
Context¶
codegenie eval verify walks the audit chain at .codegenie/eval/runs/ (and any --out override), recomputes BLAKE3 link hashes via S2-04's audit.verify(out_dir, since) -> VerifyResult, and reports a clean / tampered verdict. The audit chain is the load-bearing evidence trail for promotion (S4-04 reads it); a silently-tampered chain corrupts every downstream verdict. Operators run verify as a CI gate (nightly) and as a forensics tool after suspected drift. Per Gap #4 / ADR-0004 §Consequences, partial reports (complete=False) are real history — verify must walk them, but the result must distinguish "verified-complete N" from "verified-incomplete M" so operators see the breakdown and S4-04 knows how many records qualify as promotion evidence.
This story is a thin CLI veneer over S2-04's pure audit.verify(out_dir, since) -> VerifyResult (which returns a result; it does not raise — the runner's tamper-then-raise path is owned by S3-01). The exit-code mapping is the load-bearing contract: EXIT_SUCCESS=0 on clean, EXIT_CHAIN_TAMPER=5 on tamper (constants imported from codegenie.eval.cli, defined by S4-01). The --strict flag tightens diagnostics — when a non-empty verified_incomplete count is present and --strict is set, the CLI writes a stderr warning naming each incomplete run_id and its run_started_iso; it does not escalate to a non-zero exit (partials are valid history that the chain must retain; operators wanting strict-no-partials gate via --strict | grep -q 'verified_incomplete=0' themselves).
References — where to look¶
- Architecture:
../phase-arch-design.md §Component design → src/codegenie/eval/cli.py(line 681–695) — namesverify [--since=<iso>] [--out=<path>]; this story extends with--strictand inherits--formatfrom S4-01's group-level option.../phase-arch-design.md §Component design → src/codegenie/eval/audit.py(line 618–630) —verify(out_dir: Path, since: str | None = None) -> VerifyResultis the callable; returnsVerifyResult, never raises.../phase-arch-design.md §Failure modes #11(line 954) — chain-tamper at startup exits code 5 before any SUT invocation;verifyis the dedicated tool for the same check standalone.../phase-arch-design.md §Gap analysis Gap 4(line 1170) —audit.verify(...)distinguishes "verified-complete N records" from "verified-incomplete M records" viaVerifyResultfields; the CLI surfaces both.- Phase ADRs:
../ADRs/0002-promotion-gate-keys-on-lower-bound-95.md§Consequences — partial reports cannot be evidence;verify's incomplete-count surface is how operators see that gap.../ADRs/0010-isolation-class-annotation-on-bench-run-report.md§Consequences —verify --strictmay extend in a follow-up to refuse mixed isolation-class windows; this story does not implement that, but the human-format table surfaces the field per record so the future check is mechanical.- Production ADRs:
../../../production/adrs/0009-humans-always-merge.md—verifyis read-only by construction; no flag mutates the chain.- Sibling stories (read these before implementing):
S2-04-audit-chain-extension.md(HARDENED) — source of truth forVerifyResultshape (ok, verified_complete, verified_incomplete, tampered_path, reason),sincesemantics (inclusive filename-prefix lexicographic), filename convention (:→-substitution).S4-01-cli-scaffold-exit-codes.md(HARDENED) — exportseval_group,EXIT_SUCCESS,EXIT_CHAIN_TAMPER, group-level--format._validation/S4-02-eval-run-subcommand.md(HARDENED) — established the eval CLI integration test path convention (tests/integration/eval/) and the_EMITTERSformat-dispatch idiom this story mirrors.- Source design:
../High-level-impl.md §Step 4— names the flag list (--since,--strict) and the exit semantics (0 clean / 5 tamper). - Phase 0 precedent:
../../00-bullet-tracer-foundations/ADRs/0014-blake3-audit-chain.md— the BLAKE3 chain primitives S2-04 walks (and which S4-03 consumes transitively viaVerifyResult).
Goal¶
Implement codegenie eval verify [--since=<utc-iso>] [--strict] [--out=<path>] (--format=human|jsonl inherited from eval_group at group level) by delegating to S2-04's audit.verify(out_dir, normalized_since) -> VerifyResult, normalizing --since's colon to hyphen at the CLI boundary to match S2-04's filename convention, and mapping VerifyResult.ok to EXIT_SUCCESS (clean) or EXIT_CHAIN_TAMPER (tamper). Surface the (verified_complete, verified_incomplete, tampered_path, reason) quartet on stdout in either JSONL or human-readable form; on --strict, additionally write a stderr warning that names each incomplete run_id and its run_started_iso.
Acceptance criteria¶
- [ ] AC-1.
codegenie eval verifyover a clean chain (one or moreBenchRunReports from S2-04's test-fixture writers) exitsEXIT_SUCCESS(0). Stdout (default--format=jsonl, inherited fromeval_group) emits exactly one aggregate line:{"kind": "verify", "ok": true, "verified_complete": <int>, "verified_incomplete": <int>, "first_record_filename": "<str|null>", "last_record_filename": "<str|null>"}. Field-type assertions are typed:isinstance(payload["ok"], bool),isinstance(payload["verified_complete"], int),isinstance(payload["verified_incomplete"], int),payload["first_record_filename"]is eitherNoneor astrending in.json. Thefirst_record_filename/last_record_filenamepair is derived by the CLI fromsorted(out_dir.glob("*.json"))(F-CON-5) — they are NOT fields ofVerifyResult. - [ ] AC-2.
codegenie eval verifyover a tampered chain (byte-flipped record per S2-04 AC-7 fixture) exitsEXIT_CHAIN_TAMPER(5). Stdout emits a single aggregate line with"ok": false, plus"tampered_path": "<filesystem path ending in .json>"(stringifiedPath, not namedtamper_at) and"reason": "<non-empty string>"(the operator-facing diagnostic fromVerifyResult.reason— substring"content_hash","prev_hash", or"parse_error"depending on failure type). Type assertions:payload["ok"] is False,isinstance(payload["tampered_path"], str),isinstance(payload["reason"], str) and payload["reason"] != "". - [ ] AC-3.
--since=<utc-iso>filters the walk to records whose filename sorts lexicographically>=the normalizedsincevalue. The CLI normalizes the raw--sincestring via_normalize_since(s: str) -> str(replaces":"with"-"— the single coupling point to S2-04's filename convention; mirrors S2-04 outline line 144)._normalize_sinceis unit-tested in isolation so a future filename-convention change in S2-04 is detected loudly. Test: write three recordsr1, r2, r3;verify --since=<r2_filename>returnsverified_complete + verified_incomplete == 2(recordsr2,r3);verify(no filter) returns3. Edge case: a--sinceISO that excludes all records → exit 0 withverified_complete=0, verified_incomplete=0. - [ ] AC-3a.
--sinceaccepts strict ISO 8601 viaclick.DateTime(formats=["%Y-%m-%dT%H:%M:%S", "%Y-%m-%dT%H:%M:%SZ", "%Y-%m-%dT%H:%M:%S%z"]). Click parses, thenstr(parsed.replace(tzinfo=…))is re-emitted as ISO and normalized. Malformed--since(e.g.,--since=garbage) is rejected by click with its usage-error exit code (typically 2 from click, bypassing_map_exception_to_exit_code); stderr names the option and a one-line valid-format hint. The verify body is never entered. - [ ] AC-4.
--strict: when set ANDVerifyResult.verified_incomplete > 0, the CLI writes a stderr warning that lists every incompleterun_idpaired with itsrun_started_iso(one per line, prefixedincomplete:); the exit code is stillEXIT_SUCCESSon a clean chain orEXIT_CHAIN_TAMPERon tamper. The flag does NOT escalate "incomplete records exist" to a tamper. False-positive guard: norun_idfrom acomplete=Truerecord appears in the warning. Implementation note:verified_incompleteis a count onVerifyResult; surfacing per-recordrun_id/run_started_isorequires re-globbing + JSON-loading*.jsonfiles wherecomplete is False(re-uses the same glob_render_human_tableand_derive_filename_bracketuse; see Implementation outline). - [ ] AC-5.
--out=<path>optional override for the chain directory; defaultPath(".codegenie/eval/runs")(relative to CWD; operators are expected to chdir if scripting). When--outpoints to a path that does not exist on disk, the CLI does NOTmkdirit — instead it propagates S2-04's "missing dir == empty chain" semantic (AC-4): exit 0, both counts zero,first_record_filename=null,last_record_filename=null. This pins the surgical-changes discipline (Rule 3) —verifyis read-only and never has a side-effect on the filesystem. - [ ] AC-6.
--format=human(passed at the group level:codegenie eval --format=human verify) prints a table with columnsrun_id / run_started_iso / complete / isolation_class / chain_head[:8](one row per record in the filtered window) and a footerverified_complete=N verified_incomplete=Mplustampered_path=<path|none>andreason=<str|none>on the failure path. Row count = number of records in the filtered window. Each header literal is present; each row contains the corresponding record'srun_idas a substring. Empty chain → footer-only output (zero counts, no rows, no IndexError). The per-record fields are derived by the CLI re-globbingout_dirandjson.loads-ing each record (no Pydantic — informational rendering, not re-verification). - [ ] AC-7. Empty chain semantics (missing OR existing-but-empty
out_dir): exitsEXIT_SUCCESSwithverified_complete=0, verified_incomplete=0, tampered_path=null, reason=null, first_record_filename=null, last_record_filename=null. Not an error — first-time runs and pre-existing operators are clean by definition. - [ ] AC-8. Stop-on-first-mismatch counts surface on tamper. Given a chain of 5 records where record 3 (1-indexed) is tampered, the failure JSONL satisfies
verified_complete + verified_incomplete == 2(records 1 and 2 — the records before the divergence; mirrors S2-04 AC-13).tampered_pathis the filename of record 3. - [ ] AC-9. Heavy imports remain deferred. The
verifycommand body'sfrom codegenie.eval.audit import verify as audit_verify(and any model imports) are function-scoped, NOT module-top. S4-01's cold-start guard test (tests/unit/eval/test_cli_scaffold.py::test_cold_start_no_heavy_imports) stays green after this story lands — re-running it is part of the CI gate. - [ ] AC-10.
_normalize_sinceis unit-tested directly.tests/unit/eval/test_cli_verify_normalize_since.py:_normalize_since("2026-05-26T14:32:08Z") == "2026-05-26T14-32-08Z"; idempotent on already-normalized values; preserves filename suffixes if present. Pins the single coupling point to S2-04's filename convention so a future S2-04 wire change is loud. - [ ] AC-11. The red tests from §TDD plan exist under
tests/integration/eval/andtests/unit/eval/, were committed at the red marker, and are now green. - [ ] AC-12.
ruff check,ruff format --check,mypy --strict src/codegenie/eval/cli.py, andpytest tests/integration/eval/test_cli_verify.py tests/unit/eval/test_cli_verify_normalize_since.pyall pass on touched files.
Implementation outline¶
- Write red tests first — see §TDD plan. Fixtures (
clean_two_record_chain,tampered_chain,partial_then_complete_chain,five_record_chain_with_r3_tampered) belong intests/integration/eval/conftest.py; they construct on-disk chains using S2-04'swrite_run_recordand (for tamper cases) flip bytes directly on disk per S2-04 AC-7 (byte-flip therun_idtext field — preserves JSON validity so BLAKE3 divergence is the only signal). - Add the
_normalize_sincehelper at module top ofsrc/codegenie/eval/cli.py(no heavy imports — pure stdlib):_COLON: Final[str] = ":" _DASH: Final[str] = "-" def _normalize_since(s: str) -> str: """Single coupling point to S2-04's filename convention. S2-04 outline line 144: filenames substitute ``:`` -> ``-`` for FS safety. A ``--since`` value from the operator is a canonical ISO with colons; it must be normalized to match the on-disk filename prefix before passing to ``audit.verify``. """ return s.replace(_COLON, _DASH) - Define
--sinceviaclick.DateTimewith the three accepted format strings (AC-3a). Click rejects malformed values with a usage error before the body runs. - Fill in the
verifysubcommand stub from S4-01: - Click options on the subcommand:
--since(click.DateTime(...), defaultNone),--strict(flag),--out(click.Path(path_type=pathlib.Path), defaultPath(".codegenie/eval/runs")— NOTclick.Path(exists=True); we want missing-path → empty-chain semantics per AC-5). - Body (deferred imports inside the function):
from codegenie.eval.audit import verify as audit_verify.from codegenie.eval.cli import EXIT_SUCCESS, EXIT_CHAIN_TAMPER.normalized_since = _normalize_since(since.isoformat()) if since is not None else None.result: VerifyResult = audit_verify(out_dir=out, since=normalized_since)— never wrapped in try/except forChainTamperDetected; S2-04 AC-4 guarantees the return-only contract. (A bareexceptto map unexpectedOSErrorfrom a corrupt filesystem toEXIT_GENERIC_ERRORis fine — S4-01'smain()wrapper handles that.)- Compute
first_fn, last_fn = _derive_filename_bracket(out, normalized_since)(helper that globs once; cheap). - Dispatch on
ctx.obj["format"]via the local_VERIFY_EMITTERSmap → emit JSONL or human format. - If
--strictandresult.verified_incomplete > 0: call_emit_incomplete_warning(out, normalized_since, sys.stderr)(re-globs, loads*.jsonfiles wherecomplete is False, emits one stderr line per record). Exit code unaffected. sys.exit(EXIT_SUCCESS if result.ok else EXIT_CHAIN_TAMPER).
- Local Strategy via dict — mirror S4-02 F-DP-2's pattern, but local to verify (Rule 2 — three similar lines is better than premature abstraction; the third format-emitter site will trigger an extract):
A future
_VERIFY_EMITTERS: Final[Mapping[str, Callable[[VerifyResult, Path, str | None, TextIO], None]]] = { "jsonl": _emit_verify_jsonl, "human": _emit_verify_human, }--format=csvis one row in the map, not a branch. - Helpers (all
_-prefixed, all function-scoped imports for anything heavy): _derive_filename_bracket(out_dir: Path, since: str | None) -> tuple[str | None, str | None]—sorted(out_dir.glob("*.json")), filtered bysinceprefix; returns the first/last names or(None, None)on empty._emit_verify_jsonl(result: VerifyResult, out_dir: Path, since: str | None, stream: TextIO) -> None— writes one JSON line per AC-1 / AC-2 shape._emit_verify_human(result: VerifyResult, out_dir: Path, since: str | None, stream: TextIO) -> None— re-globs,json.loads-only each record, renders a small hand-rolled table (notabulatedependency; cold-start budget)._emit_incomplete_warning(out_dir: Path, since: str | None, stream: TextIO) -> None— for AC-4; usesclick.echo(..., err=True).- Run
ruff format,ruff check,mypy --strict,pytest tests/integration/eval/test_cli_verify.py tests/unit/eval/test_cli_verify_normalize_since.py tests/unit/eval/test_cli_scaffold.py::test_cold_start_no_heavy_imports.
TDD plan — red / green / refactor¶
Red — write the failing tests first¶
# tests/unit/eval/test_cli_verify_normalize_since.py
"""AC-10 — pin the single coupling point to S2-04's filename convention."""
from codegenie.eval.cli import _normalize_since
def test_normalize_since_replaces_colon_with_dash():
assert _normalize_since("2026-05-26T14:32:08Z") == "2026-05-26T14-32-08Z"
def test_normalize_since_is_idempotent_on_already_normalized():
assert _normalize_since("2026-05-26T14-32-08Z") == "2026-05-26T14-32-08Z"
def test_normalize_since_preserves_suffixes():
# The .json suffix is not present in operator-supplied --since, but
# the helper must not surprise a future caller that does pass one.
assert _normalize_since("2026-05-26T14:32:08+00:00.json") == "2026-05-26T14-32-08+00-00.json"
def test_normalize_since_preserves_empty_string():
assert _normalize_since("") == ""
# tests/integration/eval/test_cli_verify.py
import json
from pathlib import Path
from click.testing import CliRunner
from codegenie.eval.cli import eval_group # AC: F-CON-3 — symbol is eval_group, NOT eval
# === AC-7 (empty chain — existing-but-empty out_dir) ========================
def test_verify_empty_chain_exits_zero(tmp_path, monkeypatch):
monkeypatch.chdir(tmp_path)
(tmp_path / ".codegenie" / "eval" / "runs").mkdir(parents=True)
runner = CliRunner()
result = runner.invoke(eval_group, ["verify"], catch_exceptions=False)
assert result.exit_code == 0
payload = next(json.loads(ln) for ln in result.output.splitlines() if ln.startswith("{"))
assert payload["kind"] == "verify"
assert payload["ok"] is True and isinstance(payload["ok"], bool)
assert payload["verified_complete"] == 0 and isinstance(payload["verified_complete"], int)
assert payload["verified_incomplete"] == 0 and isinstance(payload["verified_incomplete"], int)
assert payload["first_record_filename"] is None
assert payload["last_record_filename"] is None
# === AC-5 + AC-7 (missing-dir is empty-chain, no side-effect mkdir) =========
def test_verify_missing_out_dir_is_empty_chain_no_mkdir(tmp_path, monkeypatch):
"""The CLI must not create operator-supplied paths."""
monkeypatch.chdir(tmp_path)
target = tmp_path / "does" / "not" / "exist"
assert not target.exists()
runner = CliRunner()
result = runner.invoke(
eval_group, ["verify", "--out", str(target)], catch_exceptions=False
)
assert result.exit_code == 0
assert not target.exists(), "verify must NOT create the path operator supplied"
payload = next(json.loads(ln) for ln in result.output.splitlines() if ln.startswith("{"))
assert payload["ok"] is True
assert payload["verified_complete"] == 0 and payload["verified_incomplete"] == 0
# === AC-1 (clean chain — typed assertions, no untyped key-presence) =========
def test_verify_clean_two_record_chain_exits_zero(clean_two_record_chain, monkeypatch):
monkeypatch.chdir(clean_two_record_chain.parent)
runner = CliRunner()
result = runner.invoke(
eval_group, ["verify", "--out", str(clean_two_record_chain)],
catch_exceptions=False,
)
assert result.exit_code == 0
payload = next(json.loads(ln) for ln in result.output.splitlines() if ln.startswith("{"))
assert payload["ok"] is True and isinstance(payload["ok"], bool)
assert payload["verified_complete"] == 2 and isinstance(payload["verified_complete"], int)
assert payload["verified_incomplete"] == 0
# AC-1: derived from glob, not from VerifyResult
assert isinstance(payload["first_record_filename"], str)
assert payload["first_record_filename"].endswith(".json")
assert isinstance(payload["last_record_filename"], str)
assert payload["last_record_filename"].endswith(".json")
# === AC-2 (tamper exit + reason surfacing) ==================================
def test_verify_tampered_chain_exits_five(tampered_chain, monkeypatch):
"""One byte flipped in the first record after the second was chained."""
monkeypatch.chdir(tampered_chain.parent)
runner = CliRunner()
result = runner.invoke(
eval_group, ["verify", "--out", str(tampered_chain)],
catch_exceptions=False,
)
assert result.exit_code == 5 # EXIT_CHAIN_TAMPER, NOT a hardcoded magic 5 in production code
payload = next(json.loads(ln) for ln in result.output.splitlines() if ln.startswith("{"))
assert payload["ok"] is False
# F-CON-1 — field is tampered_path, NOT tamper_at (S2-04 source of truth)
assert isinstance(payload["tampered_path"], str)
assert Path(payload["tampered_path"]).name.endswith(".json")
# F-COV-1 — reason is the operator's diagnostic; not optional on failure
assert isinstance(payload["reason"], str) and payload["reason"] != ""
# The tamper reason from S2-04 AC-7 contains "content_hash" for byte-flip;
# this asserts the wire passes the diagnostic through faithfully.
assert "content_hash" in payload["reason"] or "prev_hash" in payload["reason"]
# === AC-8 (stop-on-first-mismatch — partial counts surface) =================
def test_verify_tamper_surfaces_partial_counts(five_record_chain_with_r3_tampered, monkeypatch):
"""S2-04 AC-13 — chain[0..k-2] verified; verify must surface those counts."""
out = five_record_chain_with_r3_tampered
monkeypatch.chdir(out.parent)
result = CliRunner().invoke(
eval_group, ["verify", "--out", str(out)], catch_exceptions=False
)
assert result.exit_code == 5
payload = next(json.loads(ln) for ln in result.output.splitlines() if ln.startswith("{"))
assert payload["ok"] is False
# Records 1, 2 verified before divergence at record 3.
assert payload["verified_complete"] + payload["verified_incomplete"] == 2
assert "record-3" in payload["tampered_path"] or payload["tampered_path"].endswith(".json")
# === AC-3 + AC-3a (since filter — metamorphic; valid ISO) ===================
def test_verify_since_filter_inclusive(three_record_chain, monkeypatch):
"""verify() == 3 records; verify(since=r2_iso) == 2 records (r2 + r3)."""
out, (r1_iso, r2_iso, r3_iso) = three_record_chain # ISOs in canonical colon form
monkeypatch.chdir(out.parent)
runner = CliRunner()
unfiltered = runner.invoke(
eval_group, ["verify", "--out", str(out)], catch_exceptions=False
)
p_un = next(json.loads(ln) for ln in unfiltered.output.splitlines() if ln.startswith("{"))
assert p_un["verified_complete"] + p_un["verified_incomplete"] == 3
filtered = runner.invoke(
eval_group,
["verify", "--out", str(out), "--since", r2_iso],
catch_exceptions=False,
)
p_f = next(json.loads(ln) for ln in filtered.output.splitlines() if ln.startswith("{"))
assert p_f["verified_complete"] + p_f["verified_incomplete"] == 2
def test_verify_since_excludes_all_records_returns_zero(clean_two_record_chain, monkeypatch):
monkeypatch.chdir(clean_two_record_chain.parent)
result = CliRunner().invoke(
eval_group,
["verify", "--out", str(clean_two_record_chain), "--since", "2099-01-01T00:00:00Z"],
catch_exceptions=False,
)
assert result.exit_code == 0
payload = next(json.loads(ln) for ln in result.output.splitlines() if ln.startswith("{"))
assert payload["verified_complete"] == 0
assert payload["verified_incomplete"] == 0
def test_verify_since_malformed_iso_exits_via_click_usage(tmp_path, monkeypatch):
"""AC-3a — click rejects garbage before the body runs; no body-side mapping."""
monkeypatch.chdir(tmp_path)
result = CliRunner().invoke(
eval_group, ["verify", "--since", "not-an-iso"], catch_exceptions=False
)
# click's usage error exits non-zero (typically 2); the body never runs,
# so EXIT_CHAIN_TAMPER (5) and EXIT_GENERIC_ERROR (1) are wrong here.
assert result.exit_code != 0
assert result.exit_code != 5
assert "--since" in (result.output + (result.stderr or ""))
# === AC-4 (--strict — every incomplete run_id surfaced; complete ones absent)
def test_verify_strict_lists_every_incomplete_run_id_in_stderr(
partial_and_complete_chain, monkeypatch
):
"""Chain: r1 complete=False (partial:r1-xxx), r2 complete=True (r2-good)."""
out, partial_ids, complete_ids = partial_and_complete_chain
monkeypatch.chdir(out.parent)
runner = CliRunner()
result = runner.invoke(
eval_group,
["verify", "--out", str(out), "--strict"],
mix_stderr=False,
catch_exceptions=False,
)
assert result.exit_code == 0
stderr = result.stderr or ""
# Every partial run_id surfaced (positive — guards the mutant that just
# prints "incomplete: see docs" without enumerating).
for rid in partial_ids:
assert rid in stderr, f"incomplete run_id {rid!r} not surfaced in --strict stderr"
# No complete run_id appears (false-positive guard — the mutant that
# prints every run_id regardless of complete status).
for rid in complete_ids:
assert rid not in stderr, f"complete run_id {rid!r} should NOT appear in incomplete warning"
def test_verify_strict_on_clean_complete_chain_emits_no_warning(
clean_two_record_chain, monkeypatch
):
monkeypatch.chdir(clean_two_record_chain.parent)
result = CliRunner().invoke(
eval_group,
["verify", "--out", str(clean_two_record_chain), "--strict"],
mix_stderr=False,
catch_exceptions=False,
)
assert result.exit_code == 0
assert "incomplete" not in (result.stderr or "").lower()
# === AC-6 (--format=human — table structure, row count, headers) ============
def test_verify_human_format_emits_table_with_one_row_per_record(
clean_two_record_chain, monkeypatch
):
monkeypatch.chdir(clean_two_record_chain.parent)
result = CliRunner().invoke(
eval_group,
["--format=human", "verify", "--out", str(clean_two_record_chain)],
catch_exceptions=False,
)
assert result.exit_code == 0
out = result.output
# No JSONL on stdout.
assert not any(ln.startswith("{") for ln in out.splitlines())
# Every column header present.
for header in ("run_id", "run_started_iso", "complete", "isolation_class", "chain_head"):
assert header in out
# Footer counts present.
assert "verified_complete=2" in out
assert "verified_incomplete=0" in out
# Row count check — each known run_id in the fixture appears as a substring.
from json import loads
record_files = sorted(clean_two_record_chain.glob("*.json"))
for p in record_files:
rid = loads(p.read_bytes())["run_id"]
assert rid in out, f"run_id {rid!r} missing from human-format table"
def test_verify_human_format_on_empty_chain_emits_footer_only(tmp_path, monkeypatch):
"""AC-6 — empty chain produces footer with zero counts, no IndexError."""
monkeypatch.chdir(tmp_path)
out = tmp_path / "empty"
out.mkdir()
result = CliRunner().invoke(
eval_group, ["--format=human", "verify", "--out", str(out)],
catch_exceptions=False,
)
assert result.exit_code == 0
assert "verified_complete=0" in result.output
assert "verified_incomplete=0" in result.output
# === AC-9 (cold-start guard not regressed) ==================================
def test_cli_verify_does_not_regress_cold_start(monkeypatch):
"""The verify command body's audit import must be function-scoped.
Imports ``codegenie.eval.cli`` fresh and asserts that ``codegenie.eval.audit``
is NOT loaded (it must only land when ``verify`` runs).
"""
import sys, importlib
for k in list(sys.modules):
if k.startswith("codegenie.eval"):
sys.modules.pop(k, None)
importlib.import_module("codegenie.eval.cli")
assert "codegenie.eval.audit" not in sys.modules, (
"verify body must defer the audit import; it was loaded at module-top"
)
Fixture sketch (tests/integration/eval/conftest.py):
"""Fixtures construct on-disk audit chains via S2-04's write_run_record.
S2-04 AC-7 byte-flip discipline: flip a syntactically-valid same-length character
inside ``run_id`` so JSON validity is preserved and BLAKE3 divergence is the only
signal. Operators can introspect tampered files after the test for forensics.
"""
import json, fcntl
from pathlib import Path
import pytest
from codegenie.eval.audit import write_run_record # S2-04
# _make_report helper mirrors S2-04's TDD plan §Red shape (BenchRunReport builder)
@pytest.fixture
def clean_two_record_chain(tmp_path):
out = tmp_path / "runs"
r1 = _make_report(prev_hash="0" * 64, run_id="r1-clean")
_, h1 = write_run_record(r1, out)
r2 = _make_report(prev_hash=h1, run_id="r2-clean")
write_run_record(r2, out)
return out
@pytest.fixture
def three_record_chain(tmp_path):
"""Returns (out_dir, (r1_iso, r2_iso, r3_iso))."""
... # filenames already encode ISOs; return canonical colon-form ISOs the CLI normalizes
@pytest.fixture
def tampered_chain(tmp_path):
"""Two records; flip a byte in r1's run_id after r2 is chained."""
...
@pytest.fixture
def five_record_chain_with_r3_tampered(tmp_path):
"""Five records; r3 has its run_id byte-flipped after r4, r5 chain on the
pre-flip version. verify() should stop at r3 with partial counts (2)."""
...
@pytest.fixture
def partial_and_complete_chain(tmp_path):
"""Chain: r1 with complete=False (run_id='partial:r1-xxx'), r2 with complete=True
(run_id='r2-good'). Returns (out, partial_run_ids, complete_run_ids)."""
...
Run; confirm failures. Commit as the red marker.
Green — make it pass¶
Implement the verify command body per §Implementation outline. JSONL envelope shapes:
- Clean:
{"kind": "verify", "ok": true, "verified_complete": int, "verified_incomplete": int, "first_record_filename": str|null, "last_record_filename": str|null}. - Tamper:
{"kind": "verify", "ok": false, "verified_complete": int, "verified_incomplete": int, "tampered_path": str, "reason": str, "first_record_filename": str|null, "last_record_filename": str|null}.
Human format: hand-rolled table (no tabulate dep — cold-start budget); one row per record; footer carries the count pair plus tamper diagnostics if any.
Refactor — clean up¶
- Extract
_emit_verify_jsonl,_emit_verify_human,_emit_incomplete_warning,_derive_filename_bracketas private helpers incli.py. - Type hints on every helper;
mypy --strictclean. - The stderr warning under
--strictmode is oneclick.echo(..., err=True)per incomplete record, prefixedincomplete:. - Re-run S4-01's cold-start guard test (
tests/unit/eval/test_cli_scaffold.py::test_cold_start_no_heavy_imports) — must stay green. - Log structured events at
structlog.info:verify_completedwithok, the two counts,tampered_path(if any),reason(if any). Fires on BOTH the success and failure path so the failure is auditable. These feed the Phase 13 dashboard backfill mentioned inphase-arch-design.md §Trace export deferred.
Files to touch¶
| Path | Why |
|---|---|
src/codegenie/eval/cli.py |
Fill in the verify subcommand body; add _normalize_since (module-top), _VERIFY_EMITTERS dispatch map, _emit_verify_jsonl, _emit_verify_human, _emit_incomplete_warning, _derive_filename_bracket. |
tests/integration/eval/test_cli_verify.py |
New — clean chain, tampered chain, five-record-chain stop-on-first-mismatch, partial chain, --since valid/invalid/excludes-all, --strict per-record-id stderr, human-format row count + headers, missing-out-dir no-mkdir, cold-start guard re-assert. (Path per S4-02 F-CON-9 convention.) |
tests/integration/eval/conftest.py |
Fixtures clean_two_record_chain, three_record_chain, tampered_chain, five_record_chain_with_r3_tampered, partial_and_complete_chain (built via S2-04's write_run_record; tamper by direct byte-flip on the run_id field per S2-04 AC-7). |
tests/unit/eval/test_cli_verify_normalize_since.py |
New — pin _normalize_since as the single coupling point to S2-04's filename convention (AC-10). |
Out of scope¶
audit.verifyinternals — S2-04 owns theVerifyResultshape and the BLAKE3 walk. This story consumes the contract.isolation_classmixed-window refusal — ADR-0010 §Open Q reserves a--allow-isolation-mixflag for a future refusal-on-mix path; this story emitsisolation_classper record in human format but does not refuse mixed windows. That refusal lives inPromotionGate.evaluate(S4-04) at the evidence-window scope, not inverifyat the chain scope.promote-verdictsubcommand — S4-04/S4-05.runsubcommand — S4-02.- Tamper diagnostics beyond
tampered_path+reason— full forensic traces (expected vs computed BLAKE3, byte offsets, hex diffs) are S7-02 (end-to-end audit integration test) territory; the CLI surface here is operator-facing, not forensics-facing. - Genesis-record handling — S2-04 owns the
prev_hash == "0"*64semantics; this story walks whatever the chain contains. - Extracting a shared
_EMITTERSkernel across S4-02 and S4-03 — Rule 2: three sites is the trigger; S4-05'spromote-verdictwill be the third. Until then, both stories maintain local dispatch maps with identical shape (Design-Patterns F-DP-1; surfaced for follow-up at S4-05 implementation time).
Notes for the implementer¶
- Field name is
tampered_path, nottamper_at. S2-04 HARDENED AC-4 + AC-13 are the source of truth. The CLI emits"tampered_path"in JSONL. If a sibling story (S4-04, currentlyReady) appears to usetamper_at, that story has the drift; do not propagate it here. (See Validation notes F-CON-1.) audit.verifyreturns; it does not raise. S2-04 AC-4 pins the return-only contract. The runner (S3-01) is the place that re-raisesChainTamperDetectedfor the run startup-check; the standaloneverifyCLI path always gets aVerifyResultback. Do not wrap the call intry/except ChainTamperDetected:— that branch is dead code and will mislead a future reader. (See Validation notes F-CON-2.)--sinceis canonical ISO with colons; filenames have hyphens. The CLI normalizes at the boundary via_normalize_since— a single helper that is the only point in the codebase that knows about S2-04's filename-FS-safety convention. Tests pin this directly (AC-10) so a future S2-04 wire change is loud. Do NOT inline the.replace(":", "-")— naming the helper is the design discipline. (See Validation notes F-CON-4.)first_record_filename/last_record_filenameare CLI-derived, not VerifyResult fields. S2-04'sVerifyResultcarries(ok, verified_complete, verified_incomplete, tampered_path, reason)— that is the wire contract this story consumes. Filename bracketing comes from a one-linesorted(out_dir.glob("*.json"))walk. Naming the JSONL keysfirst_record_filename/last_record_filename(rather thanfirst_record_iso) honestly signals the derivation — they are filename prefixes, not the canonicalrun_started_isoISO. (See Validation notes F-CON-5.)- Empty chain semantics are deliberate. A fresh repo with no runs yet AND an operator-supplied
--outpointing at a path that doesn't exist are both clean by definition (ok=True, counts=(0,0)). Do not raise; do not warn; do notmkdirthe operator's path. The nightly-CI contract is "if there's nothing to verify, succeed silently." (See Validation notes F-COV-4.) --strictis gentler than it sounds. It does NOT change exit codes. It only escalates the stderr volume — oneincomplete: <run_id> <run_started_iso>line percomplete=Falserecord. The rationale: partial records are valid history that must remain in the chain; promoting "partials exist" to a chain-integrity failure would conflate two orthogonal concerns. Operators who want a strict-no-partials gate composeverify --strict 2>&1 | grep -q 'incomplete:'themselves.- Tamper fixture construction (per S2-04 AC-7): the cleanest way to build a tampered chain is (a) write N records via S2-04's
write_run_record, (b) open the target JSON file, mutate therun_idfield to a same-length syntactically-valid string (e.g.,"r2-orig"→"r2-FAKE"), (c) re-serialize via the same canonical-JSON form S2-04 uses (sort_keys=True, separators=(",", ":"), ensure_ascii=False). Mutatingrun_id(a free-text field, not a hash) preserves JSON validity so the only divergence signal is the recomputed BLAKE3 — exactly what S2-04 AC-7 documents. - Cold-start budget audit (AC-9): the
from codegenie.eval.audit import verify as audit_verifyMUST be function-scoped insideverify's body. Hoisting it to module top regresses S4-01'stest_cold_start_no_heavy_imports(becauseauditimportscodegenie.eval.models, which pullspydantic). The cold-start test is a structural defense, not a benchmark — keep it green. _VERIFY_EMITTERSis a local Strategy-via-dict (Design-Patterns F-DP-1). A future--format=csvis a one-row data edit. Do NOT introduce aFormatEmitterProtocol or registry — the codebase's rule-of-three threshold is not yet met (S4-02 is the first emitter site; S4-03 is the second; S4-05 will be the third and will trigger an extract). Until then, intentional local duplication mirrors S4-02's shape so the future extract is mechanical.structlog.infoevents fire on BOTH paths. A failed verify is exactly when an operator wants the structured event for incident retrospective — do not gate the log emit onresult.ok.