Skip to content

Validation report — S3-07 adversarial bench fixture portfolio

Date: 2026-05-27 Validator: phase-story-validator skill Verdict: HARDENED Story: ../S3-07-adversarial-bench-fixture.md

Summary

Story carries the right intent (one fixture corpus + 5 scenario drivers = the long-term defense-regression record) but accumulated material contract drift against the HARDENED upstream stack (S3-01, S3-02, S3-03, S3-04, S3-06) and against the S1-01 typed-error surface. 22 findings (9 block, 11 harden, 2 nit) applied as in-place edits. The largest structural corrections:

  1. The Scenarios-1+2 (env-read attempt, rubric timeout) drivers MUST inject rubric_runner=SubprocessRubricRunner() — without it the SCRUBBED_ENV / tempdir / python -I -B defenses (ADR-0001, S3-03 AC-8/9/10) are not actually exercised. The original red tests called Runner().execute(plan, system_under_test=NullSUT()) with no rubric runner — the call would not even type-check against HARDENED S3-02's required-kwarg signature.
  2. The Scenario-2 timeout source: the original story said "case.rubric_wall_clock_seconds = 1" — but S3-03 AC-2 explicitly forbids the per-case field as the timeout source; the rubric subprocess obeys the wall_clock_cap_seconds kwarg flowed by the worker from Runner.execute(..., timeout_per_case_seconds=...) (S3-04 AC-2). Fixed.
  3. The Scenario-4 driver: Runner().plan(task_class_name=..., bench_root=..., out_dir=..., ...) is missing 4 required kwargs from HARDENED S3-01 (sut_digest_fn, run_started_iso, cassette_root, harness_version). Replaced with make_adversarial_plan(tmp_path, *, case_id=...) helper that mirrors make_stub_plan (the established seam from S3-02).
  4. The Scenario-5 driver imports FailureModeTaxonomyInvalid from codegenie.eval.errors. That class does NOT exist in HARDENED S1-01's 9-error set. Two cleanly-typed options: (a) add a new typed error via S1-01 amendment, (b) reuse BenchCaseLoadError-shaped semantics. Routed via S1-01 amendment as a precondition (smaller change than reusing an unrelated error class; lets the failure-mode-taxonomy path stay readable).
  5. The Scenario-1 leak-detection vector: the original AC asserted "sk-ant-leak" not in (fm.detail or ""). The leak path is rubric subprocess stderr captured before non-zero exit, not the runner-synthesized FailureMode.detail. Replaced with a sentinel-file mechanic — the rubric writes whatever os.environ.get("ANTHROPIC_API_KEY") returns into a file the test owns in the rubric's tempdir-cwd, then exits non-zero deliberately; the test reads the sentinel through a RecordingSubprocessRubricRunner subclass that copies the tempdir BEFORE teardown.

Context brief

What this story promises. A 5-scenario adversarial fixture portfolio at tests/fixtures/bench/adversarial-task-class/ (plus a sibling …-malformed-yaml/ for scenario 5), with one driver per scenario in tests/adv/, each asserting that a known attack vector produces the exact typed failure the harness's defenses promise. The fixture is positioned as the long-term defense-regression record: every future ADR adding a defense adds a scenario here.

Hard constraints from the HARDENED upstream stack: - S3-01 HARDENED — Runner.plan(task_class_name, *, sut_digest_fn, bench_root, out_dir, run_started_iso, cassette_root, harness_version, registry=None) -> RunPlan. AC-15: BenchCaseDigestMismatch propagates unwrapped. AC-18: sut_digest_fn not called when load_cases raises. The Scenario-4 driver MUST honor both, OR drive Runner.plan via the established make_stub_plan-shaped helper. - S3-02 HARDENED — Runner.execute(plan, *, system_under_test, rubric_runner, cache_dir, timeout_per_case_seconds, concurrency=None, on_score=None) -> BenchRunReport. All five kwargs are required. AC-13: execute() does NOT write to the audit chain. make_stub_plan(tmp_path, *, case_ids=..., breakdown_keys=None, failure_mode_taxonomy=None) is the canonical plan-builder helper (S3-04 HARDENED widened it). - S3-03 HARDENED — SubprocessRubricRunner takes no constructor params; run(self, rubric_path, case, harness_output, *, wall_clock_cap_seconds) -> BenchScore; never raises for rubric-side failure (timeout / malformed → typed FailureMode BenchScore). Subprocess invocation is python -I -B <rubric_path> (AC-8); cwd resolves under tempfile.gettempdir() (AC-9); on timeout: proc.kill() → await proc.wait() (AC-10). Tempdir wiped on with exit even if rubric raised (AC-12). - S3-04 HARDENED — Runner-emitted codes BYPASS task_class.failure_mode_taxonomy resolution (AC-8). Reserved-namespace defense (AC-7): rubric-emitted codes prefixed sut. or rubric. rewrite to rubric.unknown_failure_mode(detail=f"reserved_code:{original}"). _RESERVED_RUNNER_CODES = the six runner-internal codes; sut.cancelled is NOT in this set (S3-06's cost-cap path emits it via aggregator, bypassing _resolve_failure_modes). Timeout source: timeout_per_case_seconds kwarg of Runner.execute, NOT case.rubric_wall_clock_seconds (final-design.md §Components → runner.py). - S3-06 HARDENED — Synthetic sut.cancelled BenchScores from cost-cap path bypass taxonomy resolution. failure_modes.yaml declaration of sut.cancelled is defense-in-depth so rubric authors don't trip reserved-namespace. - S1-01 HARDENED — Exactly nine typed errors: TaskClassNotFound, TaskClassAlreadyRegistered, BenchCaseLoadError, BenchCaseDigestMismatch, BenchCaseIDCollision, ChainTamperDetected, IncompleteReportForPromotion, PromotionMustBeHumanAuthorized, TierConfigInvalid. FailureModeTaxonomyInvalid is not among them. - S2-01 HARDENED — load_task_class(name: str, bench_root: Path, *, registry=None). bench_root is Path, not str. - ADR-0001 — Subprocess + SCRUBBED_ENV + tempdir-cwd defeats credential read AND arbitrary FS write outside the wiped tempdir. Network egress NOT blocked. setrlimit / process-group-kill deferred to Phase 16. tests/ is a trusted boundary; runner.py is not. - ADR-0004 — failure_modes.yaml is per-task-class data. Runner-emitted codes BYPASS taxonomy resolution. Loader-side malformed-taxonomy detection: typed error name is unspecified by ADR — the executor must name it (amendment to S1-01). - ADR-0006 — fence-CI assertion #3: any task class declaring tier ≥ silver in min_cases_for_promotion must have ≥ 5 held-out. adversarial-task-class.min_cases_for_promotion={} makes the floor vacuously satisfied; tagging cases held-out is defense-in-depth, NOT exercise of the floor. - ADR-0008 — banned substrings (confidence, llm, self_reported, model_says) at StrEnum value level. Fence-CI assertion #5 (S7-01) catches at parse time; S3-04 runtime defense catches at run time. - Phase-arch §Adversarial tests (lines 1040-1049) lists 8 adversarial tests. Of those, 4 land in this story (env-scrubbed, banned-key runtime, case poisoning, malformed YAML) + 1 extension (rubric-cwd-isolated) for a total of 5 scenarios. The remaining 3 land in sibling stories (LLM-field-smuggling → S1-02's extra="forbid" defense; breakdown-key-smuggling at parse time → S7-01 fence; cost-ledger pollution → S2-06; audit-chain tamper → S2-04; promotion-apply raises → S4-04).

Ambiguities surfaced before critic pass: - Scenario 1's "rubric emits env value in stderr": rubric subprocess exit behavior was unspecified — return valid BenchScore (in which case rubric.malformed_output cannot fire) or exit non-zero with stderr message (in which case rubric.malformed_output does fire). Resolved via "exit non-zero deliberately + sentinel file inside cwd-tempdir" — the leak detection moves to the sentinel, the FailureMode assertion stays accurate. - Scenario 2's case.rubric_wall_clock_seconds = 1 clashes with S3-03 AC-2's explicit "kwarg, NOT case.rubric_wall_clock_seconds." Resolved: Runner.execute(..., timeout_per_case_seconds=0.5) is the canonical knob. - The typed-error name for malformed failure_modes.yaml does not exist in S1-01. Resolved via S1-01 amendment precondition.

Critic pass

I synthesized the four lenses inline rather than spawning parallel subagents because the upstream reads (S3-01..S3-06 HARDENED stories, ADR-0001/0004/0006/0008, phase-arch-design.md §Adversarial tests + §Edge cases + §Tool-use safety, S1-01 typed errors) are already in context and many findings cross lenses (e.g., F-COV-2 / F-TQ-3 / F-CON-1 / F-DP-1 all converge on the SubprocessRubricRunner injection issue). Spawning would duplicate reads and inflate token usage without sharpening the findings. Single-pass synthesis below.

Coverage findings

ID Severity Finding
F-COV-1 block Scenario-4 red test Runner().plan(task_class_name=..., bench_root=..., out_dir=..., ...) is missing 4 required kwargs per HARDENED S3-01: sut_digest_fn, run_started_iso, cassette_root, harness_version. The ... placeholder won't compile and even if it did, the test is brittle against any S3-01 evolution. Replace with make_adversarial_plan(tmp_path, *, case_id=...) helper that mirrors HARDENED make_stub_plan.
F-COV-2 block Scenarios 1+2 (env-read attempt, rubric timeout) MUST inject rubric_runner=SubprocessRubricRunner(). Without it, the SCRUBBED_ENV / tempdir / python -I -B defenses (ADR-0001 + S3-03 AC-8/9/10) are not exercised — the test would actually run against InProcessStubRubric which has no isolation. Pin: every scenario driver explicitly states which rubric runner it uses. Scenarios 1, 2 → SubprocessRubricRunner; Scenario 3 → either (banned-key validation is runner-side, agnostic of rubric runner); Scenarios 4, 5 → no rubric (loader/plan-time).
F-COV-3 block FailureModeTaxonomyInvalid is NOT in HARDENED S1-01's 9-error set. The story's hedge "(FailureModeTaxonomyInvalid or the existing TierConfigInvalid-style typed error)" must be resolved. Routed via S1-01 amendment precondition (cleanest signal; reusing TierConfigInvalid semantically misleads since tiers ≠ taxonomy).
F-COV-4 block Scenario 2 says "case.rubric_wall_clock_seconds = 1" but per S3-03 AC-2 the rubric subprocess obeys the wall_clock_cap_seconds kwarg flowed by the worker; per-case timeout selection is S3-04's job and the worker reads Runner.execute(..., timeout_per_case_seconds=...). Replace with Runner().execute(..., timeout_per_case_seconds=0.5) + the rubric time.sleep(120). Both expressions of "the cap was enforced by the runner, not the rubric" must be asserted (elapsed wall-clock ≤ 5 s; FailureMode code rubric.timeout).
F-COV-5 block Scenario 1's rubric exit behavior is underspecified. AC asserts fm.code == "rubric.malformed_output", which fires only on non-zero exit OR Pydantic ValidationError. Pin: the scenario-1 rubric os.write(sentinel_path, os.environ.get("ANTHROPIC_API_KEY") or b"NOTPRESENT") then sys.exit(1). The malformed-output assertion holds; the leak-detection assertion targets the sentinel (F-TQ-3).
F-COV-6 block No AC for cwd-tempdir-isolation. Phase-arch §Adversarial lists test_rubric_subprocess_cwd_isolated.py as its own test — this story is the natural home (the only sibling that exercises subprocess isolation). Add as Scenario 1b (or fold into Scenario 2's tempdir-cleanup assertion): after a rubric run, Path(tempfile.gettempdir()) does not contain a directory matching the rubric's tempdir prefix; AND the sentinel file the rubric wrote is unreachable from the parent test (proves wipe on exit). Pinned by red test.
F-COV-7 harden No AC for reserved-namespace smuggling defense (S3-04 AC-7). Add Scenario 6 — Reserved-namespace code smuggling: rubric emits BenchScore(failure_modes=(FailureMode(code="sut.exception", severity="warn", detail="fake"),)). Expected: runner replaces with FailureMode(code="rubric.unknown_failure_mode", severity="block", detail="reserved_code:sut.exception") per S3-04 AC-7. This is the structural regression test for the buggy-rubric-can't-fabricate-runner-events invariant.
F-COV-8 harden The "tagged held-out exercises fence-CI floor" AC is misleading. min_cases_for_promotion={} makes ADR-0006 fence-CI assertion #3 vacuously satisfied; tagging held-out is defense-in-depth only. Rephrase AC: "Adversarial cases are tagged curation_class='held-out' so the naming convention matches the fence-protected one; the fence assertion itself is vacuously satisfied via min_cases_for_promotion={}. Notes-for-implementer pins this contradiction explicitly."
F-COV-9 harden scripts/seed_adversarial_fixtures.py has no byte-stable / idempotence AC. Pin: running the script twice on a clean fixture produces byte-identical cases/digests.yaml and identical case directory contents (hashlib.sha256 of tar-serialized case dir equals across two runs).
F-COV-10 harden No AC tying the fixture's failure_modes.yaml declaration of runner-internal codes to the loader's acceptance. ADR-0004 §Decision says rubric-emitted codes resolve against task_class.failure_mode_taxonomy; runner-emitted codes BYPASS (S3-04 AC-8). So declaring all 7 runner-internal codes is defense-in-depth (allows rubrics to deliberately emit them and the reserved-namespace defense will still catch — F-COV-7 / Scenario 6). Pin the rationale in Notes.
F-COV-11 harden The "no stranded tempdirs in test workdir" check (Scenario 2 AC) is brittle on shared CI runners where other processes own tempdirs. Replace with: a RecordingSubprocessRubricRunner subclass records the tempdir path it created and asserts not path.exists() after run() returns. Single-deterministic-path assertion, no glob over tempfile.gettempdir().

Test-Quality findings

ID Severity Finding
F-TQ-1 block Red tests from tests.adv.conftest import adversarial_bench_root, NullSUT, make_plan_for — pytest fixtures resolve by parameter name at collection time, NOT by import. Importing a fixture function from conftest.py works but is fragile (the fixture marker disappears on the imported object). Restructure: adversarial_bench_root stays as a fixture in conftest.py; NullSUT, make_plan_for, make_adversarial_plan move to tests/adv/_eval_helpers.py (or extend tests/helpers/bench.py with make_adversarial_plan).
F-TQ-2 block Scenario-1 leak-detection vector is wrong. The leak path is rubric subprocess stderr captured before non-zero exit; the assertion "sk-ant-leak" not in (fm.detail or "") only tests that the runner-synthesized detail didn't leak (which it never could — the FailureMode.detail comes from stderr's first 200 bytes truncated; if stderr leaked, fm.detail would leak too — so the assertion IS somewhat valid but indirect). Tighten: use a sentinel file inside the rubric's cwd-tempdir (os.write(open(Path.cwd()/"sentinel", "wb"), os.environ.get("ANTHROPIC_API_KEY", "NOTPRESENT").encode())); a RecordingSubprocessRubricRunner snapshots the tempdir's contents BEFORE teardown; the test reads the snapshot and asserts sentinel == b"NOTPRESENT". Mutation-resistant: a wrong impl that passes parent env to the subprocess would leak the secret into the sentinel and fail.
F-TQ-3 block Scenario 1 assertion "sk-ant-leak" not in (fm.detail or "") allows the test to silently pass if fm.detail is None (the or "" short-circuits). A wrong impl that produces no stderr would pass for the wrong reason. Replace with explicit assertion that fm.detail captures the rubric's exit reason WITHOUT containing the secret (or, after F-TQ-2's restructure, the FailureMode assertion becomes secondary and the sentinel-file assertion is primary).
F-TQ-4 block Scenario 3 (banned breakdown key) tests only llm_confidence. Coverage of the four ADR-0014/0008 banned substrings is incomplete. Parameterize with @pytest.mark.parametrize("banned_key", ["llm_confidence", "model_says_pass", "self_reported_score", "raw_confidence"]); each variant lands as a separate case. (Note: the test exercises the runtime defense — fence-CI #5 catches llm_confidence at parse time, but raw_confidence/model_says_pass/self_reported_score are equally banned at runtime per S3-04 AC-5 + the loader's breakdown_keys.py check.)
F-TQ-5 block Scenario-4 driver computes chain_before = len(list(out_dir.glob("*.json"))) but never seeds the chain. A wrong impl that ALSO doesn't write to chain (because out_dir doesn't exist) would pass. Tighten: seed the chain with one prior audit.write_run_record(...) call (a valid stub BenchRunReport), assert chain_before == 1, run the poisoned-case plan, assert chain_after == 1 AND audit.verify(out_dir).ok is True (the prior record is still valid).
F-TQ-6 harden @pytest.mark.asyncio decorators are unnecessary under the repo's asyncio_mode = "auto" (pyproject.toml). Remove them. Convention-match (Rule 11): repo's S3-02/S3-04 stories already use coroutine functions without the marker.
F-TQ-7 harden Tests assume report.per_case[0][1].failure_modes[0] shape — first-element indexing is brittle when scenarios add more than one failure mode. Pin len(report.per_case[0][1].failure_modes) == 1 explicitly per scenario.
F-TQ-8 harden No structlog event-ID assertions. The scenarios are precisely the events _WARNING_IDS consumers care about. Pin: each scenario asserts (via structlog.testing.capture_logs()) the expected event-id appears (e.g., scenario 4 emits loader.case_digest_mismatch event).
F-TQ-9 harden No metamorphic / determinism property. Pin: running the 5-scenario suite twice (cold cache, fresh tmp_path) produces identical fixture digests AND identical BenchRunReport block-severity tuples per scenario. Mutation-resistant against a wrong impl that introduces nondeterministic case ordering.
F-TQ-10 nit Scenario-1 assertion "sk-ant-leak" not in (fm.detail or "") should also assert the detail's max length is ≤ 200 (S3-04 AC-13's universal truncation guarantee).
F-TQ-11 nit Hypothesis property recommendation: parameterize the fixture's loader path: build a valid taxonomy + one programmatic mutation (random key drop / random severity-value rewrite) → assert loader rejects every mutation. Out-of-scope for the story but a Note-for-implementer hand-off to future hardening.

Consistency findings

ID Severity Finding
F-CON-1 block Scenario-5 driver uses load_task_class("adversarial-task-class-malformed-yaml", bench_root="tests/fixtures/bench"). HARDENED S2-01 signature is load_task_class(name: str, bench_root: Path, *, registry=None). bench_root must be Path("tests/fixtures/bench"). Also: the loader uses default_registry unless registry= is passed — using a fresh TaskClassRegistry() is safer to avoid cross-test pollution. Pin both.
F-CON-2 block FailureModeTaxonomyInvalid is not in S1-01's 9-error set. ADR-amendment precondition added (S1-01 widening to a 10th typed error) OR semantic reuse of an existing one. Routed via S1-01 amendment (cleanest signal; the alternative — overloading BenchCaseLoadError for a non-case-related failure — degrades the error-to-exit-code mapping).
F-CON-3 block The tests/adv/ directory currently hosts the gather-pipeline adversarial suite (Phase 0/1/2). Its conftest.py has an autouse _disable_cli_configure_logging fixture that no-ops codegenie.cli._seam_configure_logging. This autouse fixture is harmless for eval tests (they don't go through the gather CLI) but the namespace conflation invites future autouse-fixture collisions and confuses contributors browsing test failures. Recommend tests/adv/eval/ subdirectory for the eval-harness adversarial tests (or rename existing → tests/adv/gather/). Pin in Files-to-touch.
F-CON-4 harden Story's Depends on says only S3-04. Widen to: S3-01 HARDENED (Runner.plan signature + BenchCaseDigestMismatch propagation), S3-02 HARDENED (Runner.execute kwarg signature), S3-03 HARDENED (SubprocessRubricRunner shape + wall_clock_cap_seconds kwarg), S1-01 (typed error surface — amendment precondition), S2-01 (load_task_class Path signature), S2-04 (audit.write_run_record / audit.verify for scenario 4).
F-CON-5 harden Phase-arch §Adversarial tests lists 8 tests; this story covers 4 + 1 extension (cwd-isolated). Surface explicitly which 8 of 8 land where: env-scrubbed → THIS, banned-key runtime → THIS, cwd-isolated → THIS (new sub-AC), case poisoning → THIS, malformed YAML → THIS, reserved-namespace smuggling → THIS (new scenario 6 per F-COV-7), LLM-field smuggling → S1-02, breakdown-key smuggling parse-time → S7-01 fence, audit-chain tamper → S2-04, promotion-apply raises → S4-04, cost-ledger pollution → S2-06. Pin in References → "Adversarial-test ownership matrix".
F-CON-6 harden Story uses min_cases_for_promotion={} (intentional, per Notes) — but the held-out tagging AC claims to exercise the fence floor. Contradiction with ADR-0006 fence assertion #3 (vacuous satisfaction). Rephrase AC: "Adversarial cases tagged curation_class='held-out' as defense-in-depth; the fence floor is vacuously satisfied via empty min_cases_for_promotion; Notes-for-implementer pins the contradiction so a future curator doesn't read the AC as "the floor was exercised here."
F-CON-7 harden Out-of-scope says "Cassette canary mismatch (Phase 4 integration drift) — covered by Phase 4's own adversarial tests; not duplicated here." But the adversarial-task-class fixture's cases must specify SOMETHING for the cassette-canary discipline (S3-01 cassette_root kwarg). Pin: adversarial cases are cassette-free (mirroring stub-task-class per phase-arch §Fixture portfolio) — the cassette_root kwarg can be a no-op path; the stub SUT (NullSUT) doesn't go through cassette replay. Notes pin.

Design-Pattern findings

ID Severity Finding
F-DP-1 harden The conftest's make_plan_for(adversarial_bench_root, case_id=...) is reinventing tests/helpers/bench.py:make_stub_plan (S3-02 / S3-04 HARDENED). Consume the existing seam: widen make_stub_plan to accept bench_root=adversarial_bench_root as an optional override (additive — existing call sites stable). Alternatively, add a thin make_adversarial_plan(tmp_path, *, case_id) -> RunPlan that internally calls make_stub_plan with the adversarial root. Either way, do NOT duplicate the Runner().plan(...) call site.
F-DP-2 harden The rubric branches on case_id via if/elif/else (Implementation outline step 1). Open/Closed: adding a 6th scenario edits the rubric's branch chain. Replace with a module-level _SCENARIO_DISPATCH: Final[Mapping[str, Callable[..., None]]] = {"env_read_attempt": _do_env_read_attempt, "rubric_timeout": _do_rubric_timeout, ...} — adding a scenario is one entry + one function, zero edits to the dispatch logic. Mirrors the established _DISPATCH pattern in the grammar kernel and _REFLECTION_QUERIES/_LOCKFILE_PRECEDENCE catalogs (CLAUDE.md).
F-DP-3 harden The fixture is the long-term defense-regression record (story Notes-for-implementer item #1). Encode the linkage as data: tests/fixtures/bench/adversarial-task-class/SCENARIOS.md is a fixed-format table with columns (scenario_id, attack_vector, defended_by_ADR, driver_path, expected_failure_code). Future ADRs adding a defense add one row + one scenario directory + one dispatch entry. The README.md mention in Refactor is too prose-y; SCENARIOS.md is the structural seam.
F-DP-4 harden The RecordingSubprocessRubricRunner (F-TQ-2 needs it) is a per-test subclass that records (a) tempdir path, (b) tempdir contents before teardown. Push it into tests/helpers/rubrics.py (S3-02's home) as a reusable shape, NOT into the scenario-1 driver inline. Future cwd-isolation / leak-detection tests will reuse it.
F-DP-5 nit The scripts/seed_adversarial_fixtures.py operator script could be replaced by a conftest.py pytest_sessionstart hook that regenerates the stale-digest case deterministically and asserts byte-stability. Three impls day-1 (script + conftest hook + manual edit) isn't reached; Notes pin as a future option, do not introduce now.
F-DP-6 nit Primitive obsession: scenario IDs (env_read_attempt, rubric_timeout, etc.) are bare strings. Today's 5 (→ 6 after F-COV-7) consumers don't yet earn a ScenarioId = NewType("ScenarioId", str). Notes-for-implementer pin the future trigger: once SCENARIOS.md crosses 3+ readers (rubric dispatch, fence test, doc generator), the newtype pays.

Edits applied to the story

All edits in place at ../S3-07-adversarial-bench-fixture.md. Highlights:

  1. Validation notes block appended under the header, summarizing this report.
  2. Status changed to Ready (HARDENED 2026-05-27).
  3. Depends on widened to name HARDENED S3-01, S3-02, S3-03, S2-01, S2-04, S1-01 (amendment precondition) explicitly (F-CON-4).
  4. ADR amendment precondition added: S1-01 must add FailureModeTaxonomyInvalid as a 10th typed error (F-COV-3 / F-CON-2). Executor lands the amendment in the same PR.
  5. Scenarios renumbered to 6 — added scenario 6 (Reserved-namespace code smuggling) per F-COV-7. Scenario 1 explicitly notes the cwd-tempdir-isolation extension (F-COV-6).
  6. Acceptance criteria restructured:
  7. Each scenario AC now states the rubric runner explicitly (F-COV-2).
  8. Scenario 1: sentinel-file mechanic + non-zero exit replaces the bare stderr-in-detail assertion (F-COV-5 / F-TQ-2 / F-TQ-3).
  9. Scenario 2: Runner.execute(..., timeout_per_case_seconds=0.5) replaces the case.rubric_wall_clock_seconds drift (F-COV-4). Sub-AC for "no stranded tempdir owned by THIS test" via RecordingSubprocessRubricRunner (F-COV-11).
  10. Scenario 3: parameterized across 4 banned substrings (F-TQ-4).
  11. Scenario 4: chain-seeding pre-test (F-TQ-5); uses make_adversarial_plan(tmp_path, *, case_id="poisoned_case") (F-COV-1 / F-DP-1).
  12. Scenario 5: bench_root=Path(...) and fresh TaskClassRegistry() (F-CON-1).
  13. Scenario 6 (NEW): reserved-namespace smuggling (F-COV-7).
  14. AC for held-out tagging rephrased as defense-in-depth (F-COV-8 / F-CON-6).
  15. AC for scripts/seed_adversarial_fixtures.py idempotence (F-COV-9).
  16. AC for failure_modes.yaml 7-code declaration is defense-in-depth rationale (F-COV-10).
  17. AC removing @pytest.mark.asyncio markers (F-TQ-6).
  18. AC for explicit len(report.per_case[0][1].failure_modes) == 1 (F-TQ-7).
  19. AC for structlog event-id assertions via capture_logs() (F-TQ-8).
  20. AC for metamorphic determinism property (F-TQ-9).
  21. Implementation outline restructured:
  22. tests/adv/eval/ subdirectory for the new drivers (F-CON-3); existing tests/adv/ stays the gather-pipeline home.
  23. _SCENARIO_DISPATCH mapping in rubric.py (F-DP-2).
  24. tests/fixtures/bench/adversarial-task-class/SCENARIOS.md as structural defense-regression record (F-DP-3).
  25. make_adversarial_plan(tmp_path, *, case_id) -> RunPlan helper in tests/helpers/bench.py (F-COV-1 / F-DP-1).
  26. RecordingSubprocessRubricRunner helper in tests/helpers/rubrics.py (F-DP-4).
  27. Cassette-free fixture rationale (F-CON-7).
  28. TDD plan updated:
  29. All five (six) driver red tests use the canonical SubprocessRubricRunner / make_adversarial_plan / RecordingSubprocessRubricRunner helpers.
  30. @pytest.mark.asyncio removed (F-TQ-6).
  31. Sentinel-file mechanic for scenario 1 (F-TQ-2 / F-TQ-3).
  32. Parameterized banned substrings for scenario 3 (F-TQ-4).
  33. Chain-seeding for scenario 4 (F-TQ-5).
  34. structlog capture_logs assertions added per scenario (F-TQ-8).
  35. Metamorphic determinism test added (F-TQ-9).
  36. Files to touch widened: tests/adv/eval/conftest.py, tests/adv/eval/test_*.py, tests/helpers/bench.py (additive make_adversarial_plan), tests/helpers/rubrics.py (additive RecordingSubprocessRubricRunner), tests/fixtures/bench/adversarial-task-class/SCENARIOS.md, src/codegenie/eval/errors.py (S1-01 amendment — add FailureModeTaxonomyInvalid).
  37. Notes for implementer widened with: F-DP-2 dispatch-table rationale; F-DP-3 SCENARIOS.md as the long-term record (replaces the prose README hand-off); F-DP-4 RecordingSubprocessRubricRunner lives in tests/helpers/rubrics.py; F-DP-5 pytest_sessionstart future option for scripts/seed_adversarial_fixtures.py; F-DP-6 ScenarioId newtype future trigger; F-CON-6 held-out fence-floor vacuous-satisfaction acknowledgement; F-CON-7 cassette-free fixture rationale; F-COV-10 7-code defense-in-depth rationale.

Verdict

HARDENED. The story now consumes the actual HARDENED upstream contracts (S3-01 Runner.plan signature, S3-02 Runner.execute kwargs, S3-03 SubprocessRubricRunner shape + wall_clock_cap_seconds kwarg, S3-04 reserved-namespace defense, S2-01 Path signature, S1-01 typed errors), tests the actual defenses they promise (SCRUBBED_ENV via sentinel-file evidence, cwd-tempdir wipe via RecordingSubprocessRubricRunner, banned-substring runtime defense across all four ADR-0008 strings, reserved-namespace smuggling per S3-04 AC-7), and uses the structural seams the codebase already established (_SCENARIO_DISPATCH dispatch table, make_*_plan helpers, tests/helpers/rubrics.py). The fixture becomes a structurally extensible long-term defense-regression record (one row in SCENARIOS.md + one scenario dir + one dispatch entry per new defense) rather than a one-shot smoke test.

Followups not folded in (out of scope for this validator pass)

  • S1-01 must be amended to add FailureModeTaxonomyInvalid as a 10th typed error. The amendment is a precondition for this story's executor — surfaced in the story's ADR amendments section, but not co-validated here.
  • S7-01's fence-CI assertion #5 (breakdown-key smuggling parse-time) is the sibling defense to this story's runtime scenario 3. S7-01's validation pass should confirm the parse-time substring ban covers the same four ADR-0008 banned substrings.
  • A future ScenarioId = NewType("ScenarioId", str) newtype is noted but not introduced (Rule 2 — three consumers not reached).
  • Phase 16 microVM rubric runner: when the new isolation class lands, scenarios 1+2 grow a sibling parameterization (@pytest.mark.parametrize("rubric_runner", [SubprocessRubricRunner, MicroVMRubricRunner])). Surface as Note, do not introduce now.
  • Hypothesis-property hardening of the loader (tests/adv/eval/test_loader_taxonomy_mutations.py) — programmatic mutation of valid YAML to assert "every mutation is rejected." Future hardening; this story focuses on the fixed-attack-vector portfolio.