Validation report — S3-07 adversarial bench fixture portfolio¶
Date: 2026-05-27
Validator: phase-story-validator skill
Verdict: HARDENED
Story: ../S3-07-adversarial-bench-fixture.md
Summary¶
Story carries the right intent (one fixture corpus + 5 scenario drivers = the long-term defense-regression record) but accumulated material contract drift against the HARDENED upstream stack (S3-01, S3-02, S3-03, S3-04, S3-06) and against the S1-01 typed-error surface. 22 findings (9 block, 11 harden, 2 nit) applied as in-place edits. The largest structural corrections:
- The Scenarios-1+2 (env-read attempt, rubric timeout) drivers MUST inject
rubric_runner=SubprocessRubricRunner()— without it the SCRUBBED_ENV / tempdir /python -I -Bdefenses (ADR-0001, S3-03 AC-8/9/10) are not actually exercised. The original red tests calledRunner().execute(plan, system_under_test=NullSUT())with no rubric runner — the call would not even type-check against HARDENED S3-02's required-kwarg signature. - The Scenario-2 timeout source: the original story said "
case.rubric_wall_clock_seconds = 1" — but S3-03 AC-2 explicitly forbids the per-case field as the timeout source; the rubric subprocess obeys thewall_clock_cap_secondskwarg flowed by the worker fromRunner.execute(..., timeout_per_case_seconds=...)(S3-04 AC-2). Fixed. - The Scenario-4 driver:
Runner().plan(task_class_name=..., bench_root=..., out_dir=..., ...)is missing 4 required kwargs from HARDENED S3-01 (sut_digest_fn,run_started_iso,cassette_root,harness_version). Replaced withmake_adversarial_plan(tmp_path, *, case_id=...)helper that mirrorsmake_stub_plan(the established seam from S3-02). - The Scenario-5 driver imports
FailureModeTaxonomyInvalidfromcodegenie.eval.errors. That class does NOT exist in HARDENED S1-01's 9-error set. Two cleanly-typed options: (a) add a new typed error via S1-01 amendment, (b) reuseBenchCaseLoadError-shaped semantics. Routed via S1-01 amendment as a precondition (smaller change than reusing an unrelated error class; lets the failure-mode-taxonomy path stay readable). - The Scenario-1 leak-detection vector: the original AC asserted
"sk-ant-leak" not in (fm.detail or ""). The leak path is rubric subprocess stderr captured before non-zero exit, not the runner-synthesized FailureMode.detail. Replaced with a sentinel-file mechanic — the rubric writes whateveros.environ.get("ANTHROPIC_API_KEY")returns into a file the test owns in the rubric's tempdir-cwd, then exits non-zero deliberately; the test reads the sentinel through aRecordingSubprocessRubricRunnersubclass that copies the tempdir BEFORE teardown.
Context brief¶
What this story promises. A 5-scenario adversarial fixture portfolio at tests/fixtures/bench/adversarial-task-class/ (plus a sibling …-malformed-yaml/ for scenario 5), with one driver per scenario in tests/adv/, each asserting that a known attack vector produces the exact typed failure the harness's defenses promise. The fixture is positioned as the long-term defense-regression record: every future ADR adding a defense adds a scenario here.
Hard constraints from the HARDENED upstream stack:
- S3-01 HARDENED — Runner.plan(task_class_name, *, sut_digest_fn, bench_root, out_dir, run_started_iso, cassette_root, harness_version, registry=None) -> RunPlan. AC-15: BenchCaseDigestMismatch propagates unwrapped. AC-18: sut_digest_fn not called when load_cases raises. The Scenario-4 driver MUST honor both, OR drive Runner.plan via the established make_stub_plan-shaped helper.
- S3-02 HARDENED — Runner.execute(plan, *, system_under_test, rubric_runner, cache_dir, timeout_per_case_seconds, concurrency=None, on_score=None) -> BenchRunReport. All five kwargs are required. AC-13: execute() does NOT write to the audit chain. make_stub_plan(tmp_path, *, case_ids=..., breakdown_keys=None, failure_mode_taxonomy=None) is the canonical plan-builder helper (S3-04 HARDENED widened it).
- S3-03 HARDENED — SubprocessRubricRunner takes no constructor params; run(self, rubric_path, case, harness_output, *, wall_clock_cap_seconds) -> BenchScore; never raises for rubric-side failure (timeout / malformed → typed FailureMode BenchScore). Subprocess invocation is python -I -B <rubric_path> (AC-8); cwd resolves under tempfile.gettempdir() (AC-9); on timeout: proc.kill() → await proc.wait() (AC-10). Tempdir wiped on with exit even if rubric raised (AC-12).
- S3-04 HARDENED — Runner-emitted codes BYPASS task_class.failure_mode_taxonomy resolution (AC-8). Reserved-namespace defense (AC-7): rubric-emitted codes prefixed sut. or rubric. rewrite to rubric.unknown_failure_mode(detail=f"reserved_code:{original}"). _RESERVED_RUNNER_CODES = the six runner-internal codes; sut.cancelled is NOT in this set (S3-06's cost-cap path emits it via aggregator, bypassing _resolve_failure_modes). Timeout source: timeout_per_case_seconds kwarg of Runner.execute, NOT case.rubric_wall_clock_seconds (final-design.md §Components → runner.py).
- S3-06 HARDENED — Synthetic sut.cancelled BenchScores from cost-cap path bypass taxonomy resolution. failure_modes.yaml declaration of sut.cancelled is defense-in-depth so rubric authors don't trip reserved-namespace.
- S1-01 HARDENED — Exactly nine typed errors: TaskClassNotFound, TaskClassAlreadyRegistered, BenchCaseLoadError, BenchCaseDigestMismatch, BenchCaseIDCollision, ChainTamperDetected, IncompleteReportForPromotion, PromotionMustBeHumanAuthorized, TierConfigInvalid. FailureModeTaxonomyInvalid is not among them.
- S2-01 HARDENED — load_task_class(name: str, bench_root: Path, *, registry=None). bench_root is Path, not str.
- ADR-0001 — Subprocess + SCRUBBED_ENV + tempdir-cwd defeats credential read AND arbitrary FS write outside the wiped tempdir. Network egress NOT blocked. setrlimit / process-group-kill deferred to Phase 16. tests/ is a trusted boundary; runner.py is not.
- ADR-0004 — failure_modes.yaml is per-task-class data. Runner-emitted codes BYPASS taxonomy resolution. Loader-side malformed-taxonomy detection: typed error name is unspecified by ADR — the executor must name it (amendment to S1-01).
- ADR-0006 — fence-CI assertion #3: any task class declaring tier ≥ silver in min_cases_for_promotion must have ≥ 5 held-out. adversarial-task-class.min_cases_for_promotion={} makes the floor vacuously satisfied; tagging cases held-out is defense-in-depth, NOT exercise of the floor.
- ADR-0008 — banned substrings (confidence, llm, self_reported, model_says) at StrEnum value level. Fence-CI assertion #5 (S7-01) catches at parse time; S3-04 runtime defense catches at run time.
- Phase-arch §Adversarial tests (lines 1040-1049) lists 8 adversarial tests. Of those, 4 land in this story (env-scrubbed, banned-key runtime, case poisoning, malformed YAML) + 1 extension (rubric-cwd-isolated) for a total of 5 scenarios. The remaining 3 land in sibling stories (LLM-field-smuggling → S1-02's extra="forbid" defense; breakdown-key-smuggling at parse time → S7-01 fence; cost-ledger pollution → S2-06; audit-chain tamper → S2-04; promotion-apply raises → S4-04).
Ambiguities surfaced before critic pass:
- Scenario 1's "rubric emits env value in stderr": rubric subprocess exit behavior was unspecified — return valid BenchScore (in which case rubric.malformed_output cannot fire) or exit non-zero with stderr message (in which case rubric.malformed_output does fire). Resolved via "exit non-zero deliberately + sentinel file inside cwd-tempdir" — the leak detection moves to the sentinel, the FailureMode assertion stays accurate.
- Scenario 2's case.rubric_wall_clock_seconds = 1 clashes with S3-03 AC-2's explicit "kwarg, NOT case.rubric_wall_clock_seconds." Resolved: Runner.execute(..., timeout_per_case_seconds=0.5) is the canonical knob.
- The typed-error name for malformed failure_modes.yaml does not exist in S1-01. Resolved via S1-01 amendment precondition.
Critic pass¶
I synthesized the four lenses inline rather than spawning parallel subagents because the upstream reads (S3-01..S3-06 HARDENED stories, ADR-0001/0004/0006/0008, phase-arch-design.md §Adversarial tests + §Edge cases + §Tool-use safety, S1-01 typed errors) are already in context and many findings cross lenses (e.g., F-COV-2 / F-TQ-3 / F-CON-1 / F-DP-1 all converge on the SubprocessRubricRunner injection issue). Spawning would duplicate reads and inflate token usage without sharpening the findings. Single-pass synthesis below.
Coverage findings¶
| ID | Severity | Finding |
|---|---|---|
| F-COV-1 | block | Scenario-4 red test Runner().plan(task_class_name=..., bench_root=..., out_dir=..., ...) is missing 4 required kwargs per HARDENED S3-01: sut_digest_fn, run_started_iso, cassette_root, harness_version. The ... placeholder won't compile and even if it did, the test is brittle against any S3-01 evolution. Replace with make_adversarial_plan(tmp_path, *, case_id=...) helper that mirrors HARDENED make_stub_plan. |
| F-COV-2 | block | Scenarios 1+2 (env-read attempt, rubric timeout) MUST inject rubric_runner=SubprocessRubricRunner(). Without it, the SCRUBBED_ENV / tempdir / python -I -B defenses (ADR-0001 + S3-03 AC-8/9/10) are not exercised — the test would actually run against InProcessStubRubric which has no isolation. Pin: every scenario driver explicitly states which rubric runner it uses. Scenarios 1, 2 → SubprocessRubricRunner; Scenario 3 → either (banned-key validation is runner-side, agnostic of rubric runner); Scenarios 4, 5 → no rubric (loader/plan-time). |
| F-COV-3 | block | FailureModeTaxonomyInvalid is NOT in HARDENED S1-01's 9-error set. The story's hedge "(FailureModeTaxonomyInvalid or the existing TierConfigInvalid-style typed error)" must be resolved. Routed via S1-01 amendment precondition (cleanest signal; reusing TierConfigInvalid semantically misleads since tiers ≠ taxonomy). |
| F-COV-4 | block | Scenario 2 says "case.rubric_wall_clock_seconds = 1" but per S3-03 AC-2 the rubric subprocess obeys the wall_clock_cap_seconds kwarg flowed by the worker; per-case timeout selection is S3-04's job and the worker reads Runner.execute(..., timeout_per_case_seconds=...). Replace with Runner().execute(..., timeout_per_case_seconds=0.5) + the rubric time.sleep(120). Both expressions of "the cap was enforced by the runner, not the rubric" must be asserted (elapsed wall-clock ≤ 5 s; FailureMode code rubric.timeout). |
| F-COV-5 | block | Scenario 1's rubric exit behavior is underspecified. AC asserts fm.code == "rubric.malformed_output", which fires only on non-zero exit OR Pydantic ValidationError. Pin: the scenario-1 rubric os.write(sentinel_path, os.environ.get("ANTHROPIC_API_KEY") or b"NOTPRESENT") then sys.exit(1). The malformed-output assertion holds; the leak-detection assertion targets the sentinel (F-TQ-3). |
| F-COV-6 | block | No AC for cwd-tempdir-isolation. Phase-arch §Adversarial lists test_rubric_subprocess_cwd_isolated.py as its own test — this story is the natural home (the only sibling that exercises subprocess isolation). Add as Scenario 1b (or fold into Scenario 2's tempdir-cleanup assertion): after a rubric run, Path(tempfile.gettempdir()) does not contain a directory matching the rubric's tempdir prefix; AND the sentinel file the rubric wrote is unreachable from the parent test (proves wipe on exit). Pinned by red test. |
| F-COV-7 | harden | No AC for reserved-namespace smuggling defense (S3-04 AC-7). Add Scenario 6 — Reserved-namespace code smuggling: rubric emits BenchScore(failure_modes=(FailureMode(code="sut.exception", severity="warn", detail="fake"),)). Expected: runner replaces with FailureMode(code="rubric.unknown_failure_mode", severity="block", detail="reserved_code:sut.exception") per S3-04 AC-7. This is the structural regression test for the buggy-rubric-can't-fabricate-runner-events invariant. |
| F-COV-8 | harden | The "tagged held-out exercises fence-CI floor" AC is misleading. min_cases_for_promotion={} makes ADR-0006 fence-CI assertion #3 vacuously satisfied; tagging held-out is defense-in-depth only. Rephrase AC: "Adversarial cases are tagged curation_class='held-out' so the naming convention matches the fence-protected one; the fence assertion itself is vacuously satisfied via min_cases_for_promotion={}. Notes-for-implementer pins this contradiction explicitly." |
| F-COV-9 | harden | scripts/seed_adversarial_fixtures.py has no byte-stable / idempotence AC. Pin: running the script twice on a clean fixture produces byte-identical cases/digests.yaml and identical case directory contents (hashlib.sha256 of tar-serialized case dir equals across two runs). |
| F-COV-10 | harden | No AC tying the fixture's failure_modes.yaml declaration of runner-internal codes to the loader's acceptance. ADR-0004 §Decision says rubric-emitted codes resolve against task_class.failure_mode_taxonomy; runner-emitted codes BYPASS (S3-04 AC-8). So declaring all 7 runner-internal codes is defense-in-depth (allows rubrics to deliberately emit them and the reserved-namespace defense will still catch — F-COV-7 / Scenario 6). Pin the rationale in Notes. |
| F-COV-11 | harden | The "no stranded tempdirs in test workdir" check (Scenario 2 AC) is brittle on shared CI runners where other processes own tempdirs. Replace with: a RecordingSubprocessRubricRunner subclass records the tempdir path it created and asserts not path.exists() after run() returns. Single-deterministic-path assertion, no glob over tempfile.gettempdir(). |
Test-Quality findings¶
| ID | Severity | Finding |
|---|---|---|
| F-TQ-1 | block | Red tests from tests.adv.conftest import adversarial_bench_root, NullSUT, make_plan_for — pytest fixtures resolve by parameter name at collection time, NOT by import. Importing a fixture function from conftest.py works but is fragile (the fixture marker disappears on the imported object). Restructure: adversarial_bench_root stays as a fixture in conftest.py; NullSUT, make_plan_for, make_adversarial_plan move to tests/adv/_eval_helpers.py (or extend tests/helpers/bench.py with make_adversarial_plan). |
| F-TQ-2 | block | Scenario-1 leak-detection vector is wrong. The leak path is rubric subprocess stderr captured before non-zero exit; the assertion "sk-ant-leak" not in (fm.detail or "") only tests that the runner-synthesized detail didn't leak (which it never could — the FailureMode.detail comes from stderr's first 200 bytes truncated; if stderr leaked, fm.detail would leak too — so the assertion IS somewhat valid but indirect). Tighten: use a sentinel file inside the rubric's cwd-tempdir (os.write(open(Path.cwd()/"sentinel", "wb"), os.environ.get("ANTHROPIC_API_KEY", "NOTPRESENT").encode())); a RecordingSubprocessRubricRunner snapshots the tempdir's contents BEFORE teardown; the test reads the snapshot and asserts sentinel == b"NOTPRESENT". Mutation-resistant: a wrong impl that passes parent env to the subprocess would leak the secret into the sentinel and fail. |
| F-TQ-3 | block | Scenario 1 assertion "sk-ant-leak" not in (fm.detail or "") allows the test to silently pass if fm.detail is None (the or "" short-circuits). A wrong impl that produces no stderr would pass for the wrong reason. Replace with explicit assertion that fm.detail captures the rubric's exit reason WITHOUT containing the secret (or, after F-TQ-2's restructure, the FailureMode assertion becomes secondary and the sentinel-file assertion is primary). |
| F-TQ-4 | block | Scenario 3 (banned breakdown key) tests only llm_confidence. Coverage of the four ADR-0014/0008 banned substrings is incomplete. Parameterize with @pytest.mark.parametrize("banned_key", ["llm_confidence", "model_says_pass", "self_reported_score", "raw_confidence"]); each variant lands as a separate case. (Note: the test exercises the runtime defense — fence-CI #5 catches llm_confidence at parse time, but raw_confidence/model_says_pass/self_reported_score are equally banned at runtime per S3-04 AC-5 + the loader's breakdown_keys.py check.) |
| F-TQ-5 | block | Scenario-4 driver computes chain_before = len(list(out_dir.glob("*.json"))) but never seeds the chain. A wrong impl that ALSO doesn't write to chain (because out_dir doesn't exist) would pass. Tighten: seed the chain with one prior audit.write_run_record(...) call (a valid stub BenchRunReport), assert chain_before == 1, run the poisoned-case plan, assert chain_after == 1 AND audit.verify(out_dir).ok is True (the prior record is still valid). |
| F-TQ-6 | harden | @pytest.mark.asyncio decorators are unnecessary under the repo's asyncio_mode = "auto" (pyproject.toml). Remove them. Convention-match (Rule 11): repo's S3-02/S3-04 stories already use coroutine functions without the marker. |
| F-TQ-7 | harden | Tests assume report.per_case[0][1].failure_modes[0] shape — first-element indexing is brittle when scenarios add more than one failure mode. Pin len(report.per_case[0][1].failure_modes) == 1 explicitly per scenario. |
| F-TQ-8 | harden | No structlog event-ID assertions. The scenarios are precisely the events _WARNING_IDS consumers care about. Pin: each scenario asserts (via structlog.testing.capture_logs()) the expected event-id appears (e.g., scenario 4 emits loader.case_digest_mismatch event). |
| F-TQ-9 | harden | No metamorphic / determinism property. Pin: running the 5-scenario suite twice (cold cache, fresh tmp_path) produces identical fixture digests AND identical BenchRunReport block-severity tuples per scenario. Mutation-resistant against a wrong impl that introduces nondeterministic case ordering. |
| F-TQ-10 | nit | Scenario-1 assertion "sk-ant-leak" not in (fm.detail or "") should also assert the detail's max length is ≤ 200 (S3-04 AC-13's universal truncation guarantee). |
| F-TQ-11 | nit | Hypothesis property recommendation: parameterize the fixture's loader path: build a valid taxonomy + one programmatic mutation (random key drop / random severity-value rewrite) → assert loader rejects every mutation. Out-of-scope for the story but a Note-for-implementer hand-off to future hardening. |
Consistency findings¶
| ID | Severity | Finding |
|---|---|---|
| F-CON-1 | block | Scenario-5 driver uses load_task_class("adversarial-task-class-malformed-yaml", bench_root="tests/fixtures/bench"). HARDENED S2-01 signature is load_task_class(name: str, bench_root: Path, *, registry=None). bench_root must be Path("tests/fixtures/bench"). Also: the loader uses default_registry unless registry= is passed — using a fresh TaskClassRegistry() is safer to avoid cross-test pollution. Pin both. |
| F-CON-2 | block | FailureModeTaxonomyInvalid is not in S1-01's 9-error set. ADR-amendment precondition added (S1-01 widening to a 10th typed error) OR semantic reuse of an existing one. Routed via S1-01 amendment (cleanest signal; the alternative — overloading BenchCaseLoadError for a non-case-related failure — degrades the error-to-exit-code mapping). |
| F-CON-3 | block | The tests/adv/ directory currently hosts the gather-pipeline adversarial suite (Phase 0/1/2). Its conftest.py has an autouse _disable_cli_configure_logging fixture that no-ops codegenie.cli._seam_configure_logging. This autouse fixture is harmless for eval tests (they don't go through the gather CLI) but the namespace conflation invites future autouse-fixture collisions and confuses contributors browsing test failures. Recommend tests/adv/eval/ subdirectory for the eval-harness adversarial tests (or rename existing → tests/adv/gather/). Pin in Files-to-touch. |
| F-CON-4 | harden | Story's Depends on says only S3-04. Widen to: S3-01 HARDENED (Runner.plan signature + BenchCaseDigestMismatch propagation), S3-02 HARDENED (Runner.execute kwarg signature), S3-03 HARDENED (SubprocessRubricRunner shape + wall_clock_cap_seconds kwarg), S1-01 (typed error surface — amendment precondition), S2-01 (load_task_class Path signature), S2-04 (audit.write_run_record / audit.verify for scenario 4). |
| F-CON-5 | harden | Phase-arch §Adversarial tests lists 8 tests; this story covers 4 + 1 extension (cwd-isolated). Surface explicitly which 8 of 8 land where: env-scrubbed → THIS, banned-key runtime → THIS, cwd-isolated → THIS (new sub-AC), case poisoning → THIS, malformed YAML → THIS, reserved-namespace smuggling → THIS (new scenario 6 per F-COV-7), LLM-field smuggling → S1-02, breakdown-key smuggling parse-time → S7-01 fence, audit-chain tamper → S2-04, promotion-apply raises → S4-04, cost-ledger pollution → S2-06. Pin in References → "Adversarial-test ownership matrix". |
| F-CON-6 | harden | Story uses min_cases_for_promotion={} (intentional, per Notes) — but the held-out tagging AC claims to exercise the fence floor. Contradiction with ADR-0006 fence assertion #3 (vacuous satisfaction). Rephrase AC: "Adversarial cases tagged curation_class='held-out' as defense-in-depth; the fence floor is vacuously satisfied via empty min_cases_for_promotion; Notes-for-implementer pins the contradiction so a future curator doesn't read the AC as "the floor was exercised here." |
| F-CON-7 | harden | Out-of-scope says "Cassette canary mismatch (Phase 4 integration drift) — covered by Phase 4's own adversarial tests; not duplicated here." But the adversarial-task-class fixture's cases must specify SOMETHING for the cassette-canary discipline (S3-01 cassette_root kwarg). Pin: adversarial cases are cassette-free (mirroring stub-task-class per phase-arch §Fixture portfolio) — the cassette_root kwarg can be a no-op path; the stub SUT (NullSUT) doesn't go through cassette replay. Notes pin. |
Design-Pattern findings¶
| ID | Severity | Finding |
|---|---|---|
| F-DP-1 | harden | The conftest's make_plan_for(adversarial_bench_root, case_id=...) is reinventing tests/helpers/bench.py:make_stub_plan (S3-02 / S3-04 HARDENED). Consume the existing seam: widen make_stub_plan to accept bench_root=adversarial_bench_root as an optional override (additive — existing call sites stable). Alternatively, add a thin make_adversarial_plan(tmp_path, *, case_id) -> RunPlan that internally calls make_stub_plan with the adversarial root. Either way, do NOT duplicate the Runner().plan(...) call site. |
| F-DP-2 | harden | The rubric branches on case_id via if/elif/else (Implementation outline step 1). Open/Closed: adding a 6th scenario edits the rubric's branch chain. Replace with a module-level _SCENARIO_DISPATCH: Final[Mapping[str, Callable[..., None]]] = {"env_read_attempt": _do_env_read_attempt, "rubric_timeout": _do_rubric_timeout, ...} — adding a scenario is one entry + one function, zero edits to the dispatch logic. Mirrors the established _DISPATCH pattern in the grammar kernel and _REFLECTION_QUERIES/_LOCKFILE_PRECEDENCE catalogs (CLAUDE.md). |
| F-DP-3 | harden | The fixture is the long-term defense-regression record (story Notes-for-implementer item #1). Encode the linkage as data: tests/fixtures/bench/adversarial-task-class/SCENARIOS.md is a fixed-format table with columns (scenario_id, attack_vector, defended_by_ADR, driver_path, expected_failure_code). Future ADRs adding a defense add one row + one scenario directory + one dispatch entry. The README.md mention in Refactor is too prose-y; SCENARIOS.md is the structural seam. |
| F-DP-4 | harden | The RecordingSubprocessRubricRunner (F-TQ-2 needs it) is a per-test subclass that records (a) tempdir path, (b) tempdir contents before teardown. Push it into tests/helpers/rubrics.py (S3-02's home) as a reusable shape, NOT into the scenario-1 driver inline. Future cwd-isolation / leak-detection tests will reuse it. |
| F-DP-5 | nit | The scripts/seed_adversarial_fixtures.py operator script could be replaced by a conftest.py pytest_sessionstart hook that regenerates the stale-digest case deterministically and asserts byte-stability. Three impls day-1 (script + conftest hook + manual edit) isn't reached; Notes pin as a future option, do not introduce now. |
| F-DP-6 | nit | Primitive obsession: scenario IDs (env_read_attempt, rubric_timeout, etc.) are bare strings. Today's 5 (→ 6 after F-COV-7) consumers don't yet earn a ScenarioId = NewType("ScenarioId", str). Notes-for-implementer pin the future trigger: once SCENARIOS.md crosses 3+ readers (rubric dispatch, fence test, doc generator), the newtype pays. |
Edits applied to the story¶
All edits in place at ../S3-07-adversarial-bench-fixture.md. Highlights:
Validation notesblock appended under the header, summarizing this report.- Status changed to
Ready (HARDENED 2026-05-27). - Depends on widened to name HARDENED S3-01, S3-02, S3-03, S2-01, S2-04, S1-01 (amendment precondition) explicitly (F-CON-4).
- ADR amendment precondition added: S1-01 must add
FailureModeTaxonomyInvalidas a 10th typed error (F-COV-3 / F-CON-2). Executor lands the amendment in the same PR. - Scenarios renumbered to 6 — added scenario 6 (Reserved-namespace code smuggling) per F-COV-7. Scenario 1 explicitly notes the cwd-tempdir-isolation extension (F-COV-6).
- Acceptance criteria restructured:
- Each scenario AC now states the rubric runner explicitly (F-COV-2).
- Scenario 1: sentinel-file mechanic + non-zero exit replaces the bare stderr-in-detail assertion (F-COV-5 / F-TQ-2 / F-TQ-3).
- Scenario 2:
Runner.execute(..., timeout_per_case_seconds=0.5)replaces thecase.rubric_wall_clock_secondsdrift (F-COV-4). Sub-AC for "no stranded tempdir owned by THIS test" viaRecordingSubprocessRubricRunner(F-COV-11). - Scenario 3: parameterized across 4 banned substrings (F-TQ-4).
- Scenario 4: chain-seeding pre-test (F-TQ-5); uses
make_adversarial_plan(tmp_path, *, case_id="poisoned_case")(F-COV-1 / F-DP-1). - Scenario 5:
bench_root=Path(...)and freshTaskClassRegistry()(F-CON-1). - Scenario 6 (NEW): reserved-namespace smuggling (F-COV-7).
- AC for
held-outtagging rephrased as defense-in-depth (F-COV-8 / F-CON-6). - AC for
scripts/seed_adversarial_fixtures.pyidempotence (F-COV-9). - AC for
failure_modes.yaml7-code declaration is defense-in-depth rationale (F-COV-10). - AC removing
@pytest.mark.asynciomarkers (F-TQ-6). - AC for explicit
len(report.per_case[0][1].failure_modes) == 1(F-TQ-7). - AC for structlog event-id assertions via
capture_logs()(F-TQ-8). - AC for metamorphic determinism property (F-TQ-9).
- Implementation outline restructured:
tests/adv/eval/subdirectory for the new drivers (F-CON-3); existingtests/adv/stays the gather-pipeline home._SCENARIO_DISPATCHmapping inrubric.py(F-DP-2).tests/fixtures/bench/adversarial-task-class/SCENARIOS.mdas structural defense-regression record (F-DP-3).make_adversarial_plan(tmp_path, *, case_id) -> RunPlanhelper intests/helpers/bench.py(F-COV-1 / F-DP-1).RecordingSubprocessRubricRunnerhelper intests/helpers/rubrics.py(F-DP-4).- Cassette-free fixture rationale (F-CON-7).
- TDD plan updated:
- All five (six) driver red tests use the canonical
SubprocessRubricRunner/make_adversarial_plan/RecordingSubprocessRubricRunnerhelpers. @pytest.mark.asyncioremoved (F-TQ-6).- Sentinel-file mechanic for scenario 1 (F-TQ-2 / F-TQ-3).
- Parameterized banned substrings for scenario 3 (F-TQ-4).
- Chain-seeding for scenario 4 (F-TQ-5).
- structlog capture_logs assertions added per scenario (F-TQ-8).
- Metamorphic determinism test added (F-TQ-9).
- Files to touch widened:
tests/adv/eval/conftest.py,tests/adv/eval/test_*.py,tests/helpers/bench.py(additivemake_adversarial_plan),tests/helpers/rubrics.py(additiveRecordingSubprocessRubricRunner),tests/fixtures/bench/adversarial-task-class/SCENARIOS.md,src/codegenie/eval/errors.py(S1-01 amendment — addFailureModeTaxonomyInvalid). - Notes for implementer widened with: F-DP-2 dispatch-table rationale; F-DP-3 SCENARIOS.md as the long-term record (replaces the prose README hand-off); F-DP-4
RecordingSubprocessRubricRunnerlives intests/helpers/rubrics.py; F-DP-5pytest_sessionstartfuture option forscripts/seed_adversarial_fixtures.py; F-DP-6ScenarioIdnewtype future trigger; F-CON-6 held-out fence-floor vacuous-satisfaction acknowledgement; F-CON-7 cassette-free fixture rationale; F-COV-10 7-code defense-in-depth rationale.
Verdict¶
HARDENED. The story now consumes the actual HARDENED upstream contracts (S3-01 Runner.plan signature, S3-02 Runner.execute kwargs, S3-03 SubprocessRubricRunner shape + wall_clock_cap_seconds kwarg, S3-04 reserved-namespace defense, S2-01 Path signature, S1-01 typed errors), tests the actual defenses they promise (SCRUBBED_ENV via sentinel-file evidence, cwd-tempdir wipe via RecordingSubprocessRubricRunner, banned-substring runtime defense across all four ADR-0008 strings, reserved-namespace smuggling per S3-04 AC-7), and uses the structural seams the codebase already established (_SCENARIO_DISPATCH dispatch table, make_*_plan helpers, tests/helpers/rubrics.py). The fixture becomes a structurally extensible long-term defense-regression record (one row in SCENARIOS.md + one scenario dir + one dispatch entry per new defense) rather than a one-shot smoke test.
Followups not folded in (out of scope for this validator pass)¶
- S1-01 must be amended to add
FailureModeTaxonomyInvalidas a 10th typed error. The amendment is a precondition for this story's executor — surfaced in the story's ADR amendments section, but not co-validated here. - S7-01's fence-CI assertion #5 (breakdown-key smuggling parse-time) is the sibling defense to this story's runtime scenario 3. S7-01's validation pass should confirm the parse-time substring ban covers the same four ADR-0008 banned substrings.
- A future
ScenarioId = NewType("ScenarioId", str)newtype is noted but not introduced (Rule 2 — three consumers not reached). - Phase 16
microVMrubric runner: when the new isolation class lands, scenarios 1+2 grow a sibling parameterization (@pytest.mark.parametrize("rubric_runner", [SubprocessRubricRunner, MicroVMRubricRunner])). Surface as Note, do not introduce now. - Hypothesis-property hardening of the loader (
tests/adv/eval/test_loader_taxonomy_mutations.py) — programmatic mutation of valid YAML to assert "every mutation is rejected." Future hardening; this story focuses on the fixed-attack-vector portfolio.