Story S5-02 — vuln-remediation rubric (subprocess entrypoint) + bench-author unit tests¶
Step: Step 5 — Backfill bench/vuln-remediation/ with ≥10 cases + rubric + taxonomies
Status: HARDENED (phase-story-validator, 2026-07-25 — third pass: full body-vs-validation-notes sync + AC-4/AC-1..AC-11 test-file-routing reconciliation)
Effort: M
Depends on: S5-01 HARDENED (ships the VulnRemediationRubric stub class, BreakdownKey StrEnum with exactly four members, failure_modes.yaml with 11 block + 3 warn + 2 info codes; S5-02 replaces the stub's score method body byte-for-byte), S1-02 HARDENED (BenchScore / BenchCase / FailureMode Pydantic wire-type shapes — score: float [0,1], wall_clock_ms: int >= 0, breakdown: dict[str, float] typed-at-the-edge, failure_modes: tuple[FailureMode, ...], FailureMode.severity: Literal["block","warn","info"], FailureMode.detail: str | None), S1-04 (Rubric Protocol — class with one score(self, case, harness_output) -> BenchScore method; mypy --strict structural check), S2-01 HARDENED (load_task_class + autouse conftest is the import surface for the hyphenated bench/vuln-remediation/ directory). Transitively: S3-03 (the runner that subprocess-invokes this rubric; produces the envelope shape).
ADRs honored: ADR-0001 (subprocess entrypoint; if __name__ == "__main__" JSON-in/JSON-out; ≤ 60 s budget; SCRUBBED_ENV mirrors Phase 5 ADR-0012; in-process bench-author tests bypass the boundary), ADR-0004 (every emitted failure_mode.code is constrained by the taxonomy; unknown codes resolve at the runner to rubric.unknown_failure_mode; severities are the rubric's responsibility at emit time, taken from failure_modes.yaml), ADR-0008 (every emitted BenchScore.breakdown key must be a BreakdownKey value; runner rejects unknown keys as rubric.unknown_breakdown_key), Phase 5 ADR-0012 (env-allowlist SCRUBBED_ENV pattern reused for the subprocess invocation), Phase 5 ADR-0014 (banned-substring source-of-truth for breakdown keys — shared with ADR-0008)
Validation notes¶
Validated: 2026-06-04 (initial), 2026-06-05 (second pass sync — incomplete), 2026-07-25 (third pass — full sync) Verdict: HARDENED Findings addressed: 23 initial + 9 third-pass sync = 32 total
Third-pass changes (2026-07-25 — full audit log: _validation/S5-02-vuln-rubric-and-unit-tests.md):
- F-CON-SYNC-1 (BLOCK) — Implementation outline §4 rewritten to ship
bench/vuln-remediation/tests/conftest.py(autouseload_task_class(...)fixture) instead of the previously-still-referencedtests/__init__.py. The 2026-06-05 second pass claimed to make this edit but §4 still saidtests/__init__.py. - F-CON-SYNC-2 (BLOCK) — Files-to-touch table row for
tests/__init__.pyremoved; replaced withtests/conftest.pyrow. Second pass's Validation-notes bullet said this had happened but the table still had the__init__.pyrow. - F-CON-SYNC-3 (BLOCK) — Red-section TDD-plan code block replaced with the six tightened tests matching AC-4 exactly (
result.score == 1.0,result.failure_modes == (), exact set equality,== declarednot<=). The 2026-06-05 pass tightened the ACs but left the Red block using the pre-hardening thin assertions — an executor following §Red verbatim would commit a red marker that trivially passes and then re-tighten under §Green with no traceable red→green delta. - F-CON-SYNC-4 (BLOCK) — Green §1 and Refactor §2 rewritten to declare
_SEVERITY_FOR_EMITTED_CODE: Final[Mapping[str, Literal["block","warn","info"]]]per AC-10 / F-DP-1. Previously §1 said "read the YAML at module load" and §2 said "Lift the YAML severity load to module import time" — both directly contradicted the AC-10 hardcoded-severity pin the second pass added. - F-CON-SYNC-5 (HARDEN) — Implementation outline §3 (
__main__entrypoint) rewritten to instantiate_HarnessOutput.model_validate(payload["harness_output"])per AC-2 / F-DP-2. The_HarnessOutput,_ValidatorSignals,_RecipeSignalsmodel shapes are now pinned inline in §3. Previously §3 showed only raw-dict access, contradicting AC-2's Pydantic-validation requirement. - F-CON-SYNC-6 (BLOCK) — Implementation outline §5 (subprocess-test) expanded to pin the parent-process sentinel protocol AC-3 requires (
ANTHROPIC_API_KEY=parent-sentinel,AWS_ACCESS_KEY_ID=parent-sentinel,HOME=/parent-home,USER=parent-user) and the stderr debug-line assertion path. Previously §5 said only "mirror the runner's contract." - F-COV-SYNC-1 (BLOCK) — AC-4 "exactly six" reconciled with AC-1/AC-2/AC-5..AC-11. AC-4 rewritten to say exactly six core-condition tests in
bench/vuln-remediation/tests/test_rubric_unit.py. AC-1/AC-2/AC-5/AC-6/AC-7/AC-8/AC-9/AC-10/AC-11 tests are explicitly file-routed: (a) AC-1/AC-5/AC-6/AC-7/AC-8/AC-10/AC-11 → sametest_rubric_unit.pyfile as additional tests beyond the six core-condition set; (b) AC-2'stest_main_exits_nonzero_on_malformed_envelope_jsonand AC-3's subprocess-SCRUBBED_ENV tests →tests/integration/test_rubric_subprocess_vuln.py; (c) AC-9's AST test →bench/vuln-remediation/tests/test_rubric_static.py(mirrors S5-01'stest_breakdown_keys_static.pypattern); (d) AC-12 extends the existingtests/unit/test_eval_package_imports_no_llm_sdk.py. AC-4 wording changed from "exactly the following six (no fewer, no extras for this story's red→green window)" to "the following six core-condition tests (plus the additional tests pinned by AC-1/AC-5..AC-11 in the same file, and the tests pinned by AC-2/AC-3/AC-9/AC-12 in the files named in those ACs)." - F-CON-SYNC-7 (HARDEN) — Files-to-touch table extended with three rows:
bench/vuln-remediation/tests/test_rubric_static.py(AC-9),tests/unit/test_eval_package_imports_no_llm_sdk.py(AC-12 glob extension — modify existing file, not new), and moved AC-2/AC-3 subprocess entrypoint tests explicitly totests/integration/test_rubric_subprocess_vuln.py. - F-TQ-SYNC-1 (HARDEN) — Notes-for-implementer pins the
passed-derivation rule:passed = (score == 1.0), computed post-mean, NOT derived directly fromharness_outputsub-conditions. Prevents an implementation that flipspassed = all(harness_output_conditions)— indistinguishable observably from the correct form for the AC-1/AC-6 rows tested here, but silently divergent under future extension.
Second-pass claim-vs-body drift diagnosis (root cause): the 2026-06-05 pass tightened ACs and appended detailed Validation-notes bullets but the corresponding Implementation-outline / Red-TDD / Refactor / Files-to-touch edits were skipped. This third pass forcibly reconciles: body IS notes now.
Initial-pass changes (2026-06-04):
- Status line updated to
HARDENED (phase-story-validator, 2026-06-04)(F-CON-9). - Depends-on rewritten to name S5-01 HARDENED, S1-02 HARDENED (wire-type shapes), S1-04 (Rubric Protocol), S2-01 HARDENED (loader is the import surface) (F-CON-8).
- AC-1 dual-surface contract pinned (BLOCK): S5-01 HARDENED ships the rubric as a
class VulnRemediationRubricto satisfy the S1-04RubricProtocol structural check. The original TDD plan'sfrom bench.vuln_remediation.rubric import scorewould fail because no module-levelscoreexists in the S5-01 stub. AC-1 + Implementation outline §2/§3 now pin the dual surface: a module-level pure functionscore(case, harness_output) -> BenchScore(used by both the in-process tests and the__main__entrypoint directly) PLUS theVulnRemediationRubricclass whosescore(self, case, harness_output)method delegatesreturn score(case, harness_output). The class is the Protocol-conformance surface (S1-04, S5-01); the function is the test/entrypoint surface (this story). S5-01's stub body is replaced byte-for-byte (F-CON-1). bench/vuln-remediation/tests/__init__.pyhard-banned (BLOCK): the original Implementation outline §4 creates this file as an "empty package marker." That directly contradicts S5-01 F-CON-5's hard ban: S2-01 HARDENED uses PEP 420 implicit namespace packages — no__init__.pyfiles anywhere underbench/. Tests discovery does not need it (pytest's rootdir-based discovery + a siblingconftest.pyis sufficient). Implementation outline §4 rewritten to shipbench/vuln-remediation/tests/conftest.pyinstead — an autouseload_task_class("vuln-remediation", bench_root=...)fixture that registersbench.vuln_remediation.*in sys.modules sofrom bench.vuln_remediation.rubric import scoreresolves. Files-to-touch row updated (F-CON-2).- Bench-author conftest is the import bridge (BLOCK): without the conftest, standard Python
from bench.vuln_remediation.rubric import scorecannot resolve the hyphenated on-disk directory under standard Python import machinery (same issue as S5-01 F-CON-3). The conftest callsload_task_class("vuln-remediation", bench_root=REPO_ROOT / "bench")before any test imports, populating sys.modules under the underscore key viaspec_from_file_location. Implementation outline §4 pins the conftest body; AC-1 pins the import path resolves; Notes-for-implementer documents why the conftest exists and what would break without it (F-CON-3). - Semantic-symmetry inversions pinned as ACs (BLOCK): S5-01 HARDENED documented the four breakdown↔failure-mode pairs (
cve.dropped↔validator.cve_not_dropped,validator.build_passed↔validator.build_failed,validator.tests_passed↔validator.tests_failed,recipe.applied↔recipe.semantic_drift) but explicitly deferred enforcement to this story ("the rubric (S5-02) owns the score↔failure inversion"). The original AC-4 only tested two inversions (validator.tests_failed,validator.cve_not_dropped); a wrong implementation that emittedrubric.unknown_failure_modefor the other two would pass the existing tests. New AC-5 explicitly pins all four inversions via a parametrized test (test_each_falsy_breakdown_condition_emits_its_paired_failure_code) (F-COV-1). - Static AST ban on non-determinism (BLOCK): the Notes-for-implementer line on determinism ("no
time.time(), norandom.random(), noos.environreads, nouuid.uuid4()") is not enforceable — a contributor addingtime.time()to the rubric would silently break the audit chain's byte-stability and cause S5-06's cache hit-rate to fall below 95% with no surfaced cause. New AC-9 +test_rubric_module_has_no_nondeterministic_imports_or_callsASTsrubric.pyand rejects:import time/time.X,import random/random.X,import uuid/uuid.X,os.environaccess,datetime.now(,datetime.utcnow(. Mirrors S5-01'stest_breakdown_key_values_are_ast_constant_stringspattern (F-COV-2). - Subprocess SCRUBBED_ENV ACs pinned (BLOCK): the original AC-3 only asserted "completes within 60 wall-clock seconds." ADR-0001 §Decision and Phase 5 ADR-0012 require:
ANTHROPIC_API_KEY,AWS_*,HOME,USERall absent inside the rubric subprocess. The original integration test could pass while leaving any of these reachable. New AC-3 expanded: integration test sets each env var to a sentinel value in the parent process, runs the rubric subprocess withSCRUBBED_ENV, and asserts via the rubric's debug-emit-on-stderr thatos.environ.get("ANTHROPIC_API_KEY")isNone,os.environ.get("AWS_ACCESS_KEY_ID")isNone,os.environ.get("HOME")isNone,os.environ.get("USER")isNone. Mirrorstests/adv/test_rubric_subprocess_env_scrubbed.pyfrom arch line 296 (F-COV-7). - Malformed-JSON exit-code AC (BLOCK): ADR-0001 §Consequences pins
rubric.malformed_outputas the runner's reaction to a rubric non-zero exit. The original story does not pin that the rubric's__main__exits non-zero on bad input — it couldtry/exceptand emit a passingBenchScore. New AC-2 +test_main_exits_nonzero_on_malformed_envelope_jsonfeeds the subprocessb"not-json"on stdin and asserts the process exits with code != 0 (F-COV-8). - Test assertions tightened against trivial mutants (HARDEN):
- AC-4(a)
result.score >= 0.95→result.score == 1.0(kills a "score = 0.95" hardcoded-return mutant) (F-TQ-1). - AC-4(a)
all(fm.severity != "block")→result.failure_modes == ()(kills "emit info-severity on full-pass" mutant) (F-TQ-2). - AC-4(b)/(c)
"X" in {fm.code ...}→ exact set equality{fm.code for fm in result.failure_modes} == {"X"}(kills "also emit a spurious failure" mutant) (F-TQ-3). - AC-4(d)
set(result.breakdown.keys()) <= declared→== declared(kills "ship only a subset" mutant; Implementation outline §2 already produces all four keys, so equality is the right contract) (F-TQ-4). - Mean-formula mutation test (HARDEN): new
test_half_pass_yields_score_exactly_half_kills_min_max_mutants— forharness_outputwith exactly 2 of 4 sub-conditions true, assertsresult.score == 0.5. Kills mutants wheremeanis replaced bymin(0.0),max(1.0),len(failing)(2.0), orsum(2.0) (F-TQ-5). - BenchCase-invariance property test (HARDEN): new
test_score_invariant_under_unrelated_case_field_mutations— the rubric readsharness_output, notcase.*(Notes-for-implementer line 233 is explicit). For a fixed harness_output, mutatingcase.case_id,case.difficulty,case.dispositionmust not change the resultingBenchScore.model_dump_json(). Catches accidental reads ofcase.*that would couple the rubric to case shape and break cache hit-rate (F-TQ-6). - Canonical
failure_modesordering AC (HARDEN): for the determinism contract (audit chain byte-stability), thefailure_modestuple ordering must be canonical for any given falsy-condition set. New AC-7 +test_failure_modes_tuple_is_sorted_by_codepins lexicographic ordering byfm.code. Without this, two equivalent runs could emit(A, B)vs(B, A)→model_dump_jsonbytes differ → cache miss + audit chain divergence. Implementation outline §2 amended to sort the emit set (F-COV-3). - Missing-harness-key behavior pinned (HARDEN): new AC-8 +
test_missing_harness_output_key_propagates_keyerrorfeedsharness_output = {"validator": {}, "recipe": {"applied": True}}and assertspytest.raises(KeyError). Pins the "fail loud" behavior Notes-for-implementer line 234 already documents but no test enforces. A defensiveharness_output.get("validator", {}).get("build_passed", False)implementation would silently downgrade a missing key toFalse(failing run) instead of surfacing the contract violation (F-COV-4). - Exactly six unit tests (HARDEN): AC-4 changed from "at least 5" to "exactly 6 (named):" pinning the test set so the executor cannot under-cover. The six tests cover the original five plus the deterministic-replay one already in §TDD plan (F-COV-9).
- Severities hardcoded; YAML consistency test (HARDEN, F-DP-1): the Refactor section's instruction to "lift the YAML severity load to module import time (single I/O); cache as
_TAXONOMY" creates a second YAML reader (S5-01'sregistration.pyalready has_severity_taxonomy_from_yaml). By end of Phase 6.5 this is the 4th YAML reader (S5-01 reg + this story's rubric + S6-01 migration reg + S6-03 migration rubric) — well past rule-of-three. Pin a tighter alternative: the rubric emits exactly four block codes (validator.build_failed,validator.tests_failed,validator.cve_not_dropped,recipe.semantic_drift), all known at bench-author time. Implementation outline §2 now declares_SEVERITY_FOR_EMITTED_CODE: Final[Mapping[str, Literal["block","warn","info"]]]hardcoded with these four → "block". A consistency test (test_hardcoded_severities_match_failure_modes_yaml) reads the YAML and asserts each hardcoded severity matches; YAML drift surfaces at PR time. Avoids brittle import-time I/O (cwd-relative path resolution under subprocess), avoids the second-loader rule-of-three trigger, and pins the load-bearing severity decision at the bench-author's source-of-truth. The rule-of-three lift (when the third task class lands in Phase 15) becomes a kernel_load_failure_mode_taxonomyinsrc/codegenie/eval/loader.pyshared by the registration path; rubrics continue to hardcode their emit-set severities. Notes-for-implementer explains the tradeoff. - HarnessOutput envelope endorsement (HARDEN, F-DP-2): the Refactor instruction "destructure with pydantic BaseModel for the envelope" is currently a Notes aside. Promoted to Implementation outline §3 as a concrete
_HarnessOutput(BaseModel, frozen=True, extra="forbid")withvalidator: _ValidatorSignalsandrecipe: _RecipeSignalssub-models. This (a) gives mypy --strict a real shape to verify; (b) surfaces SUT contract drift at the envelope-parse layer instead of atKeyErrordeep inscore(); (c) makes the contract between Phase 6'sVulnRemediationSut.run_caseand the rubric explicit and testable. Pinned in Notes-for-implementer as a Phase 6→Phase 6.5 contract-surface decision. Local-only (not a shared model) per Rule 2 — Phase 7's migration rubric will have a different SUT contract. - Adversarial mutant catalog added (HARDEN, F-TQ-7): Notes-for-implementer surfaces the six named mutants this §TDD kills, mirroring S5-01's pattern.
Files to touchaligned:bench/vuln-remediation/tests/__init__.pyrow dropped;bench/vuln-remediation/tests/conftest.pyrow added.- ADRs honored expanded to name Phase 5 ADR-0012 (env-allowlist source-of-truth for SCRUBBED_ENV) and Phase 5 ADR-0014 (substring-ban source-of-truth shared with ADR-0008).
Design endorsements (no edit; surfaced in Notes-for-implementer):
- Functional-core / imperative-shell — already followed (pure score() + __main__ shell). Reaffirmed.
- Open/Closed seam at bench/{task-class}/rubric.py — Phase 7's migration rubric copies this pattern verbatim. Reaffirmed.
- Strategy pattern for condition→failure-code mapping — _CONDITION_FAILURE_PAIRS: Final[tuple[tuple[BreakdownKey, str], ...]] keeps the rubric loop-driven and Open/Closed at the inversion table.
No NEEDS RESEARCH items — every pattern is precedented in this repo (Phase 5 ADR-0012 env-allowlist test discipline, S5-01 HARDENED AST-walk pattern, S1-02 boundary-inclusivity test pattern).
Context¶
The rubric is control-plane code: it produces BenchScore, which feeds the promotion gate, which determines whether a task class graduates. ADR-0001 makes the rubric a subprocess entrypoint specifically because it lives under bench/**, a CODEOWNERS-gated path that any contributor may PR — the runner therefore never imports it. The bench-author writes the rubric to a precise contract: read a JSON envelope (containing the BenchCase shape + the SUT's harness_output) from stdin; emit a BenchScore JSON to stdout; terminate in ≤ 60 s; produce no other side effects.
The trusted boundary distinction is load-bearing: bench/vuln-remediation/tests/test_rubric_unit.py may import the rubric module directly (in-process) and test its score(...) function with hand-built fixtures. The harness runner never imports it. This split is what makes the rubric simultaneously (a) testable with normal pytest ergonomics during bench-author development and (b) safe to invoke across a process boundary in production runs.
References — where to look¶
- Architecture:
../phase-arch-design.md §Component design → src/codegenie/eval/rubric.py— theRubricProtocol (score(case, harness_output) -> BenchScore); the bench-author'sscore(...)function must satisfy it for in-process unit tests, even though the runner crosses a subprocess boundary.../phase-arch-design.md §Control flow— the subprocess invocation shape (subprocess.run(rubric.py, env=SCRUBBED, stdin=JSON, timeout)).../phase-arch-design.md §Edge cases #3, #4, #5— non-zero exit, timeout, malformed JSON: all becomeFailureMode(severity="block")at the runner; the rubric does not need to handle them, but must not swallow internal exceptions and emit a misleadingly-passing score.../phase-arch-design.md §Harness engineering → Tracing strategy— the rubric is allowed to emitstructlogJSON on stderr; stdout is reserved for theBenchScoreenvelope.- Phase ADRs:
../ADRs/0001-rubric-execution-isolation-via-subprocess.md §Decision, §Consequences—if __name__ == "__main__":entrypoint is the bench-author's load-bearing surface; bench-author tests verify bothscore(...)(in-process) and the subprocess CLI (python rubric.py < stdin > stdout).../ADRs/0004-per-task-class-failure-modes-taxonomy.md §Consequences— the rubric emitsfailure_mode_code: str; the runner resolves it against the taxonomy. Unknown codes becomerubric.unknown_failure_mode(block-severity) — fail loud on drift.../ADRs/0008-breakdown-keys-strenum-with-substring-ban.md §Consequences—BenchScore.breakdowndict keys must beBreakdownKeyvalues; mismatched keys producerubric.unknown_breakdown_keyat runtime.- Source design:
../High-level-impl.md §Step 5— the rubric scores recipe-applied + validator-passed + cve-dropped signals; specific scoring formula is rubric-author judgment.
Goal¶
Implement bench/vuln-remediation/rubric.py as a deterministic subprocess entrypoint that reads a JSON envelope from stdin, emits a BenchScore JSON to stdout in ≤ 60 s per case, and is covered by in-process bench-author unit tests in bench/vuln-remediation/tests/test_rubric_unit.py.
Acceptance criteria¶
-
[ ] AC-1 (dual-surface contract: module-level
score+ Protocol-conforming class).bench/vuln-remediation/rubric.pydefines BOTH (a) a module-level pure functionscore(case: BenchCase, harness_output: Mapping[str, Any]) -> BenchScore(used by in-process tests and the__main__entrypoint directly), AND (b)class VulnRemediationRubricwhosescore(self, case, harness_output)method body is exactlyreturn score(case, harness_output). The class is the S1-04RubricProtocol-conformance surface S5-01 registered against; the function is the test/entrypoint surface this story exercises. Testtest_class_score_method_delegates_to_module_level_scoresweeps the five-row condition matrix from AC-4(f) and asserts byte-equality ofmodel_dump_json()across both call paths for every row:VulnRemediationRubric().score(case, ho).model_dump_json() == score(case, ho).model_dump_json(). (Note:is-identity would false-fail — Pydanticfrozen=Truemodels from two separate calls are distinct instances; byte-equality across the full matrix is what kills a divergent-class-method mutant.) The returnedBenchScoreis frozen, hasbreakdownkeys drawn exactly fromBreakdownKeyvalues, andfailure_modes[*].codevalues drawn exactly from the codes declared infailure_modes.yaml. -
[ ] AC-2 (
__main__entrypoint shape + non-zero exit on malformed input).rubric.pyhas anif __name__ == "__main__":block that: readssys.stdin.buffer.read(), parses it as JSON, validates into a typed envelope (BenchCase+harness_output) via the local_HarnessOutputPydantic model, callsscore(...), writes the resultingBenchScoreas JSON tosys.stdout.buffer, and exits 0 on success. Onjson.JSONDecodeErrororpydantic.ValidationErrorthe process exits non-zero (sys.exit(2)); testtest_main_exits_nonzero_on_malformed_envelope_jsonfeedsb"not-json"on stdin viasubprocess.runand assertsreturncode != 0. The rubric does not wrapscore(...)in a broadtry/exceptthat would emit a misleadingly-passingBenchScoreon internal failure (ADR-0001 §Consequencesrubric.malformed_outputis the runner's reaction to non-zero exit). -
[ ] AC-3 (subprocess SCRUBBED_ENV + ≤60 s wall-clock).
tests/integration/test_rubric_subprocess_vuln.pyrunspython bench/vuln-remediation/rubric.pyviasubprocess.runwithenv=SCRUBBED_ENV(containing onlyPYTHONPATH,PYTHONHASHSEED=0, minimalPATHper ADR-0001 §Decision and Phase 5 ADR-0012 env-allowlist precedent) andcwd=tempfile.TemporaryDirectory(). The parent process setsANTHROPIC_API_KEY=parent-sentinel,AWS_ACCESS_KEY_ID=parent-sentinel,HOME=/parent-home,USER=parent-userbefore spawn; the rubric writes a debug line to stderr (stdout is reserved for theBenchScore) reportingos.environ.get("ANTHROPIC_API_KEY"),os.environ.get("AWS_ACCESS_KEY_ID"),os.environ.get("HOME"),os.environ.get("USER"); the test asserts each isNonein the rubric's environment. Wall-clock ≤ 60 s on a representative envelope (the four-positive-condition envelope from AC-4(a)). -
[ ] AC-4 (six core-condition in-process unit tests — pinned, named, tightened).
bench/vuln-remediation/tests/test_rubric_unit.pyexists and contains the following six core-condition tests (plus the additional tests pinned by AC-1, AC-5, AC-6, AC-7, AC-8, AC-10, AC-11 in the same file — see §Test-file routing below — and the tests pinned by AC-2, AC-3, AC-9, AC-12 in the files named in those ACs). Test-file routing (executor discipline):bench/vuln-remediation/tests/test_rubric_unit.pyholds AC-1 + AC-4(a–f) + AC-5 + AC-6 + AC-7 + AC-8 + AC-10 + AC-11 (unit-level, in-process).tests/integration/test_rubric_subprocess_vuln.pyholds AC-2 (test_main_exits_nonzero_on_malformed_envelope_json) + AC-3 (SCRUBBED_ENV subprocess tests).bench/vuln-remediation/tests/test_rubric_static.pyholds AC-9 (test_rubric_module_has_no_nondeterministic_imports_or_calls) — mirrors S5-01'stest_breakdown_keys_static.pysibling pattern.tests/unit/test_eval_package_imports_no_llm_sdk.pyis extended (not new) per AC-12. The six core-condition tests are:- (a)
test_full_pass_yields_score_one_passed_true_no_failure_modes— when all four sub-conditions areTrue:result.passed is True,result.score == 1.0(exact, not>= 0.95),result.failure_modes == ()(exact, not "no block-severity"). - (b)
test_tests_failed_emits_exactly_validator_tests_failed_block— when onlyvalidator.tests_passed=False:result.passed is False,{fm.code for fm in result.failure_modes} == {"validator.tests_failed"}(exact set), and that oneFailureMode.severity == "block". - (c)
test_cve_not_dropped_emits_exactly_validator_cve_not_dropped_block— when onlyvalidator.cve_dropped=False:{fm.code for fm in result.failure_modes} == {"validator.cve_not_dropped"}, that one severity is"block". - (d)
test_breakdown_keys_equal_full_declared_breakdown_key_set— for anyharness_output,set(result.breakdown.keys()) == {m.value for m in BreakdownKey}(exact equality, not subset — Implementation outline §2 produces all four keys always). - (e)
test_score_is_deterministic_under_repeated_invocation— same inputs → byte-identicalmodel_dump_json()across 10 invocations (tighter than the original 2). - (f)
test_rubric_emits_only_declared_failure_mode_codes_and_breakdown_keys— sweep the five-row condition matrix from the original §TDD plan; every emittedfm.code∈ YAML-declared codes; every breakdown key ∈ BreakdownKey values.
- (a)
-
[ ] AC-5 (semantic-symmetry inversions — all four pairs enforced). A parametrized test
test_each_falsy_breakdown_condition_emits_its_paired_failure_codepins all four S5-01-documented inversions:validator.build_passed=False→{"validator.build_failed"}(block);validator.tests_passed=False→{"validator.tests_failed"}(block);cve.dropped=False→{"validator.cve_not_dropped"}(block);recipe.applied=False→{"recipe.semantic_drift"}(block). For each case, all other sub-conditions areTrueand the only emitted failure-code set is the singleton paired code. Kills the "emitrubric.unknown_failure_modefor everything" mutant that the original 2-of-4 coverage missed. -
[ ] AC-6 (mean-formula + half-pass mutant kill).
test_half_pass_yields_score_exactly_half_kills_min_max_mutants— forharness_outputwith exactly 2 of the 4 sub-conditionsTrue(e.g.,build_passed=True, tests_passed=False, cve_dropped=True, recipe.applied=False): assertsresult.score == 0.5(exact). Kills mean→min (0.0), mean→max (1.0), mean→sum (2.0), mean→len(failing) (2.0) mutants.result.failure_modeshas exactly two entries with{"validator.tests_failed", "recipe.semantic_drift"}(defense-in-depth on AC-5). -
[ ] AC-7 (canonical
failure_modestuple ordering).test_failure_modes_tuple_is_sorted_by_code— for a multi-failure case (all four sub-conditionsFalse),result.failure_modesis atupleand[fm.code for fm in result.failure_modes] == sorted([fm.code for fm in result.failure_modes])(lexicographic bycode). Required formodel_dump_json()byte-stability across runs — without this, S5-06's cache-hit-rate test silently degrades and the audit chain diverges (ADR-0001 audit-comparability). -
[ ] AC-8 (fail-loud on missing harness_output keys).
test_missing_harness_output_key_propagates_keyerror— feedsharness_output = {"validator": {}, "recipe": {"applied": True}}(missingbuild_passed,tests_passed,cve_dropped) and assertspytest.raises(KeyError). Pins the "let it propagate; runner recordsrubric.malformed_output" behavior Notes-for-implementer documents. A defensiveharness_output.get("validator", {}).get("build_passed", False)would silently flip a contract violation into a failing-but-passing score. -
[ ] AC-9 (static AST ban on non-determinism inside
rubric.py).test_rubric_module_has_no_nondeterministic_imports_or_callsparsesbench/vuln-remediation/rubric.pyviaast.parseand asserts: noimport time/from time, noimport random/from random, noimport uuid/from uuid, noos.environattribute access, nodatetime.now(/datetime.utcnow(/datetime.today(calls. Mirrors S5-01'stest_breakdown_key_values_are_ast_constant_stringsAST-walk pattern. Required forBenchScore.model_dump_json()byte-stability (audit chain + S5-06 cache hit rate); without it, a contributor addingwall_clock_ms = int((time.perf_counter() - t0) * 1000)would break determinism with no surfaced cause until S5-06 fails N stories later. Note: the__main__block'swall_clock_msreporting must use a deterministic fixed sentinel (e.g.,wall_clock_ms = 0) — see Implementation outline §3. -
[ ] AC-10 (hardcoded-severity ↔ YAML consistency).
rubric.pydeclares_SEVERITY_FOR_EMITTED_CODE: Final[Mapping[str, Literal["block","warn","info"]]]mapping the four codes the rubric can emit (validator.build_failed,validator.tests_failed,validator.cve_not_dropped,recipe.semantic_drift) to their severities (all"block"). Testtest_hardcoded_severities_match_failure_modes_yamlreadsfailure_modes.yamland assertsyaml_taxonomy[code]["severity"] == _SEVERITY_FOR_EMITTED_CODE[code]for each code. YAML drift (e.g., a contributor downgradingvalidator.tests_failedtowarnwithout amending the rubric) surfaces at PR time. Avoids the brittle import-time YAML-read alternative the original Refactor section prescribed (see F-DP-1). -
[ ] AC-11 (BenchCase-invariance — rubric reads
harness_output, notcase).test_score_invariant_under_unrelated_case_field_mutations— fixesharness_outputand mutatescase.case_id,case.difficulty,case.dispositionacross three constructions; asserts all three produce the samemodel_dump_json(). Pins the "rubric scores fromharness_outputdirectly" Notes-for-implementer contract; catches accidental coupling that would break cache hit-rate. -
[ ] AC-12 (no LLM SDK + fence-CI extension). The rubric does not import any LLM SDK (
anthropic,openai,langchain,langgraph,transformers,torch,sentence-transformers).tests/unit/test_eval_package_imports_no_llm_sdk.pyis extended to walkbench/**/rubric.py(the existing walk currently only coverssrc/codegenie/eval/**/*.py); the extension is a single-glob addition. Test stays green for this story's rubric. -
[ ] AC-13 (red→green pipeline, lint, typecheck, fence-CI). Red test from §TDD plan exists, was committed at red marker, now green.
ruff check,ruff format --check,mypy --strict bench/vuln-remediation/rubric.py bench/vuln-remediation/tests/test_rubric_unit.py bench/vuln-remediation/tests/conftest.py tests/integration/test_rubric_subprocess_vuln.py, andpytest bench/vuln-remediation/tests/ tests/integration/test_rubric_subprocess_vuln.pyall pass. S7-01's fence-CI assertions (#4 literal name; #5 BreakdownKey substring ban; #6 taxonomy validity) all stay green on the modified files.
Implementation outline¶
- Write the red test
bench/vuln-remediation/tests/test_rubric_unit.pyfirst — see §TDD plan. - Implement
score(case, harness_output)as a pure function AND theVulnRemediationRubricclass per AC-1 (dual-surface contract). Parseharness_outputvia a file-local_HarnessOutput(BaseModel, frozen=True, extra="forbid")with sub-models_ValidatorSignals(build_passed: bool, tests_passed: bool, cve_dropped: bool)and_RecipeSignals(applied: bool)— surfaces SUT contract drift at the envelope-parse layer instead ofKeyErrordeep in scoring. Computebreakdownas{BreakdownKey.X.value: 1.0 if condition else 0.0, ...}producing all four BreakdownKey values (AC-4(d) equality). Computescore = mean(breakdown.values()); computepassed = (score == 1.0)(derived post-mean, NOTall(harness_output_conditions)— see Notes-for-implementer). Computefailure_modesby mapping each falsy condition to its paired failure code via_CONDITION_FAILURE_PAIRS: Final[tuple[tuple[BreakdownKey, str], ...]]and looking up severity from_SEVERITY_FOR_EMITTED_CODE: Final[Mapping[str, Literal["block","warn","info"]]](hardcoded — the four codes the rubric can emit:validator.build_failed,validator.tests_failed,validator.cve_not_dropped,recipe.semantic_drift, all"block"). Sort the emittedfailure_modestuple lexicographically byfm.codeper AC-7.VulnRemediationRubric.score(self, case, harness_output)body is exactlyreturn score(case, harness_output)(AC-1 delegation). - Implement the
if __name__ == "__main__":entrypoint:from typing import Final, Literal, Mapping from pydantic import BaseModel, ConfigDict class _ValidatorSignals(BaseModel): model_config = ConfigDict(frozen=True, extra="forbid") build_passed: bool tests_passed: bool cve_dropped: bool class _RecipeSignals(BaseModel): model_config = ConfigDict(frozen=True, extra="forbid") applied: bool class _HarnessOutput(BaseModel): model_config = ConfigDict(frozen=True, extra="forbid") validator: _ValidatorSignals recipe: _RecipeSignals _SEVERITY_FOR_EMITTED_CODE: Final[Mapping[str, Literal["block", "warn", "info"]]] = { "validator.build_failed": "block", "validator.tests_failed": "block", "validator.cve_not_dropped": "block", "recipe.semantic_drift": "block", } if __name__ == "__main__": import json import sys from pydantic import ValidationError from codegenie.eval.models import BenchCase, BenchScore try: payload = json.loads(sys.stdin.buffer.read()) case = BenchCase.model_validate(payload["case"]) envelope = _HarnessOutput.model_validate(payload["harness_output"]) except (json.JSONDecodeError, ValidationError, KeyError): sys.exit(2) # `wall_clock_ms` is a deterministic fixed sentinel (0) per AC-9 — no # `time.perf_counter()` inside the module (breaks byte-stability). result = score(case, envelope.model_dump()) sys.stdout.buffer.write(result.model_dump_json().encode("utf-8")) sys.exit(0) - Implement
bench/vuln-remediation/tests/conftest.pyas an autouse fixture that callsload_task_class("vuln-remediation", bench_root=REPO_ROOT / "bench")before any test imports, sofrom bench.vuln_remediation.rubric import scoreresolves via S2-01's hyphen→underscorespec_from_file_locationbridge. Do NOT createbench/vuln-remediation/tests/__init__.py— S5-01 F-CON-5 hard-banned it; S2-01 HARDENED uses PEP 420 implicit namespace packages. The autouse conftest is the sole import bridge for this hyphenated leaf. - Write
tests/integration/test_rubric_subprocess_vuln.pyto exercise the subprocess path. Set parent-process env sentinels (ANTHROPIC_API_KEY=parent-sentinel,AWS_ACCESS_KEY_ID=parent-sentinel,HOME=/parent-home,USER=parent-user) viamonkeypatch.setenv(...)BEFOREsubprocess.run("python", str(RUBRIC_PATH), env=SCRUBBED_ENV, cwd=tempfile.TemporaryDirectory(), input=<envelope-JSON-bytes>, capture_output=True, timeout=60). Rubric writes a debug line on stderr (stdout reserved forBenchScore) reportingos.environ.get("<var>")for each; test parses stderr and asserts each reported value isNone. Second test in the same file:test_main_exits_nonzero_on_malformed_envelope_jsonfeedsb"not-json"on stdin and assertsreturncode != 0(AC-2). Assert wall-clock ≤ 60 s on the four-positive-condition envelope (AC-3 budget). - Write
bench/vuln-remediation/tests/test_rubric_static.pyfor AC-9: parsebench/vuln-remediation/rubric.pyviaast.parse; walk forimport time,from time import ...,import random,from random import ...,import uuid,from uuid import ...,Attribute(value=Name(id="os"), attr="environ"),Call(func=Attribute(attr="now" | "utcnow" | "today")). Assert none present. Mirrors S5-01'stest_breakdown_key_values_are_ast_constant_stringssibling pattern. - Extend
tests/unit/test_eval_package_imports_no_llm_sdk.py(existing file per S1-05) — addbench/**/rubric.pyto its glob; the AST-walk for LLM-SDK imports is structurally identical.
TDD plan — red / green / refactor¶
Red — write the failing test first¶
Test file path: bench/vuln-remediation/tests/test_rubric_unit.py
# bench/vuln-remediation/tests/test_rubric_unit.py
"""In-process bench-author tests. The runner crosses a subprocess boundary;
these tests bypass that boundary because bench/**/tests/ is a trusted edge
(per ADR-0001 §Decision)."""
import json
from datetime import UTC, datetime
from pathlib import Path
import pytest
from bench.vuln_remediation.breakdown_keys import BreakdownKey
from bench.vuln_remediation.rubric import score
from codegenie.eval.models import BenchCase, BenchScore
def _make_case(case_id: str = "001-cve-2025-12345-rag-corpus-derived") -> BenchCase:
return BenchCase(
case_id=case_id,
task_class="vuln-remediation",
disposition="positive",
difficulty="easy",
source="curated",
curation_class="rag-corpus-derived",
commit_sha=None,
added_at=datetime(2026, 5, 12, tzinfo=UTC),
last_validated_at=datetime(2026, 5, 12, tzinfo=UTC),
input_path=Path("/tmp/fake/input"),
expected_path=Path("/tmp/fake/expected"),
cassette_path=Path("/tmp/fake/cassette"),
cassette_canary_pin="0" * 32,
case_digest="blake3:" + "0" * 64,
)
def test_full_pass_yields_score_one_passed_true_no_failure_modes():
"""AC-4(a) — full-pass row. Tight assertions kill 'score=0.95 hardcoded'
and 'emit info-severity on full-pass' mutants."""
case = _make_case()
harness_output = {
"validator": {"build_passed": True, "tests_passed": True, "cve_dropped": True},
"recipe": {"applied": True},
}
result = score(case, harness_output)
assert isinstance(result, BenchScore)
assert result.passed is True
assert result.score == 1.0 # exact, not >= 0.95 (F-TQ-1)
assert result.failure_modes == () # exact, not "no block-severity" (F-TQ-2)
def test_tests_failed_emits_exactly_validator_tests_failed_block():
"""AC-4(b) — exact-set kills 'also emit a spurious failure' mutant."""
case = _make_case()
harness_output = {
"validator": {"build_passed": True, "tests_passed": False, "cve_dropped": True},
"recipe": {"applied": True},
}
result = score(case, harness_output)
assert result.passed is False
assert {fm.code for fm in result.failure_modes} == {"validator.tests_failed"}
only = next(iter(result.failure_modes))
assert only.severity == "block"
def test_cve_not_dropped_emits_exactly_validator_cve_not_dropped_block():
"""AC-4(c) — exact-set kills 'also emit a spurious failure' mutant."""
case = _make_case()
harness_output = {
"validator": {"build_passed": True, "tests_passed": True, "cve_dropped": False},
"recipe": {"applied": True},
}
result = score(case, harness_output)
assert result.passed is False
assert {fm.code for fm in result.failure_modes} == {"validator.cve_not_dropped"}
only = next(iter(result.failure_modes))
assert only.severity == "block"
def test_breakdown_keys_equal_full_declared_breakdown_key_set():
"""AC-4(d) — exact equality kills 'ship only a subset' mutant. Implementation
outline §2 produces all four keys always regardless of condition state."""
case = _make_case()
harness_output = {
"validator": {"build_passed": True, "tests_passed": False, "cve_dropped": False},
"recipe": {"applied": True},
}
result = score(case, harness_output)
assert set(result.breakdown.keys()) == {m.value for m in BreakdownKey} # == not <= (F-TQ-4)
def test_score_is_deterministic_under_repeated_invocation():
"""AC-4(e) — audit chain byte-stability. Ten invocations, not two (tighter)."""
case = _make_case()
harness_output = {
"validator": {"build_passed": True, "tests_passed": True, "cve_dropped": True},
"recipe": {"applied": True},
}
dumps = [score(case, harness_output).model_dump_json() for _ in range(10)]
assert len(set(dumps)) == 1 # all ten byte-identical
def test_rubric_emits_only_declared_failure_mode_codes_and_breakdown_keys():
"""AC-4(f) — five-row condition matrix; every emitted code ∈ YAML-declared set,
every breakdown key ∈ BreakdownKey values. Defense-in-depth on top of runner."""
import yaml
yaml_text = (Path(__file__).parent.parent / "failure_modes.yaml").read_text()
declared_codes = set(yaml.safe_load(yaml_text).keys())
declared_keys = {m.value for m in BreakdownKey}
case = _make_case()
for build, tests, cve, recipe in [
(True, True, True, True),
(False, True, True, True),
(True, False, True, True),
(True, True, False, True),
(True, True, True, False),
]:
result = score(case, {
"validator": {"build_passed": build, "tests_passed": tests, "cve_dropped": cve},
"recipe": {"applied": recipe},
})
for fm in result.failure_modes:
assert fm.code in declared_codes, f"undeclared code: {fm.code}"
assert set(result.breakdown.keys()) == declared_keys # == not <= (F-TQ-4)
Additional tests in the same file pinned by AC-1 / AC-5 / AC-6 / AC-7 / AC-8 / AC-10 / AC-11 (see AC bodies for full assertions; each is a distinct named test):
test_class_score_method_delegates_to_module_level_score(AC-1) — sweeps the five-row matrix and asserts byte-equality ofmodel_dump_json()betweenVulnRemediationRubric().score(case, ho)andscore(case, ho).test_each_falsy_breakdown_condition_emits_its_paired_failure_code(AC-5) — parametrized across the four semantic-symmetry inversions.test_half_pass_yields_score_exactly_half_kills_min_max_mutants(AC-6) —result.score == 0.5for 2-of-4-true.test_failure_modes_tuple_is_sorted_by_code(AC-7) — lexicographic sort on multi-failure row.test_missing_harness_output_key_propagates_keyerror(AC-8) —pytest.raises((KeyError, ValidationError))on malformed envelope.test_hardcoded_severities_match_failure_modes_yaml(AC-10) — YAML/_SEVERITY_FOR_EMITTED_CODEconsistency.test_score_invariant_under_unrelated_case_field_mutations(AC-11) — case-shape-invariance.
Tests routed to sibling files (AC-4 §Test-file routing): test_rubric_static.py (AC-9 AST-ban), tests/integration/test_rubric_subprocess_vuln.py (AC-2 malformed-JSON + AC-3 SCRUBBED_ENV).
Run it; confirm ModuleNotFoundError: No module named 'bench.vuln_remediation.rubric' or ImportError: cannot import name 'score'. Commit as red marker.
Green — smallest impl shape¶
- Implement
score(case, harness_output) -> BenchScore: - Parse
harness_outputvia_HarnessOutput.model_validate(harness_output)(Pydantic; frozen;extra="forbid"). A missing sub-key surfaces aspydantic.ValidationError, which AC-8 pins as "let it propagate" (do NOT swallow). - Build
breakdownby iterating the fixed_CONDITION_FAILURE_PAIRStable — each pair(BreakdownKey.X, failure_code)yieldsbreakdown[BreakdownKey.X.value] = 1.0 if condition else 0.0. All four BreakdownKey values appear in every result. - Compute
score = sum(breakdown.values()) / len(breakdown). - Compute
passed = (score == 1.0)— derived from the score post-mean, NOT fromall(harness_output_conditions). Notes-for-implementer pins the rationale. - For each falsy condition, emit
FailureMode(code=paired_code, severity=_SEVERITY_FOR_EMITTED_CODE[paired_code], detail=None). Severity is looked up from the hardcoded_SEVERITY_FOR_EMITTED_CODE: Final[Mapping[str, Literal["block","warn","info"]]](four entries, all"block") — AC-10's YAML-consistency test guards drift. - Sort emitted
failure_modeslexicographically byfm.code(AC-7 byte-stability). - Return
BenchScore(passed=..., score=..., breakdown=..., failure_modes=tuple(sorted_fms), wall_clock_ms=0, cost_usd=0.0). - Implement the
__main__entrypoint as in §Implementation outline §3 — exit 2 onjson.JSONDecodeError | pydantic.ValidationError | KeyError. - Run the test suite; iterate until green.
Refactor — clean up¶
- Keep the condition-to-code mapping as the module-level
_CONDITION_FAILURE_PAIRS: Final[tuple[tuple[BreakdownKey, str], ...]]used in §Green — loop-driven, Open/Closed at the inversion table (adding a fifth breakdown key = one tuple row + one YAML entry + one_SEVERITY_FOR_EMITTED_CODErow; zero branching-code edits). - Keep
_SEVERITY_FOR_EMITTED_CODEhardcoded at module top-level. Do NOT lift the YAML at import time — F-DP-1 pins the rationale (avoids brittle cwd-relative I/O under subprocess, avoids the second-loader rule-of-three trigger, keeps subprocess cold-start under the ADR-0001 spawn budget). The rule-of-three lift target (Phase 15's third task class) issrc/codegenie/eval/loader.py::_load_failure_mode_taxonomyper arch line 564. - Add a module docstring naming ADR-0001, ADR-0004, ADR-0008 and the "trusted boundary distinction" between in-process bench-author tests and the runner's subprocess invocation.
mypy --strictclean:harness_outputaccess goes through_HarnessOutput.model_validate(...), not raw dict indexing. The Pydantic model IS the type contract.wall_clock_msis a deterministic fixed sentinel0(AC-9 banstime.perf_counter()/time.time());cost_usd = 0.0(no LLM calls in the rubric).
Files to touch¶
| Path | Why | Status |
|---|---|---|
bench/vuln-remediation/rubric.py |
Replace S5-01 stub body byte-for-byte — module-level score() function + VulnRemediationRubric class (Protocol-conformance delegate) + _HarnessOutput/_ValidatorSignals/_RecipeSignals Pydantic models + _CONDITION_FAILURE_PAIRS + _SEVERITY_FOR_EMITTED_CODE + __main__ subprocess entrypoint |
Modify (S5-01 shipped stub) |
bench/vuln-remediation/tests/conftest.py |
Autouse load_task_class("vuln-remediation", bench_root=REPO_ROOT / "bench") fixture — the sole import bridge for the hyphenated leaf (S2-01 F-CON-8 + PEP 420). Do NOT create bench/vuln-remediation/tests/__init__.py — S5-01 F-CON-5 hard-banned it |
New |
bench/vuln-remediation/tests/test_rubric_unit.py |
Six core-condition tests (AC-4(a–f)) + AC-1 class-delegation + AC-5 semantic-symmetry + AC-6 half-pass + AC-7 sort + AC-8 fail-loud + AC-10 YAML consistency + AC-11 case-invariance | New |
bench/vuln-remediation/tests/test_rubric_static.py |
AC-9 AST-ban on non-determinism inside rubric.py — mirrors S5-01's test_breakdown_keys_static.py sibling pattern |
New |
tests/integration/test_rubric_subprocess_vuln.py |
AC-2 (test_main_exits_nonzero_on_malformed_envelope_json — feeds b"not-json", asserts returncode != 0) + AC-3 (subprocess-CLI test with SCRUBBED_ENV + parent-sentinel protocol; asserts wall-clock ≤ 60 s on representative envelope) |
New |
tests/unit/test_eval_package_imports_no_llm_sdk.py |
AC-12 — extend the existing AST-walk glob to include bench/**/rubric.py; the LLM-SDK ban is structurally identical across src/codegenie/eval/** and bench/**/rubric.py |
Modify (existing per S1-05) |
Out of scope¶
- The runner-side subprocess invocation. S3-03 owns
asyncio.create_subprocess_exec(...)withSCRUBBED_ENVandTemporaryDirectory()cwd; this story honors that contract but does not modify it. - Cases. S5-03 (RAG-corpus-derived) and S5-04 (held-out) land cases that exercise the rubric. The unit tests here use hand-built
BenchCaseobjects. - The
score(...)formula tuning. The story commits to "mechanical againstharness_output" — fine-grained weights are bench-author judgment; do not over-engineer in this story. - Cassette validation. The rubric does not re-verify cassettes;
harness_outputis whatever the SUT emitted, and trust in it is delegated to Phase 4's canary mechanism. - The integration-test wall-clock budget (≤ 60 s). The story asserts the case-level budget at this size; portfolio-scale budgets (≤ 12 min cold cache) are S5-05's concern.
Notes for the implementer¶
- The rubric must work in a stdlib-only subprocess context. No transitive imports of
codegenie.eval.runner, no FS access outside thecwdTemporaryDirectory. Read theBenchCasepaths only if you need to (most rubrics don't — they score fromharness_outputdirectly). - Do not catch
Exceptionand emit a misleading "passed" score. Ifscore(...)fails internally, let the exception propagate; the runner will recordrubric.malformed_output(block-severity) — that is the correct fail-loud behavior. - The
__main__entrypoint must accept the envelope shape the runner produces (S3-03 owns the producer side). The contract is:{"case": <BenchCase JSON>, "harness_output": <whatever the SUT emitted>}. If S3-03's envelope shape differs, that is a contract bug — fix at the harness level, not by adapting the rubric. - Coverage:
pytest --cov=bench.vuln_remediation.rubric --cov-fail-under=90should hit ≥ 90 % line, ≥ 80 % branch. The__main__block is hard to cover in pytest; usesubprocess.runin the integration test to exercise it. tests/unit/test_eval_package_imports_no_llm_sdk.pycurrently walkssrc/codegenie/eval/**/*.py(per ADR-0008 / S1-05). Extend its AST walk tobench/**/rubric.pyin this story — the ban is structurally identical, and rubrics are a logical extension of the no-LLM-SDK package boundary.- The "deterministic" property the audit chain depends on means: no
time.time(), norandom.random(), noos.environreads, nouuid.uuid4(). If the rubric needs a per-case identifier, usecase.case_id. If you find yourself reaching for randomness or wall-clock, you are doing something the rubric should not do. passedderivation: computepassed = (score == 1.0)post-mean, NOTpassed = all(harness_output_conditions). The two are observably equivalent on the AC-1 (full-pass) and AC-6 (half-pass) rows tested here, but they diverge under future extension — e.g., if a fifth breakdown key with a partial-credit scoring rule is added,all(conditions)silently misreportspassed. The mean-then-compare form tiespassedto the invariant the promotion gate consumes (score ≥ 0.95 lower-bound), not to a specific condition-set shape.- Adversarial mutant catalog (this §TDD kills): (1)
score = 0.95hardcoded return → AC-4(a) fails on exact== 1.0; (2) emit info-severityFailureModeon full-pass → AC-4(a) fails onfailure_modes == (); (3) emit spurious failure alongside the paired one → AC-4(b/c) + AC-5 fail on exact-set equality; (4)set(breakdown) <= declaredallowing subset ship → AC-4(d) fails on==; (5)mean→min(0.0) /max(1.0) /sum(2.0) /len(failing)mutants → AC-6 fails onscore == 0.5; (6) unsortedfailure_modestuple → AC-7 fails on lexicographic sort; (7)harness_output.get(k, False)swallowing missing keys → AC-8 fails onpytest.raises; (8)import timeforwall_clock_ms→ AC-9 AST-walk fails; (9) severity drift between hardcoded_SEVERITY_FOR_EMITTED_CODEand YAML → AC-10 consistency test fails; (10) rubric readingcase.difficultyto modify score → AC-11 invariance fails; (11) marker-class replacingVulnRemediationRubric→ AC-1 delegation byte-equality fails. - Contract-surface decision (Phase 6 → Phase 6.5): the
_HarnessOutputenvelope is a LOCAL Pydantic model (bench/vuln-remediation/rubric.py), NOT a shared model insrc/codegenie/eval/models.py. Rule 2 — Phase 7's migration rubric will have a different SUT contract (different signals:dockerfile_migrated,image_digest_matches, etc.). Shared envelopes lock two task classes to the same shape prematurely. The contract between Phase 6'sVulnRemediationSut.run_caseand this rubric is:_HarnessOutputis the wire-shape validator; SUT drift surfaces aspydantic.ValidationErrorat the rubric's__main__boundary →sys.exit(2)→ runner recordsrubric.malformed_output(block).