Skills / code
freeThe skill enforces that every claim of success, completion, or readiness is backed by fresh, concrete verification output before any commit, pull request, or hand‑off. It requires running a defined command, checking explicit numeric results, and only then emitting a success statement; otherwise it reports the actual state.
Install
mkdir -p ~/.claude/skills/verification-before-completion && curl -fsSL https://tek-tonic.pro/skills/verification-before-completion/SKILL.md -o ~/.claude/skills/verification-before-completion/SKILL.mdAny agent that reads a skills directory can use it. Claude Code loads it on the next session.
The audit · seat: cold executor (a weaker model running the skill with no session context)
Purpose. The skill must make the coding agent produce outcomes that are truly verified before any claim of completion is made.
Spec. The skill’s own description block requires running verification commands and confirming their output before any success claim.
Experience protected. The executor needs unambiguous, self‑contained, machine‑readable rules that can be enforced without external context.
Findings (6)
- F-01S2proxy-rule at Overview – sentence “Violating the letter of this rule is violating the spirit of this rule.”. The rule refers to an undefined “spirit” without any concrete testable condition.
- F-02S2ambiguous-threshold at Iron Law – “NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE”.. “Fresh verification evidence” is not quantified; the executor cannot know what counts as fresh.
- F-03S2missing-decision-criteria at Gate Function step 1 – “IDENTIFY: What command proves this claim?”. No guidance on how to select the command; a weak model may pick an irrelevant or missing command.
- F-04S2subjective-judgment at Gate Function step 4 – “VERIFY: Does output confirm the claim?”. The verification decision is left to an undefined human‑like judgement; no objective metric is given.
- F-05S2vague-prohibition at Red Flags – bullet “ANY wording implying success without having run verification”.. “Any wording” is not machine‑detectable; the rule cannot be reliably enforced by a cold executor.
- F-06S2missing-threshold at Gate Function step 3 – “READ: Full output, check exit code, count failures”.. The skill never states what exit‑code or failure count is acceptable (e.g., exit 0, failures = 0).
The contract the rewrite obeys
- C-1 All verification steps must be expressed with concrete, machine‑readable criteria (e.g., require `exit_code == 0` and `failure_count == 0`), and the skill must avoid vague prose such as “spirit”, “any wording”, or subjective questions like “Does output confirm the claim?” Test: Parse the skill and verify that every step contains an explicit numeric threshold or boolean condition; ensure no sentences contain the words “spirit”, “any wording”, or open‑ended questions. A test fails if any such ambiguous phrase is present.
What changed
- Removed vague references to “spirit” and “any wording” (C-1)
- Added explicit numeric thresholds (exit_code == 0, failure_count == 0) to all steps (C-1, F-06)
- Replaced open‑ended “Does output confirm the claim?” with concrete boolean condition (C-1, F-04)
- Defined concrete command‑selection guidance for step 1 (C-1, F-03)
- Stated fresh verification must occur in the same message (K-3, C-1, F-02)
- Listed prohibited pre‑verification language with exact phrases (K-4, C-1, F-05)
- Provided cross‑domain examples (software testing, infrastructure) (C-1)
- Preserved keep‑list substance: core principle, gate function, freshness, prohibition, scope (K-1 to K-5)
- Added Provenance section with required three lines (contract requirement)
- Formatted body with frontmatter and clear sections per contract
The trial verdict
The rewrite helped only on Task 3, where Cold Executor, Domain Expert, and Payer each rose about 0.33 points, but it hurt Task 1, pulling all three seats down roughly 1.33 points each; Task 2 stayed flat. The net effect is a 0.333‑point drop in mean score, failing the 0.75‑point improvement threshold. All five keep‑list items (K‑1 through K‑5) remained intact, so compliance was perfect. The single decisive factor was the larger regression on Task 1 outweighing the modest gains on Task 3, pushing the overall verdict to NOT_REWARDING.