Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.
The Phase 5 gate for the whole
plugin: a checkpoint that trains
cleanly and beats its task metric
still doesn't ship without
clearing all four stages below.
eval-harness-first built the
suite re-run here — this skill is
where that suite's baseline
decides something.
Input: a trained checkpoint,
eval/baseline-<model>.json from
eval-harness-first, and the
frozen eval/drift-suite.yaml.
Output format:
promotion-report.md — the
four-stage evidence plus a
terminal PROMOTE or REJECT
verdict that /finetune Phase 5
and /promote-checkpoint consume
directly.
Each stage gates the next — a failure at stage 2 means stage 3 doesn't run. Stages 2 and 3 share one expensive inference pass, so running them concurrently and applying gate order at verdict time is licensed on a deterministic arena (nothing saved by serializing); a judge-based arena should still wait for stage 2 first — that's where the real savings are.
trace-to-training-data's
Hygiene section exists to
prevent), and scan for label
noise. A checkpoint trained on
leaked goldens invalidates
every later stage.eval-harness-first's
eval/drift-suite.yaml —
MMLU/GSM8K/IFEval plus 200–500
domain-adjacent items — against
the checkpoint and diff against
baseline-<model>.json per
benchmark against the Drift
Budget table below.references/gate-templates.md
when every grader in the
harness is deterministic (no
LLM-judge; position
randomization N/A there).
A holdout win that
loses the live arena does not
ship — stage-2 numbers and
stage-3 judgments must agree; a
win on frozen goldens and a
loss in paired comparison is a
real signal, not a discrepancy
to explain away.| Drift (pts) | Verdict | |---|---| | ≤1 | Noise — proceed | | 2–5 | Rerun with seed variation before deciding | | >5 | HARD FAIL — no exception for task gains |
The >5pt row governs regardless of the others: a checkpoint that gained 8 points on the target task and lost 6 points of general capability still fails here — task improvement never buys back a drift-budget breach.
Item count derives from the
budget, not convenience: the
strict n for a half-width under
half the 5pt hard-fail threshold
is ~1,300 at typical accuracy
(p≈0.7); n=200 is a pragmatic
floor (±6pt half-width at that
same p, n=50 ±13pt) — report the
half-width with every verdict,
and treat a margin smaller than
it as REJECT (uncertain), not
PASS/HARD FAIL. Full math and a
5-run cautionary example:
references/gate-templates.md.
RERUN is not a verdict. A
2–5pt drift only ever produces a
PROMOTE or REJECT after the
seed-variation rerun completes —
PROMOTE requires landing back
at ≤1pt (noise); any rerun still
1pt — 2–5pt band or >5pt breach alike — resolves stage 2 to a hard
REJECT. No report may reach the Verdict section with stage 2 still showingRERUN.
Unmanaged LoRA fine-tuning loses real general capability, and stage 2 is what catches it:
If a checkpoint hits the >5pt
hard fail in stage 2, work this
escalation ladder in order — the
one canonical order this skill
and references/gate-templates.md
both point to:
lora-qlora-recipes and
preference-optimization tune
for the training run, applied
here in reverse.This order is a default, not a
law: remediation guidance from
a single before/after run pair
is a hypothesis — label it
low-confidence once any lever
produces a reversal, and prefer
a seed-variation repeat over
trusting the next rung blindly.
A lever that clears the drift
breach but drops a
success-criterion metric below
target is a two-sided tradeoff
for a human, not a reason to
keep descending the ladder. Full
reasoning and the 5-run
trajectory behind both caveats:
references/gate-templates.md.
Disclose drift-suite instruction reuse. A replay row copying the drift harness's exact instruction phrasing (not just disjoint source items) makes that benchmark's post-replay score an upper bound — flag it instruction-familiar, or re-probe with a paraphrase, before treating a near-budget pass as clean.
promotion-report.md covers all
four stages as sections and
must end with a terminal
verdict: PROMOTE or REJECT,
the evidence that produced it,
and exactly one top remediation
when the verdict is REJECT.
Template: references/gate-templates.md.
The terminal contract other
skills parse:
## Verdict
REJECT
Evidence: domain-adjacent drift
suite dropped 6.2pt (threshold:
>5pt hard fail) despite +8pt on
the target task.
Top remediation: swap the
replay-mix fraction from 10%
toward 20%, holding step count
constant.
REJECT hands
the remediation back to a human
decision at
finetuning-method-selection or
the relevant training skill.eval-harness-first — owns the
drift suite and baseline this
skill re-runs and diffs
against; no baseline-<model>.json
means nothing to gate against.quantized-export — the only
valid next step after a
PROMOTE verdict.preference-optimization and
lora-qlora-recipes — own the
LR and rank levers in the
Catastrophic Forgetting
escalation path; this skill
diagnoses the breach, those
skills own the config that
caused it.dataset-curation — owns the
replay-mix construction recipe
the escalation ladder's first
rung applies.Complete promotion-report.md
template with all four stages,
the drift-suite scoring table,
the paired-arena protocol (item
count, position randomization,
win-rate threshold), and a
replay-mix configuration example:
references/gate-templates.md.
Copy a source-pinned command for your client. You run it yourself.
Destination: .claude/skills/checkpoint-promotion · pinned to the source commit
# Run from your project root
git clone https://github.com/wshobson/agents.git .skillboard-tmp
git -C .skillboard-tmp checkout 38e19c20d2b154510b0e624a2e3e186b19b5c527
mkdir -p ".claude/skills"
cp -r ".skillboard-tmp/plugins/llm-finetuning/skills/checkpoint-promotion" ".claude/skills/"
rm -rf .skillboard-tmpReview the source before running. This copies files into your project; it is not a one-click install and does not verify runtime safety.
sudo apt update && sudo apt install -y gitnpm install -g @anthropic-ai/claude-code# Run from your project root
git clone https://github.com/wshobson/agents.git .skillboard-tmp
git -C .skillboard-tmp checkout 38e19c20d2b154510b0e624a2e3e186b19b5c527
mkdir -p ".claude/skills"
cp -r ".skillboard-tmp/plugins/llm-finetuning/skills/checkpoint-promotion" ".claude/skills/"
rm -rf .skillboard-tmpDestination: .claude/skills/checkpoint-promotion
Scanner static-checks@0.1.0 · commit 38e19c20d2b1. Static checks cannot prove runtime safety – review the source and the exact diff before installing. How checks work.
Uses encoding/eval patterns that can hide executable content.
Evidence: [redacted]· fingerprint 1bbd174404efbce9