LLM Finetuning
Skill 111 of 183
Fine-tune vision-language models (VLMs) with supervised learning on image+text data.
5 minutes · 1,029 words · 9 sections
Install
npx skills add wshobson/agents --skill vision-sftnpx skills add wshobson/agents/plugin marketplace add wshobson/agentsThe first command installs just this skill, by the name in its SKILL.md; the second installs the whole repository.
This skill assumes finetuning-method-selection
already routed here: the data shape is
image+text demonstrations, not preference pairs
or a verifiable reward signal, and the base is a
vision-language model rather than a text-only
one. lora-qlora-recipes covers the text-only
LoRA/QLoRA recipe this skill specializes for the
vision tower and projector; read that skill first
if the LoRA fundamentals (rank, alpha, target
modules) aren’t already familiar.
Input: an image+text dataset and a VLM base
model already picked from the model catalog.
Output format: a validated adapter config —
which components are frozen, LoRA target modules,
and a min_pixels/max_pixels budget — that
llm-finetuning-training-engineer consumes
directly when it generates a runnable script.
| Situation | Default |
|---|---|
| Adapting behavior on familiar images | Frozen tower+projector, LoRA r=8–16, α=16–32 |
| Visual domain shift | Unfreeze last-6 ViT layers, vision LR 5–10x lower |
| Doesn’t fit in bf16 at target rank | QLoRA — frozen vision tower only |
fast_inference=True | finetune_vision_layers=False |
| Loss normal, eval not improving | Check the Two Silent Killers below first |
Freeze the vision tower and the projector. Put
LoRA on the LLM only, all-linear (the same
attention + MLP target list as text-only SFT —
see lora-qlora-recipes), at r=8–16,
α=16–32. This is the settled default for
adapting a VLM’s behavior without disturbing how
it sees.
# freeze tower + projector; LoRA on LLM only
for name, param in model.named_parameters():
if "vision_tower" in name or "projector" in name:
param.requires_grad = False
target_modules = [
"q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj",
] # LLM-only, all-linear — r=8-16, alpha=16-32Unfreezing vision layers is a deliberate escalation, not a default decision — reach for it only when the domain shift is visual, not textual.
Both produce a run that trains without error and without learning: the loss curve looks normal, the model doesn’t improve, and neither throws an exception — both need an explicit pre-training check, not just a clean training log.
references/collators-and-pitfalls.md.min_pixels/max_pixels resolution budget.
This pair is the single most consequential
hyperparameter for quality and memory in VLM
SFT — more than rank, alpha, or LR. Too low
silently downsamples images below what the task
needs (small document text becomes unreadable
even though training “succeeds”); too high blows
the activation memory budget or forces too small
a batch to train stably. Set it deliberately per
dataset, don’t leave it at a framework default.UnslothVisionDataCollator is the collator
Unsloth expects for VLM SFT — it handles the
image-tag alignment and per-architecture
processor contract described in
references/collators-and-pitfalls.md. Don’t
substitute a text-only collator for VLM data.finetune_vision_layers=False is required
when fast_inference=True. vLLM cannot serve
LoRA adapters on vision layers, so a fast-
inference setup that also unfreezes vision
layers fails at serve time even if training
succeeds. If the recipe calls for unfreezing the
last-6 ViT layers (see When to Unfreeze above),
fast inference is off the table for that run —
choose one or the other, not both.Base VLM choice is out of scope for this skill —
it lives in one place, the model catalog at
finetuning-method-selection‘s
references/model-catalog.md. This skill and its
references describe recipes by architecture
family only, never by recommending one model over
another.
VLM reinforcement learning (VLM-GRPO) is
reference-only in this plugin — the fragmented
tooling and reward-hacking failure modes specific
to VLM-RL are covered in grpo-rlvr-training,
not here. This skill’s scope stops at supervised
fine-tuning.
The recurring mistake across every section above
is treating a clean loss curve as proof the run
is healthy. A normal-looking curve is consistent
with both a working run and either silent
killer, since the model trains on something
either way — just not the aligned image-text
signal when a killer is present. A flat eval score
next to a normal loss curve means re-run the
checklist in references/collators-and-pitfalls.md
before touching any hyperparameter.
references/collators-and-pitfalls.md — per-
architecture collator table, dataset-format
examples with image placeholders, a pre-
training validation checklist, and the two-
stage projector-alignment recipe as an advanced
pattern.Related skills: finetuning-method-selection
routes here; lora-qlora-recipes covers the
text-only LoRA fundamentals this skill
specializes; grpo-rlvr-training covers VLM-RL
(reference-only); dataset-curation covers
image+text dataset preparation this skill doesn’t.
Fine-tune vision-language models (VLMs) with supervised learning on image+text data. Use when adapting a VLM to a visual domain or task, configuring frozen-vision-tower LoRA, or debugging a VLM fine-tune that trains without learning.
The verbatim description from this skill’s front matter — the string an agent matches on to decide whether to load it.
main, last pushed 21 September 2026.SKILL.md, not by matching a directory convention. 51 distinct layouts observed: plugins/accessibility-compliance/skills/*/SKILL.md, plugins/agent-teams/skills/*/SKILL.md, plugins/api-scaffolding/skills/*/SKILL.md, plugins/avoid-ai-writing/skills/*/SKILL.md, plugins/backend-development/skills/*/SKILL.md, plugins/before-you-build/skills/*/SKILL.md, plugins/block-no-verify/skills/*/SKILL.md, plugins/blockchain-web3/skills/*/SKILL.md, plugins/brand-landingpage/skills/*/SKILL.md, plugins/business-analytics/skills/*/SKILL.md, plugins/cicd-automation/skills/*/SKILL.md, plugins/cloud-infrastructure/skills/*/SKILL.md, plugins/conductor/skills/*/SKILL.md, plugins/data-engineering/skills/*/SKILL.md, plugins/database-design/skills/*/SKILL.md, plugins/developer-essentials/skills/*/SKILL.md, plugins/dgx-spark-ops/skills/*/SKILL.md, plugins/documentation-generation/skills/*/SKILL.md, plugins/documentation-standards/skills/*/SKILL.md, plugins/dotnet-contribution/skills/*/SKILL.md, plugins/file-conversion/skills/*/SKILL.md, plugins/framework-migration/skills/*/SKILL.md, plugins/frontend-mobile-development/skills/*/SKILL.md, plugins/game-development/skills/*/SKILL.md, plugins/hermes-tweet/skills/*/SKILL.md, plugins/hr-legal-compliance/skills/*/SKILL.md, plugins/incident-response/skills/*/SKILL.md, plugins/javascript-typescript/skills/*/SKILL.md, plugins/kubernetes-operations/skills/*/SKILL.md, plugins/llm-application-dev/skills/*/SKILL.md, plugins/llm-finetuning/skills/*/SKILL.md, plugins/machine-learning-ops/skills/*/SKILL.md, plugins/observability-monitoring/skills/*/SKILL.md, plugins/payment-processing/skills/*/SKILL.md, plugins/plugin-eval/skills/*/SKILL.md, plugins/pptx-deck-creation/skills/*/SKILL.md, plugins/protect-mcp/skills/*/SKILL.md, plugins/python-development/skills/*/SKILL.md, plugins/quantitative-trading/skills/*/SKILL.md, plugins/reverse-engineering/skills/*/SKILL.md, plugins/review-agent-governance/skills/*/SKILL.md, plugins/security-scanning/skills/*/SKILL.md, plugins/shell-scripting/skills/*/SKILL.md, plugins/ship-mate/skills/*/SKILL.md, plugins/signed-audit-trails/skills/*/SKILL.md, plugins/skill-forge-essentials/skills/*/SKILL.md, plugins/social-publishing/skills/*/SKILL.md, plugins/startup-business-analyst/skills/*/SKILL.md, plugins/superself/skills/*/SKILL.md, plugins/systems-programming/skills/*/SKILL.md, plugins/ui-design/skills/*/SKILL.md.h1 and no skipped levels:.claude-plugin/marketplace.json by Seth Hobson, declaring 94 plugins. It is read for editorial metadata only — never as the skill index, which is always the repository tree./wshobson/agents.md, and each skill at its own .md URL.1 file · 6 KB
Everything this skill ships beside its prose. All of it is set here, as a subchapter of skill 111.
Documentation the agent loads on demand, rather than up front.