Before You Build
Block No Verify
Brand Landingpage
Business Analytics
Database Design
Documentation Generation
Documentation Standards
File Conversion
Framework Migration
Frontend Mobile Development
Game Development
Hermes Tweet
Hr Legal Compliance
Incident Response
Kubernetes Operations
Machine Learning Ops
NET Contribution
Observability Monitoring
Payment Processing
Plugin Eval
Protect MCP
Python Development · Python…
Quantitative Trading
Review Agent Governance
Ship Mate
Signed Audit Trails
Skill Forge Essentials
Social Publishing
Systems Programming
180 chapters · 328 min
LLM Finetuning
Chapter 105 of 180
Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters.
4 minutes · 938 words · 9 sections
This skill assumes the routing decision already
happened — finetuning-method-selection should
have already pointed here because the data shape
is demonstrations (SFT), not preference pairs or
a verifiable reward signal. What follows is the
current best-practice recipe for configuring the
adapter itself: which modules to target, how to
size rank and alpha, what learning rate to use,
and when QLoRA buys real headroom versus when it
just adds risk. Dataset preparation and quality
checks are a separate concern — see
dataset-curation.
Input: a routing decision (SFT via LoRA/
QLoRA) plus a target size class.
Output format: a validated adapter config —
the kwarg values below, not free-form advice —
that llm-finetuning-training-engineer consumes
directly when it generates a runnable script.
The reference recipe is “LoRA Without Regret” (Thinking Machines/Schulman, 2025-09), now the settled convention for LoRA/QLoRA SFT.
Target all-linear modules, not just attention:
target_modules = [
"q_proj", "k_proj", "v_proj", "o_proj", # attention
"gate_proj", "up_proj", "down_proj", # MLP — matters most
]The MLP layers (gate_proj, up_proj,
down_proj) matter most — attention-only
targeting was the older, weaker convention.
Dropping modules to save memory is a Failure
Mode below, not a valid optimization.
lora_alpha = 2 * r is the settled
convention (NeurIPS 2025 “intruder dimensions”
result). Don’t hand-tune alpha independently of
rank — derive it from rank every time.references/hyperparameters.md.Rank is task-shaped, not a single global default:
| Task | Rank |
|---|---|
| RL (GRPO/RLVR adapters) | 1–32 |
| General default | 16–32 |
| SFT at scale | up to ~256 |
Higher rank isn’t automatically better — it raises capacity to memorize as fast as it raises capacity to generalize. Start at the row matching the task, and only move up a row if the lower rank measurably underfits on held-out eval, not as a default hedge.
Keep effective batch size under 32. This recipe was validated at that scale — pushing effective batch higher is an untested extrapolation, not a free throughput win.
Unsloth is the reference implementation this
plugin assumes as the default fast path — except
for messages-shaped conversational SFT with
assistant_only_loss=True, where Unsloth
2026.7.x’s compiled trainer has no messages-shaped
path at all and the plain-TRL escape hatch
(references/unsloth-trl-mapping.md) is the
default for that combination, not a rare-regression
fallback. Its out-of-the-box defaults, and why
each one is set that way:
lora_dropout=0 — the optimized kernel
path assumes zero dropout; setting a nonzero
value forfeits the fused-kernel speedup.bias="none" — bias terms add adapter
parameters for negligible quality gain at this
rank range.use_gradient_checkpointing="unsloth" —
Unsloth’s checkpointing variant, not vanilla HF
checkpointing; saves roughly 30% VRAM over
no checkpointing.optim="adamw_8bit" — 8-bit AdamW cuts
optimizer-state memory with negligible quality
impact at LoRA/QLoRA adapter scale.random_state fixed — pins LoRA
initialization for reproducibility across runs;
treat it like any other seed, not a tunable.These show up together on the get_peft_model
call:
model = FastLanguageModel.get_peft_model(
model,
r=32,
target_modules=target_modules,
lora_alpha=64, # 2 * r
lora_dropout=0,
bias="none",
use_gradient_checkpointing="unsloth",
random_state=3407,
)Exact kwarg names and their plain-TRL/PEFT
equivalents, plus a full worked config including
SFTConfig: references/unsloth-trl-mapping.md
and references/hyperparameters.md.
| Situation | Default choice |
|---|---|
| Adapting behavior on demonstrations | LoRA |
| Base model doesn’t fit in bf16 at target rank | QLoRA |
| Injecting dense new domain knowledge | Full FT (see finetuning-method-selection) |
| Unsure which one | LoRA — upgrade to QLoRA only if memory forces it |
dgx-spark-ops plugin’s
spark-memory-thermal-ops skill covers the
full OOM remediation ladder (bf16 LoRA is the
next thing to try, not a further QLoRA
shrink).fp16 divergence on non-BF16 GPUs. Training
in fp16 on hardware that doesn’t have solid
BF16 support is a known source of loss spikes
and silent divergence. Force bf16=True
wherever the hardware supports it; don’t fall
back to fp16 as if it were equivalent. Check
hardware support before picking a dtype:
python -c "import torch; print(torch.cuda.is_bf16_supported())"Rank too high on a small dataset overfits. A rank picked for “SFT at scale” (up to ~256) on a dataset that doesn’t have scale behind it memorizes rather than generalizes. Match rank to the Rank by Task table above, not to the largest number available.
Removing target modules to save memory costs
quality for negligible savings. The adapter
parameters on gate_proj/up_proj/down_proj
are a small fraction of total model size — cutting
them barely moves memory but measurably hurts
quality. If memory is tight, move to QLoRA or
reduce rank/batch/pack length before trimming
target modules.
All three failure modes share a pattern: they look like a training-loop bug (loss spikes, plateaus, memorization) but are actually a config choice that contradicts the reference recipe above. Check configuration against this skill before debugging the training loop itself.
references/hyperparameters.md — full rank/
alpha/LR tables by task type, rsLoRA notes,
batch/packing interactions, and a complete
worked Unsloth config block.references/unsloth-trl-mapping.md — every
Unsloth kwarg mapped to its TRL/PEFT
equivalent, current TRL API notes, and the
escape-hatch rule for when to drop back to
plain TRL.Related skills: finetuning-method-selection
routes here; dataset-curation covers the data
side this skill doesn’t; llm-finetuning-training-engineer
is the downstream consumer of the config this
skill produces.
Install this repository
npx skills add wshobson/agents/plugin marketplace add wshobson/agentsSkills install per repository, not per chapter — the CLI has no documented per-skill form, so we do not print one.
Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters. Use when writing or reviewing a LoRA/QLoRA training configuration, choosing rank/alpha/target modules, or deciding between LoRA, QLoRA, and full fine-tuning.
The verbatim description from this skill’s front matter — the string an agent matches on to decide whether to load it.
main, last pushed 5 August 2026.SKILL.md, not by matching a directory convention. 49 distinct layouts observed: plugins/accessibility-compliance/skills/*/SKILL.md, plugins/agent-teams/skills/*/SKILL.md, plugins/api-scaffolding/skills/*/SKILL.md, plugins/backend-development/skills/*/SKILL.md, plugins/before-you-build/skills/*/SKILL.md, plugins/block-no-verify/skills/*/SKILL.md, plugins/blockchain-web3/skills/*/SKILL.md, plugins/brand-landingpage/skills/*/SKILL.md, plugins/business-analytics/skills/*/SKILL.md, plugins/cicd-automation/skills/*/SKILL.md, plugins/cloud-infrastructure/skills/*/SKILL.md, plugins/conductor/skills/*/SKILL.md, plugins/data-engineering/skills/*/SKILL.md, plugins/database-design/skills/*/SKILL.md, plugins/developer-essentials/skills/*/SKILL.md, plugins/dgx-spark-ops/skills/*/SKILL.md, plugins/documentation-generation/skills/*/SKILL.md.plugins/documentation-standards/skills/*/SKILL.mdplugins/dotnet-contribution/skills/*/SKILL.mdplugins/file-conversion/skills/*/SKILL.mdplugins/framework-migration/skills/*/SKILL.mdplugins/frontend-mobile-development/skills/*/SKILL.mdplugins/game-development/skills/*/SKILL.mdplugins/hermes-tweet/skills/*/SKILL.mdplugins/hr-legal-compliance/skills/*/SKILL.mdplugins/incident-response/skills/*/SKILL.mdplugins/javascript-typescript/skills/*/SKILL.mdplugins/kubernetes-operations/skills/*/SKILL.mdplugins/llm-application-dev/skills/*/SKILL.mdplugins/llm-finetuning/skills/*/SKILL.mdplugins/machine-learning-ops/skills/*/SKILL.mdplugins/observability-monitoring/skills/*/SKILL.mdplugins/payment-processing/skills/*/SKILL.mdplugins/plugin-eval/skills/*/SKILL.mdplugins/pptx-deck-creation/skills/*/SKILL.mdplugins/protect-mcp/skills/*/SKILL.mdplugins/python-development/skills/*/SKILL.mdplugins/quantitative-trading/skills/*/SKILL.mdplugins/reverse-engineering/skills/*/SKILL.mdplugins/review-agent-governance/skills/*/SKILL.mdplugins/security-scanning/skills/*/SKILL.mdplugins/shell-scripting/skills/*/SKILL.mdplugins/ship-mate/skills/*/SKILL.mdplugins/signed-audit-trails/skills/*/SKILL.mdplugins/skill-forge-essentials/skills/*/SKILL.mdplugins/social-publishing/skills/*/SKILL.mdplugins/startup-business-analyst/skills/*/SKILL.mdplugins/systems-programming/skills/*/SKILL.mdplugins/ui-design/skills/*/SKILL.mdh1 and no skipped levels:.claude-plugin/marketplace.json by Seth Hobson, declaring 95 plugins. It is read for editorial metadata only — never as the skill index, which is always the repository tree./wshobson/agents.md, and each chapter at its own .md URL.2 files · 15 KB
Everything this skill ships beside its prose. All of it is set here, as subchapters of chapter 105.
Documentation the agent loads on demand, rather than up front.