Subchapter 24.53
references/model-evaluation/overview.mdMarkdown5 KBView on GitHub
sdk-getting-started reference first.This skill supports the evaluation feature for SageMaker Serverless Model Customization. It can evaluate any base or fine-tuned model supported by SageMaker serverless model customization — both OSS models (Llama, Mistral, Qwen, etc.) and Nova models.
Tell the user when the skill is activated:
“I can help evaluate any base or fine-tuned model supported by SageMaker serverless model customization.”
The following are supported by SageMaker and AWS but do not have a validated workflow in this reference. If the user’s request matches one of these, let them know and proceed with best-effort guidance using general AWS knowledge:
There are two evaluation types:
Do you already know which evaluation type to use?
Check conversation history, plan.md, workflow_state.json, or anything else you’ve already read.
If yes: confirm with the user.
“It sounds like you want to run [evaluation type]. Is that right?”
⏸ Wait for confirmation. If confirmed → go to Step 2.
If no: ask.
“What kind of evaluation would you like to run? I support:
- LLM-as-Judge — an LLM grades your model’s responses
- Custom Scorer — programmatic scoring (math, code, or your own logic)
Pick one, or say ‘help me decide’ if you’re not sure.”
⏸ Wait for user.
If user picks one → go to Step 2.
If user indicates uncertainty, by saying something like “help me decide,” “whatever you think,” “I’m not sure” → read references/evaluation-type-guide.md and follow its instructions. It will guide the user to a choice and then return here.
You MUST NEVER make a recommendation to the user on eval type without reading references/evaluation-type-guide.md.
Before reading the reference file, validate that the chosen evaluation type is compatible with the user’s situation. You may already know these answers from conversation context — don’t ask if you don’t need to.
aws sagemaker list-tags --resource-arn <training-job-arn> and look for the sagemaker-studio:jumpstart-model-id tag. Contains “nova” → Nova. Anything else → OSS.aws sagemaker describe-model-package --model-package-name <arn> and check the model description or source tags.If validation fails, tell the user which requirement(s) aren’t met and offer alternatives:
“[Evaluation type] won’t work because [reason].”
If the failure reason was lack of an eval dataset, inform the user that this skill’s evaluation workflows require one, but do not dead-end the conversation.
If the failure reason is something else, offer to help them pick a different evaluation type.
⏸ Wait for user.
If they say they do want help choosing a different eval type → read references/evaluation-type-guide.md.
If validation passes, read the corresponding reference file:
| User chose | Read |
|---|---|
| LLM-as-Judge | references/llmaaj-evaluation.md |
| Custom Scorer | references/custom-scorer-evaluation.md |
Follow the reference file’s instructions from the beginning.