43 skills · 279 min
Skills
Skill 18 of 43
Builds and maintains configuration-based evaluations on a workflow with the eval-config tool.
3 minutes · 757 words · 7 sections
Install
npx skills add n8n-io/n8n --skill config-evalsnpx skills add n8n-io/n8n/plugin marketplace add n8n-io/n8nThe first command installs just this skill, by the name in its SKILL.md; the second installs the whole repository.
Use this skill to attach a configuration-based evaluation to a workflow with the
eval-config tool. A config eval pairs a workflow with a name, a start node, an
end node, one or more judged metrics, and a Data Table dataset. Nothing is added
to the canvas — the config lives off-canvas via the evaluation-config API.
Config evals are the only evaluation form you work with. Do not add, read, rewire, or reason about on-canvas evaluation nodes (EvaluationTrigger, Evaluation/checkIfEvaluating/setOutputs/setMetrics). If the user asks for those, build a config eval instead and briefly say that is how you set up evaluations.
name — a human-readable evaluation name.startNodeName — the node where a run begins; it is fed one test-input row.
Must be a node with an incoming connection — never a trigger (see step 2).endNodeName — the node whose output is judged.dataTableId — a Data Table holding the test dataset. Create and populate it
with the data-tables tool first, then link it here by id.metrics — one or more judged metrics (see below).startNodeName is the first node after the trigger — the node that
receives the input the dataset varies. Never use the trigger itself: an
eval run swaps the trigger for a dataset-driven one, so the start node must
have an incoming connection or the run fails to compile. For a chat/agent
workflow this is usually the agent node (often the same as endNodeName).endNodeName is the node whose output you want scored (usually the AI agent
or the final response node).data-tables(action="list") to find an existing
dataset, or create and seed one with data-tables before creating the config.
Never invent a dataTableId; use one returned by data-tables.actualAnswer / expectedAnswer / userQuery
expressions (see Metrics).eval-config (action="create"), or update when changing an existing
config. The tool shows an approval card automatically — call it and respect
the result; do not ask for chat approval first.Each metric is LLM-judged and needs a judge model: a credentialId, a model,
and an outputType (numeric, the default, or boolean). Reuse an LLM
credential the workflow already uses when one fits.
Do not set provider unless you know the exact chat-model node type — it is
derived automatically from the credential you pass (each credential type maps to
one provider). Just pick the credential and the model.
Two presets are available:
correctness — compares the produced answer to a ground-truth answer.
Requires expectedAnswer (an n8n expression resolving to the ground-truth
value, typically a dataset column, e.g. ={{ $json.expected_output }}).helpfulness — judges the produced answer against the user’s query.
Requires userQuery (an n8n expression for the input the user asked, e.g.
={{ $json.input }}).Every metric also needs actualAnswer: an n8n expression resolving to the
workflow’s produced answer at the end node, e.g. ={{ $json.output }}.
userQuery and expectedAnswer name dataset columns (the input the user
asked; the ground-truth answer). actualAnswer names a field of the workflow’s
produced output. Write all of them as ={{ $json.<name> }} — the evaluation
reads dataset columns from the dataset row and actualAnswer from the end node
automatically. Do not reference the trigger or any node by name.
=actualAnswer, userQuery, and expectedAnswer are n8n expressions — they
read a value out of each test row at runtime. The leading = is what tells n8n
to evaluate the {{ … }} template. Without it the string is stored as literal
text: the field shows {{ $json.output }} verbatim and the judge scores that
raw string instead of the resolved value.
={{ $json.output }}, ={{ $json.expected_output }}{{ $json.output }} (no = → treated as fixed text)Only add = when the value references workflow data via {{ … }}. A genuinely
fixed constant (rare for these fields) is written as plain text without =.
Pick correctness when the dataset has a known right answer to compare against;
pick helpfulness when there is no single ground truth and quality is judged
relative to the request. Use prompt only to override the default judge prompt.
data-tables tool: one column for each input the
evaluation varies, plus a ground-truth column when using correctness.dataTableId; the eval-config tool
does not create or populate rows. If no suitable dataset exists, create one
first, then create the config.Use references/config-eval-playbook.md (opens in a new tab) for tool-call recipes, worked examples, and output shapes.
Builds and maintains configuration-based evaluations on a workflow with the eval-config tool. Use when the user asks to set up, add, view, change, or remove an evaluation, score, grade, or judge a workflow's output, or measure answer quality against a test dataset. This is the only eval form Instance AI handles — it does not touch on-canvas evaluation nodes.
The verbatim description from this skill’s front matter — the string an agent matches on to decide whether to load it.
master, last pushed 24 September 2026.SKILL.md, not by matching a directory convention. 5 distinct layouts observed: .agents/skills/*/SKILL.md, .claude/plugins/n8n/skills/*/SKILL.md, .opencode/skills/*/SKILL.md, packages/@n8n/cli/skills/*/SKILL.md, packages/@n8n/instance-ai/skills/*/SKILL.md.h1 and no skipped levels:.claude/plugins/n8n/.claude-plugin/marketplace.json by n8n, declaring 1 plugin. It is read for editorial metadata only — never as the skill index, which is always the repository tree./n8n-io/n8n.md, and each skill at its own .md URL.1 file · 5 KB
Everything this skill ships beside its prose. All of it is set here, as a subchapter of skill 18.
Documentation the agent loads on demand, rather than up front.