Omnibus
200 skills · 1230 min
Omnibus
Skill 63 of 200
Investigate the quality of PostHog MCP tool calls — error rates, latency, reach, and which tools are failing or slow.
3 minutes · 717 words · 9 sections
Install
npx skills add PostHog/skills --skill exploring-mcp-tool-qualitynpx skills add PostHog/skills/plugin marketplace add PostHog/skillsThe first command installs just this skill, by the name in its SKILL.md; the second installs the whole repository.
Any MCP server instrumented with PostHog’s MCP analytics SDK emits a
$mcp_tool_call event on the shared events table every time an agent invokes a
tool. There is no dedicated ClickHouse table — every field lives as a
$mcp_* property on events, and every tool-quality metric (error rate, latency
percentiles, reach) is an aggregation over this one event. This is the data
behind the MCP analytics dashboard and tool-quality screens.
For any MCP failure-rate headline, call posthog:metric-list before a typed tool or SQL recipe and look for mcp_tool_call_fail_pct. If it is approved and not drifted, run it with posthog:data-catalog-metric-run and use that result as the canonical headline. When the user also asks which tools drive failures, run the headline first, then use the workflows below for the breakdown and label that breakdown noncanonical. If no governed metric matches, state that the catalog has no match and label the derived rate noncanonical.
For a single tool, prefer the typed tools — posthog:query-mcp-tool-stats (calls,
errors, p50/p95, users, sessions, intents), posthog:query-mcp-tool-failures (top error
messages by harness), and posthog:query-mcp-tool-daily-stats (day-by-day trend). Each
takes a toolName + dateRange, runs the same query runner as the tool-detail
UI, and is gated behind the mcp-analytics flag — no hand-written SQL needed.
HogQL via posthog:execute-sql is the path for cross-tool questions — the
“which tool errors most” ranking below has no typed tool, so rank with SQL, then
drill into the worst tool with posthog:query-mcp-tool-stats and
posthog:query-mcp-tool-failures. The full
property schema and the established query recipes live in the shared MCP data
reference:
products/posthog_ai/skills/querying-posthog-data/references/models-mcp.md (opens in a new tab).
That reference is the single source of truth for the $mcp_* schema and the
effective-tool-name idiom used below — this skill inlines only the noncanonical
“which tool errors most” breakdown for convenience; pull the matrix, latency, and
harness recipes from the reference rather than re-deriving them. Read it before
writing queries.
Always use the effective tool name. New-SDK events wrap the real tool in
a single-exec call, so grouping on raw $mcp_tool_name collapses everything
under the wrapper. Use:
coalesce(nullIf(toString(properties.$mcp_exec_tool_call_name), ''), toString(properties.$mcp_tool_name))Always read $mcp_is_error via toBool(...) and cast
$mcp_duration_ms via toFloat(...). The properties are strings.
Always set a time range — these queries scan events otherwise.
For the “which tool errors most” breakdown, rank tools by error
rate, but guard against small-sample noise with a HAVING floor on call volume:
posthog:execute-sql
SELECT
coalesce(nullIf(toString(properties.$mcp_exec_tool_call_name), ''), toString(properties.$mcp_tool_name)) AS tool,
count() AS total_calls,
countIf(toBool(properties.$mcp_is_error)) AS errors,
round(countIf(toBool(properties.$mcp_is_error)) * 100.0 / count(), 1)
Report both rate and volume — a 100% error rate over 3 calls is rarely the
real story; a 12% rate over 50,000 calls is. Offer to pull the top
$mcp_error_message values for the worst tool (see below).
One row per tool with error rate, latency percentiles, and reach — mirrors the tool-quality screen. The ready-to-run query is in models-mcp.md (opens in a new tab) under “Tool-quality matrix”.
For one tool’s top failure buckets (grouped by harness), call
posthog:query-mcp-tool-failures with the toolName — it’s the typed equivalent of the
query below. Failures come from the same source as the error rate: errored
$mcp_tool_call events ($mcp_is_error), scoped by the effective tool name. Failures are
grouped by $mcp_error_type (a semantic bucket: internal, validation, api_4xx,
api_5xx, permission, timeout, rate_limited, missing_context) and the HTTP
$mcp_error_status when present. To see individual errored calls inside a bucket — with
the captured $mcp_error_message, session id, harness, and intent — pass the bucket’s raw
error_type/error_status to posthog:query-mcp-tool-failure-occurrences
($mcp_error_message is empty on events captured before message capture shipped):
posthog:execute-sql
SELECT
concat(
coalesce(nullIf(toString(properties.$mcp_error_type), ''), 'unknown'),
if(empty(coalesce(toString(properties.$mcp_error_status), '')), '',
concat(' (HTTP ', coalesce(toString(properties.$mcp_error_status),
$mcp_error_type is only populated on newer SDK/server paths — a chunk of errored calls
carry neither type nor status and fall into the unknown bucket.
Swap the aggregate for latency percentiles
(quantile(0.95)(toFloat(properties.$mcp_duration_ms))) and order by p95_ms.
The matrix query already returns p50_ms / p95_ms.
https://app.posthog.com/project/<project_id>/mcp-analytics/dashboardhttps://app.posthog.com/project/<project_id>/mcp-analytics/tool-qualityAlways surface a UI link so the user can verify visually.
HAVING total_calls >= N
floor stops tools with very few calls from topping the list spuriously$mcp_client_name lets you cut quality by harness (Claude Code vs Cursor vs
…); the canonical bucketing multiIf is in
models-mcp.md (opens in a new tab)products/mcp_analytics/backend/mcp_harness.py — that’s the source of truth,
and posthog:query-mcp-harness-breakdown runs it. If your hand-written SQL
disagrees with the screen, your bucketing has drifted from mcp_harness.py;
prefer the typed tool over re-deriving itexploring-mcp-sessions — drill into a
single agent run and its tool sequenceexploring-mcp-intent-clusters —
group agent goals and see which intents drive the errorsInvestigate the quality of PostHog MCP tool calls — error rates, latency, reach, and which tools are failing or slow. Use when the user asks "which MCP tool has the highest error rate?", "what's the slowest tool?", "which tools fail most often?", "how reliable is tool X?", wants a tool-quality matrix, or pastes an MCP analytics tool-quality / dashboard URL and asks what it shows.
The verbatim description from this skill’s front matter — the string an agent matches on to decide whether to load it.
main, last pushed 24 September 2026.SKILL.md, not by matching a directory convention. 2 distinct layouts observed: skills/omnibus/*/SKILL.md, skills/posthog/all/skills/*/SKILL.md.h1 and no skipped levels:.claude-plugin/marketplace.json by PostHog, declaring 6 plugins. It is read for editorial metadata only — never as the skill index, which is always the repository tree./PostHog/skills.md, and each skill at its own .md URL.