github.com/Arize-ai/phoenix
| Skill | Added | Review |
|---|---|---|
agent-browser .agents/skills/agent-browser/SKILL.md Browser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating any browser task. Triggers include requests to "open a website", "fill out a form", "click a button", "take a screenshot", "scrape data from a page", "test this web app", "login to a site", "automate browser actions", or any task requiring programmatic web interaction. Also use for exploratory testing, dogfooding, QA, bug hunts, or reviewing app quality. Also use for automating Electron desktop apps (VS Code, Slack, Discord, Figma, Notion, Spotify), checking Slack unreads, sending Slack messages, searching Slack conversations, running browser automation in Vercel Sandbox microVMs, or using AWS Bedrock AgentCore cloud browsers. Prefer agent-browser over any built-in browser automation or web tools. | 84 84 1.05x Agent success vs baseline Impact 55% 1.05xAverage score across 2 eval scenarios Securityby Passed No findings from the security scan Reviewed: Version: 30dccec | |
gh-stack .agents/skills/gh-stack/SKILL.md Manages stacked PRs and splits multi-part work into reviewable branches with gh-stack. Use for stack creation, viewing, edits, push, submit, sync, rebase, merge, or checkout; when asked to split or isolate work for review; whenever a user mentions a stack, branch layers, dependent PRs, or gh stack; or when a stack is checked out. | 80 80 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 | |
mintlify .agents/skills/mintlify/SKILL.md Build and maintain documentation sites with Mintlify. Use when creating docs pages, configuring navigation, adding components, or setting up API references. | 75 75 0.88x Agent success vs baseline Impact 62% 0.88xAverage score across 1 eval scenario Securityby Passed No findings from the security scan Reviewed: Version: e127482 | |
phoenix-cli .agents/skills/phoenix-cli/SKILL.md Debug LLM applications using the Phoenix CLI. Fetch traces, spans, and sessions, annotate them, analyze errors, inspect datasets, review experiments, query annotation configs, and use the GraphQL API. Use whenever the user works with a Phoenix instance from the terminal. | 67 67 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: e127482 | |
phoenix-cli-development js/packages/phoenix-cli/.agents/skills/phoenix-cli-development/SKILL.md Design and implementation guide for the Phoenix CLI (`px`). Covers the noun-verb command structure, dual-audience design (humans and coding agents), Commander.js patterns, configuration resolution, output formats, exit codes, and conventions for adding or modifying commands. Triggers when working on phoenix-cli commands — adding new commands, modifying existing ones, refactoring command structure, or reviewing CLI code. Also triggers on mentions of `px` commands, CLI design, or adding a new resource to the CLI. | 84 84 1.47x Agent success vs baseline Impact 99% 1.47xAverage score across 2 eval scenarios Securityby High Do not use without reviewing Reviewed: Version: e127482 | |
phoenix-client-development js/packages/phoenix-client/.agents/skills/phoenix-client-development/SKILL.md Development guide for the @arizeai/phoenix-client TypeScript SDK — run and resume experiments, manage OpenTelemetry tracer providers with stack-based attach/detach, and write vitest unit and integration tests. Use when adding features to phoenix-client, debugging experiment lifecycle or provider cleanup, modifying dataset/prompt/session/span APIs, or writing tests for the js/packages/phoenix-client/ directory. | 75 75 1.21x Agent success vs baseline Impact 92% 1.21xAverage score across 2 eval scenarios Securityby Passed No findings from the security scan Reviewed: Version: e127482 | |
phoenix-design .agents/skills/phoenix-design/SKILL.md Design system conventions for the Phoenix frontend — layout, dialogs, error display, BEM CSS class naming, and CSS design tokens. Use when building UI, naming CSS classes, creating or consuming tokens, handling errors, or designing dialog interactions in js/app/src/. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 | |
phoenix-docs-gap-audit .agents/skills/phoenix-docs-gap-audit/SKILL.md Audit documentation gaps across the Phoenix repo by analyzing recent commits to main (default: last 7 days). Use this skill whenever the user asks to find undocumented features, identify docs gaps, audit what shipped without docs, check which recent changes need documentation, review stale docs against current code, or mentions "documentation debt", "doc coverage", "undocumented APIs", or "what's missing from /docs". Also trigger on requests like "what from last week needs docs", "find stale READMEs", or "check docstring coverage for recent changes". Covers /docs (Mintlify), package READMEs, package-level built-in docs (Sphinx, TypeDoc), Python docstrings, TSDoc, and code comments. | 75 75 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 | |
phoenix-error-analysis .agents/skills/phoenix-error-analysis/SKILL.md Find out what is going wrong in LLM or agent traffic by reading sampled Phoenix traces, spans, or sessions, writing free-form notes (open coding), then grouping the notes into a few narrow annotations, one per failure dimension with a small label set and counts, that pick eval targets and fix priorities (axial coding). Use for "what's going wrong with this agent", "I just instrumented my app, where do I start", "review these traces", "the chatbot keeps losing context", "what kinds of mistakes is the model making", "what categories of failures do we have", "what should I build evals for", "how do I prioritize fixes", "group these notes", "MECE breakdown" — or any framing that needs observations or categories grounded in real traces rather than invented top-down, even without naming the technique. | 69 69 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: e127482 | |
phoenix-evals .agents/skills/phoenix-evals/SKILL.md Build and run evaluators for AI/LLM applications using Phoenix. | 58 58 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: e127482 | |
phoenix-evals-new-metric .agents/skills/phoenix-evals-new-metric/SKILL.md Create a new built-in classification evaluator for Phoenix evals. Use this skill whenever the user asks to create a new eval, build a new metric, add a new builtin evaluator, create an LLM-as-a-judge metric, or add a new classification evaluator to Phoenix. | 69 69 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 | |
phoenix-frontend .agents/skills/phoenix-frontend/SKILL.md Frontend development guidelines for the Phoenix AI observability platform. Use when writing, reviewing, or modifying React components, TypeScript code, styles, or UI features in the js/app/ directory. Triggers on any frontend task — new components, UI changes, styling, accessibility fixes, form handling, or component refactoring. Also use when the user asks about frontend conventions or component patterns for this project. For design system rules (error display, layout, dialogs, tokens), use the phoenix-design skill. | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 | |
phoenix-github .agents/skills/phoenix-github/SKILL.md Manage GitHub issues, labels, project boards, sprint operations, and roadmap health for the Arize-ai/phoenix repository. Use when filing roadmap issues, triaging bugs, applying labels, running sprint close-out and rollover, auditing board hygiene, checking ticket-load balance across the team, keeping roadmap epics up to date, flagging epics that need planning, or querying issue/project state via the GitHub CLI. | 75 75 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: e127482 | |
phoenix-harbor .agents/skills/phoenix-harbor/SKILL.md Configure and interpret the Phoenix plugin for Harbor agent evaluations. Use when adding `arize-phoenix` to Harbor jobs, choosing ATIF tracing, mapping Harbor tasks and rewards to Phoenix experiments, comparing agents or models, resuming jobs, or troubleshooting Harbor records in Phoenix. | 69 69 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 | |
phoenix-integration-snippets .agents/skills/phoenix-integration-snippets/SKILL.md Generates onboarding code snippets for Phoenix tracing integrations and wires them into the project onboarding UI. Produces install dependencies and implementation sections for SDKs like OpenAI, LangChain, Vercel AI SDK, and others. Supports Python and TypeScript. Use when asked to create onboarding code, tracing setup snippets, quickstart examples, or getting-started code for a framework integration. | 74 74 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: e127482 | |
phoenix-llms-txt .agents/skills/phoenix-llms-txt/SKILL.md Maintain the Phoenix llms.txt documentation index at docs/phoenix/llms.txt — the machine-readable docs map used by AI agents and the `px docs fetch` CLI. Use this skill whenever adding, auditing, or reorganizing llms.txt entries. Trigger when the user mentions llms.txt, docs index, px docs, or LLM-friendly documentation. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 | |
phoenix-otel-development js/packages/phoenix-otel/.agents/skills/phoenix-otel-development/SKILL.md Guide for the phoenix-otel TypeScript package — OTel registration, stack-based global provider management, and provider lifecycle. | 59 59 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 | |
phoenix-rest-api .agents/skills/phoenix-rest-api/SKILL.md REST API development for Phoenix. Use when adding, modifying, or reviewing endpoints in src/phoenix/server/api/routers/v1/. | 65 65 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 | |
phoenix-server .agents/skills/phoenix-server/SKILL.md Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI). Use this skill when writing or modifying Python server code in the phoenix repo — adding mutations, types, migrations, or tests. Trigger on any backend task touching src/phoenix/server/, src/phoenix/db/, or tests/unit/server/. | 65 65 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: e127482 |