CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-codex-plugin

Install, set up, verify, update, pin, uninstall, or troubleshoot the ToolUniverse plugin on OpenAI Codex. ALWAYS consult this skill for any of those — don't answer from memory, because the exact marketplace name (mims-harvard/ToolUniverse), the "codex plugin marketplace add" then "codex plugin add -m tooluniverse" flow, Codex's startup auto-upgrade behavior, the uvx tooluniverse MCP server, and the API-key env vars are easy to get wrong. Use it whenever someone wants to get ToolUniverse (or "the 1000+ scientific tools" / "the harvard tools") working on Codex, says the Codex plugin or its tools/skills won't load, hits a uvx or MCP-server startup error, asks how Codex updates it, wants to pin or remove it, or finds it running an old tool version — even if they never say the word "plugin". Not for the Claude Code plugin (use tooluniverse-claude-code-plugin), for running research with the tools, or for authoring new tools or skills.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable operational guide: every workflow is concrete and executable, the install/verify/update/uninstall lifecycle has explicit validation checkpoints, and troubleshooting includes an ordered, idempotent diagnosis script with feedback loops. The main costs are token redundancy between the troubleshooting table and diagnosis script, and an inline maintainer section that serves a non-user audience.

Suggestions

Move the "For plugin maintainers" section to a separate reference file (e.g., MAINTAINERS.md) or the repo docs, keeping SKILL.md focused on the user-facing flow and cutting roughly 15 lines of tokens.

Deduplicate troubleshooting: the uv-shadowing, cache-clean, and marketplace-upgrade fixes each appear in both the Troubleshooting table and the Automated diagnosis section — merge them into one canonical checklist.

Make the API key reference resolvable within this skill's own bundle (e.g., copy or link API_KEYS_REFERENCE.md under references/) so the one external pointer in the body is verifiable where the skill is installed.

DimensionReasoningScore

Conciseness

The body is dominated by executable commands and tables with almost no re-teaching of known concepts, matching anchor 4 ("efficient; minor instances of over-explanation that could be trimmed"). It falls short of anchor 5 because the "For plugin maintainers" section serves a different audience than the user-facing install flow, and the Troubleshooting table substantially duplicates the "Automated diagnosis & repair" section (e.g., the uv-shadowing and cache-clear fixes appear twice). It is well above anchor 3 since no section is padded with unnecessary explanation.

4 / 5

Actionability

Every section gives copy-paste-ready commands with expected outputs: the two-command install, "codex plugin list" with "expect: tooluniverse (enabled)", env-var exports with the crucial "then start codex from that same shell" caveat, the pin snippet with a pointer to valid versions, uninstall commands, a symptom-to-fix table, and a 7-step ordered diagnostic script that prints the FIX for each failure. This fully matches the anchor-5 "copy-paste ready... specific examples cover the common cases".

5 / 5

Workflow Clarity

The install flow is an explicit sequence (prerequisites check → two numbered install commands → restart → verify with expected output → keys → updates) with a dedicated validation section, and the diagnosis section is a checklist with a feedback loop ("Each is safe and idempotent; apply the FIX for whatever fails, then restart Codex"), matching anchor 5's "explicit validation steps; feedback loops for error recovery; checklists". Anchor 4 would require missing checkpoints, but verification is explicit at each stage.

5 / 5

Progressive Disclosure

Sections are clearly headed and easy to navigate, and detail is appropriately delegated ("Full key list: the bundled setup-tooluniverse skill → API_KEYS_REFERENCE.md" is a one-level, well-signaled pointer rather than an inlined key dump), matching anchor 4. It is not anchor 5 because no bundle files exist alongside this SKILL.md, so the lone reference points into another skill's bundle and cannot be verified here, and the maintainer-facing material is inlined rather than split out.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it enumerates the full action lifecycle with exact command-level specifics, provides comprehensive natural trigger phrasings including colloquial synonyms ("the harvard tools"), explicitly states when to use it, and cleanly fences off the sibling Claude Code skill and adjacent use cases. No vague fluff or over-claims are present.

DimensionReasoningScore

Specificity

"Install, set up, verify, update, pin, uninstall, or troubleshoot" lists multiple specific concrete actions covering the full plugin lifecycle, and it names exact technical artifacts ("mims-harvard/ToolUniverse", the "codex plugin marketplace add" then "codex plugin add -m tooluniverse" flow, the uvx tooluniverse MCP server, API-key env vars). Coverage is comprehensive rather than having the minor gaps of anchor 4, and the imperative voice matches the good examples rather than first/second person.

5 / 5

Completeness

It explicitly answers both questions: what ("Install, set up, verify, update, pin, uninstall, or troubleshoot the ToolUniverse plugin on OpenAI Codex") and when ("Use it whenever someone wants to... says... hits... asks... or finds it running an old tool version"), with concrete trigger phrases. It also adds an explicit negative scope ("Not for the Claude Code plugin..."), exceeding the anchor-5 example rather than the weaker 'when' of anchor 4.

5 / 5

Trigger Term Quality

It captures natural user phrasings comprehensively: "get ToolUniverse... working on Codex", "the 1000+ scientific tools", "the harvard tools", "won't load", "hits a uvx or MCP-server startup error", "asks how Codex updates it", "running an old tool version", and "even if they never say the word 'plugin'". This matches anchor 5's comprehensive synonym coverage; anchor 4 would require missing natural terms, and none are apparent.

5 / 5

Distinctiveness Conflict Risk

It occupies a clear niche (one named plugin on one named platform) and actively disambiguates from the closest conflict: "Not for the Claude Code plugin (use tooluniverse-claude-code-plugin), for running research with the tools, or for authoring new tools or skills". Triggers are distinct and mis-routing risk is minimal, matching anchor 5.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.