CtrlK
BlogDocsLog inGet started
Tessl Logo

tensorrt-llm

High-throughput LLM inference on NVIDIA GPUs.

48

Quality

53%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/tensorrt-llm/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, highly actionable overview with executable installation, inference, and serving examples and a clean three-file reference layer. Its weaknesses are inline explanations of well-known concepts, time-sensitive version/benchmark figures, and a lack of validation checkpoints in the serving workflow.

Suggestions

Trim or remove definitions of concepts Claude already knows (Flash Attention, Paged KV cache, tensor/pipeline parallelism) and rely on the reference guides for detail.

Move version-sensitive details (CUDA/TensorRT version pins, benchmark numbers) into the references or a dedicated section so the body stays stable.

Add an explicit verification step to the serving workflow (e.g., curl the health endpoint or send a test completion before declaring the server ready).

DimensionReasoningScore

Conciseness

The body is mostly efficient but re-explains concepts Claude already knows ("Flash Attention: Optimized attention kernels", "Paged KV cache: Efficient memory management", "Tensor parallelism (TP): Split model across GPUs") and inlines time-sensitive version and benchmark figures ("CUDA 13.2.1, TensorRT 10.x", "24,000 tokens/sec") outside any old-patterns section, matching 'includes some unnecessary explanation or could be tightened'. Not a 4 because the padding and stale-version risk exceed minor instances.

3 / 5

Actionability

Concrete, copy-paste-ready guidance throughout: docker/pip install commands, the Python LLM API example, the trtllm-serve invocation with flags, and a curl client request cover the common cases, matching 'Mostly executable guidance; concrete code or commands with minor gaps'. Not a 5 because the batch-inference and FP8 snippets reference variables without their defining imports/context and no error-handling example exists.

4 / 5

Workflow Clarity

The quick start sequences install → basic inference → serving, but there are no validation or verification checkpoints (no health check after starting the server, no output verification), matching 'sequence present but checkpoints missing or implicit'. Not a 4 because verification steps for the serving workflow are genuinely absent, not merely implicit.

3 / 5

Progressive Disclosure

A clear overview with three well-signaled, one-level-deep references (references/optimization.md, references/multi-gpu.md, references/serving.md — all verified to exist, none nesting further), matching 'Good structure; most content is appropriately placed; references mostly clear'. Not a 5 because benchmark tables and the supported-models list are inline material that could be split into a reference file.

4 / 5

Total

14

/

20

Passed

Description

45%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names a clear domain, but it is a single clause with no concrete capability list, no 'Use when' trigger guidance, and missing natural synonyms like serving or deployment. It reads more like a tagline than a triggerable skill description.

Suggestions

Add an explicit trigger clause, e.g. "Use when deploying or serving LLMs on NVIDIA GPUs (A100/H100), optimizing inference throughput or latency, or using FP8/INT4 quantization."

List 2-3 concrete capabilities in the description (e.g., quantized model serving, multi-GPU tensor parallelism, in-flight batching) to raise specificity.

Include natural synonyms users would say ("LLM serving", "inference optimization", "production inference", "TensorRT") to widen trigger coverage without adding fluff.

DimensionReasoningScore

Specificity

"High-throughput LLM inference on NVIDIA GPUs" names the domain and one performance attribute but lists no concrete actions, matching the anchor 'Names the domain but actions are minimal or generic'. It is not a 3 because no concrete capabilities (e.g., quantization, serving, batching) are stated.

2 / 5

Completeness

The 'what' is clear (high-throughput LLM inference on NVIDIA GPUs) but there is no 'Use when' clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. It is not a 4 because the 'when' is entirely missing rather than just under-specified.

3 / 5

Trigger Term Quality

"LLM inference", "NVIDIA GPUs", and "high-throughput" are natural phrases a user might say, but common variations like "serving", "deployment", "production", or "TensorRT" are absent, matching 'Some relevant keywords but missing common variations'. Keyword coverage is too thin for a 4.

3 / 5

Distinctiveness Conflict Risk

The NVIDIA GPU scope carves a niche, but the description alone overlaps with general LLM serving/inference skills such as vLLM, matching 'Somewhat specific but could still overlap with similar skills'. It is not a 4 because no distinguishing trigger separates it from those neighbors.

3 / 5

Total

11

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.