CtrlK
BlogDocsLog inGet started
Tessl Logo

tensorrt-llm

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

64

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-inference/tensorrt-llm/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

72%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable overview skill: executable quick start, honest comparison against vLLM/llama.cpp, and clean one-level-deep offloading to three real reference files. Weaknesses are the absence of any validation/verification checkpoints in the deployment and batch workflows, a small bash syntax flaw in the serving example, and time-sensitive version pins and benchmark numbers presented without deprecation framing.

Suggestions

Add explicit validation checkpoints to the workflows: after starting trtllm-serve, verify with a health check or minimal curl before load; before batch generation, smoke-test one prompt; after multi-GPU launch, confirm per-GPU memory usage with nvidia-smi.

Fix the serving snippet by moving inline comments off the backslash-continued lines (e.g. put '# Tensor parallelism (4 GPUs)' above the flag or use a trailing comment style that doesn't break continuation), so the command is truly copy-paste ready.

Move volatile version pins (tensorrt_llm==1.2.0rc3, CUDA 13.0.0, TensorRT 10.13.2) and benchmark numbers into the references or a clearly labeled versions/benchmarks section, keeping the quick start version-agnostic.

DimensionReasoningScore

Conciseness

The body is code-forward and lean — installation, inference, serving, and patterns are all shown as runnable snippets with minimal prose. Minor trims possible: the opening one-liner repeats the frontmatter, the 'Supported models' list and 'Performance benchmarks' numbers are things Claude can look up or that drift, and inline version pins ('tensorrt_llm==1.2.0rc3', 'CUDA 13.0.0, TensorRT 10.13.2') are time-sensitive without an old-patterns/deprecated framing.

4 / 5

Actionability

Mostly copy-paste ready: a complete `from tensorrt_llm import LLM, SamplingParams` example with sampling config and generation, an FP8 config, a multi-GPU config, and a full trtllm-serve command with curl client. Not a 5 because the serving snippet places inline comments after line-continuation backslashes ('--tp_size 4 \ # Tensor parallelism'), which breaks the bash continuation and would fail if pasted as-is.

4 / 5

Workflow Clarity

A logical sequence exists (install → basic inference → serving → optimization patterns → deep-dive references), but there are no validation checkpoints anywhere: no 'verify GPUs with nvidia-smi', no post-start health check for the server, and no verify-output step for batch generation. The batch-inference pattern in particular runs 100 prompts with no verification guidance, so checkpoints are missing rather than implicit — matching the 3 anchor.

3 / 5

Progressive Disclosure

Clear overview structure with well-signaled, one-level-deep references: the References section links all three existing bundle files (optimization.md, multi-gpu.md, serving.md), each with a one-line description of its scope, and the reference files contain no further nesting. Quick-start content stays inline while advanced detail lives in the bundle.

5 / 5

Total

16

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete capability list, and an explicit 'Use for...' trigger clause covering deployment, quantization, batching, and multi-GPU scenarios. The main improvement space is broader synonym coverage and slightly sharper differentiation from alternative inference engines.

DimensionReasoningScore

Specificity

The description names the domain ('Optimizes LLM inference with NVIDIA TensorRT') and several concrete capabilities: 'quantization (FP8/INT4), in-flight batching, and multi-GPU scaling' plus throughput/latency goals. Not a 5 because coverage has gaps — no mention of model building/compilation, supported model families, or serving endpoints beyond the named techniques.

4 / 5

Completeness

It explicitly answers both what ('Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency') and when ('Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization...'), with concrete trigger phrases. Both anchors are fully satisfied.

5 / 5

Trigger Term Quality

Natural user phrases are well covered: 'LLM inference', 'production deployment on NVIDIA GPUs (A100/H100)', 'faster inference than PyTorch', 'quantization', 'in-flight batching', 'multi-GPU'. Not a 5 because a few natural synonyms are missing (e.g. 'speed up inference', 'tensorrt-llm' spelled as the product name users type, 'serving LLMs' as a standalone phrase).

4 / 5

Distinctiveness Conflict Risk

The niche is distinct — TensorRT plus NVIDIA-specific hardware (A100/H100) clearly separates it from generic inference skills. Not a 5 because '10-100x faster inference than PyTorch' is a request vLLM or other serving frameworks would also satisfy, creating minor overlap risk with sibling inference-serving skills.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.