github.com/OpenLAIR/dr-claw
| Skill | Added | Review |
|---|---|---|
speculative-decoding skills/emerging-techniques/speculative-decoding/SKILL.md Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies. | 62 62 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: d51b64e | |
tensorrt-llm skills/inference-serving/tensorrt-llm/SKILL.md Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling. | 68 68 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: d51b64e | |
training-llms-megatron skills/distributed-training/megatron-core/SKILL.md Trains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies. Use when training models >1B parameters, need maximum GPU efficiency (47% MFU on H100), or require tensor/pipeline/sequence/context/expert parallelism. Production-ready framework used for Nemotron, LLaMA, DeepSeek. | 70 70 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: d51b64e | |
verl-rl-training skills/post-training/verl/SKILL.md Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends. | 54 54 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: d51b64e |