github.com/Orchestra-Research/AI-Research-SKILLs
| Skill | Added | Review |
|---|---|---|
langsmith-observability 17-observability/langsmith/SKILL.md LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systematic testing pipelines for AI applications. | 68 68 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
llama-cpp 12-inference-serving/llama-cpp/SKILL.md Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU. | 68 68 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
llama-factory 03-fine-tuning/llama-factory/SKILL.md Expert guidance for fine-tuning LLMs with LLaMA-Factory - WebUI no-code, 100+ models, 2/3/4/5/6/8-bit QLoRA, multimodal support | 40 40 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
llamaguard 07-safety-alignment/llamaguard/SKILL.md Meta's 7-8B specialized moderation model for LLM input/output filtering. 6 safety categories - violence/hate, sexual content, weapons, substances, self-harm, criminal planning. 94-95% accuracy. Deploy with vLLM, HuggingFace, Sagemaker. Integrates with NeMo Guardrails. | 56 56 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
llamaindex 14-agents/llamaindex/SKILL.md Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal support. Use for document Q&A, chatbots, knowledge retrieval, or building RAG pipelines. Best for data-centric LLM applications. | 64 64 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
llava 18-multimodal/llava/SKILL.md Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis. | 67 67 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
long-context 19-emerging-techniques/long-context/SKILL.md Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques. Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings. Covers rotary embeddings, attention biases, interpolation methods, and extrapolation strategies for LLMs. | 67 67 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
mamba-architecture 01-model-architecture/mamba/SKILL.md State-space model with O(n) complexity vs Transformers' O(n²). 5× faster inference, million-token sequences, no KV cache. Selective SSM with hardware-aware design. Mamba-1 (d_state=16) and Mamba-2 (d_state=128, multi-head). Models 130M-2.8B on HuggingFace. | 58 58 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
miles-rl-training 06-post-training/miles/SKILL.md Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inference alignment, or requiring speculative RL for maximum throughput. | 59 59 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
mlflow 13-mlops/mlflow/SKILL.md Track ML experiments, manage model registry with versioning, deploy models to production, and reproduce experiments with MLflow - framework-agnostic ML lifecycle platform | 61 61 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Version: 773a529 | |
ml-paper-writing 20-ml-paper-writing/ml-paper-writing/SKILL.md Write publication-ready ML/AI papers for NeurIPS, ICML, ICLR, ACL, AAAI, COLM. Use when drafting papers from research repos, structuring arguments, verifying citations, or preparing camera-ready submissions. For systems venues (OSDI, NSDI, ASPLOS, SOSP), use systems-paper-writing instead. | 71 71 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Version: 773a529 | |
ml-training-recipes 10-optimization/ml-training-recipes/SKILL.md Battle-tested PyTorch training recipes for all domains — LLMs, vision, diffusion, medical imaging, protein/drug discovery, spatial omics, genomics. Covers training loops, optimizer selection (AdamW, Muon), LR scheduling, mixed precision, debugging, and systematic experimentation. Use when training or fine-tuning neural networks, debugging loss spikes or OOM, choosing architectures, or optimizing GPU throughput. | 79 79 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
modal-serverless-gpu 09-infrastructure/modal/SKILL.md Serverless GPU cloud platform for running ML workloads. Use when you need on-demand GPU access without infrastructure management, deploying ML models as APIs, or running batch jobs with automatic scaling. | 64 64 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Version: 773a529 | |
model-merging 19-emerging-techniques/model-merging/SKILL.md Merge multiple fine-tuned models using mergekit to combine capabilities without retraining. Use when creating specialized models by blending domain-specific expertise (math + coding + chat), improving performance beyond single models, or experimenting rapidly with model variants. Covers SLERP, TIES-Merging, DARE, Task Arithmetic, linear merging, and production deployment strategies. | 68 68 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
model-pruning 19-emerging-techniques/model-pruning/SKILL.md Reduce LLM size and accelerate inference using pruning techniques like Wanda and SparseGPT. Use when compressing models without retraining, achieving 50% sparsity with minimal accuracy loss, or enabling faster inference on hardware accelerators. Covers unstructured pruning, structured pruning, N:M sparsity, magnitude pruning, and one-shot methods. | 63 63 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
moe-training 19-emerging-techniques/moe-training/SKILL.md Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. Use when training large-scale models with limited compute (5× cost reduction vs dense models), implementing sparse architectures like Mixtral 8x7B or DeepSeek-V3, or scaling model capacity without proportional compute increase. Covers MoE architectures, routing mechanisms, load balancing, expert parallelism, and inference optimization. | 67 67 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
nanogpt 01-model-architecture/nanogpt/SKILL.md Educational GPT implementation in ~300 lines. Reproduces GPT-2 (124M) on OpenWebText. Clean, hackable code for learning transformers. By Andrej Karpathy. Perfect for understanding GPT architecture from scratch. Train on Shakespeare (CPU) or OpenWebText (multi-GPU). | 64 64 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
nemo-curator 05-data-processing/nemo-curator/SKILL.md GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with RAPIDS. Use for preparing high-quality training datasets, cleaning web data, or deduplicating large corpora. | 70 70 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
nemo-evaluator-sdk 11-evaluation/nemo-evaluator/SKILL.md Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking. | 68 68 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Version: 773a529 | |
nemo-guardrails 07-safety-alignment/nemo-guardrails/SKILL.md NVIDIA's runtime safety framework for LLM applications. Features jailbreak detection, input/output validation, fact-checking, hallucination detection, PII filtering, toxicity detection. Uses Colang 2.0 DSL for programmable rails. Production-ready, runs on T4 GPU. | 52 52 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 |