github.com/Orchestra-Research/AI-Research-SKILLs
| Skill | Added | Review |
|---|---|---|
evaluating-cosmos-policy 18-multimodal/cosmos-policy/SKILL.md Evaluates NVIDIA Cosmos Policy on LIBERO and RoboCasa simulation environments. Use when setting up cosmos-policy for robot manipulation evaluation, running headless GPU evaluations with EGL rendering, or profiling inference latency on cluster or local GPU machines. | 72 72 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Version: 773a529 | |
evaluating-llms-harness 11-evaluation/lm-evaluation-harness/SKILL.md Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs. | 70 70 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
evolving-ai-agents 14-agents/a-evolve/SKILL.md Provides guidance for automatically evolving and optimizing AI agents across any domain using LLM-driven evolution algorithms. Use when building self-improving agents, optimizing agent prompts and skills against benchmarks, or implementing automated agent evaluation loops. | 69 69 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Version: 773a529 | |
experiment-tracking-swanlab 13-mlops/swanlab/SKILL.md Provides guidance for experiment tracking with SwanLab. Use when you need open-source run tracking, local or self-hosted dashboards, and lightweight media logging for ML workflows. | 71 71 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
faiss 15-rag/faiss/SKILL.md Facebook's library for efficient similarity search and clustering of dense vectors. Supports billions of vectors, GPU acceleration, and various index types (Flat, IVF, HNSW). Use for fast k-NN search, large-scale vector retrieval, or when you need pure similarity search without metadata. Best for high-performance applications. | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
fine-tuning-openvla-oft 18-multimodal/openvla-oft/SKILL.md Fine-tunes and evaluates OpenVLA-OFT and OpenVLA-OFT+ policies for robot action generation with continuous action heads, LoRA adaptation, and FiLM conditioning on LIBERO simulation and ALOHA real-world setups. Use when reproducing OpenVLA-OFT paper results, training custom VLA action heads (L1 or diffusion), deploying server-client inference for ALOHA, or debugging normalization, LoRA merge, and cross-GPU issues. | 79 79 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
fine-tuning-serving-openpi 18-multimodal/openpi/SKILL.md Fine-tune and serve Physical Intelligence OpenPI models (pi0, pi0-fast, pi0.5) using JAX or PyTorch backends for robot policy inference across ALOHA, DROID, and LIBERO environments. Use when adapting pi0 models to custom datasets, converting JAX checkpoints to PyTorch, running policy inference servers, or debugging norm stats and GPU memory issues. | 78 78 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
fine-tuning-with-trl 06-post-training/trl-fine-tuning/SKILL.md Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers. | 63 63 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
gguf-quantization 10-optimization/gguf/SKILL.md GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements. | 64 64 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
gptq 10-optimization/gptq/SKILL.md Post-training 4-bit quantization for LLMs with minimal accuracy loss. Use for deploying large models (70B, 405B) on consumer GPUs, when you need 4× memory reduction with <2% perplexity degradation, or for faster inference (3-4× speedup) vs FP16. Integrates with transformers and PEFT for QLoRA fine-tuning. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
grpo-rl-training 06-post-training/grpo-rl-training/SKILL.md Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training | 51 51 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Version: 773a529 | |
guidance 16-prompt-engineering/guidance/SKILL.md Control LLM output with regex and grammars, guarantee valid JSON/XML/code generation, enforce structured formats, and build multi-step workflows with Guidance - Microsoft Research's constrained generation framework | 56 56 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Version: 773a529 | |
hqq-quantization 10-optimization/hqq/SKILL.md Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deploying with vLLM or HuggingFace Transformers. | 70 70 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
huggingface-accelerate 08-distributed-training/accelerate/SKILL.md Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mixed precision (FP16/BF16/FP8). Interactive config, single launch command. HuggingFace ecosystem standard. | 64 64 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
huggingface-tokenizers 02-tokenization/huggingface-tokenizers/SKILL.md Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
implementing-llms-litgpt 01-model-architecture/litgpt/SKILL.md Implements and trains LLMs using Lightning AI's LitGPT with 20+ pretrained architectures (Llama, Gemma, Phi, Qwen, Mistral). Use when need clean model implementations, educational understanding of architectures, or production fine-tuning with LoRA/QLoRA. Single-file implementations, no abstraction layers. | 69 69 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
instructor 16-prompt-engineering/instructor/SKILL.md Extract structured data from LLM responses with Pydantic validation, retry failed extractions automatically, parse complex JSON with type safety, and stream partial results with Instructor - battle-tested structured output library | 61 61 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
knowledge-distillation 19-emerging-techniques/knowledge-distillation/SKILL.md Compress large language models using knowledge distillation from teacher to student models. Use when deploying smaller models with retained performance, transferring GPT-4 capabilities to open-source models, or reducing inference costs. Covers temperature scaling, soft targets, reverse KLD, logit distillation, and MiniLLM training strategies. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
lambda-labs-gpu-cloud 09-infrastructure/lambda-labs/SKILL.md Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training. | 61 61 Impact — No eval scenarios have been run Securityby Medium Suggest reviewing before use Version: 773a529 | |
langchain 14-agents/langchain/SKILL.md Framework for building LLM-powered applications with agents, chains, and RAG. Supports multiple providers (OpenAI, Anthropic, Google), 500+ integrations, ReAct agents, tool calling, memory management, and vector store retrieval. Use for building chatbots, question-answering systems, autonomous agents, or RAG applications. Best for rapid prototyping and production deployments. | 60 60 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Version: 773a529 |