github.com/Orchestra-Research/AI-Research-SKILLs
Skill | Added | Review |
|---|---|---|
phoenix-observability 17-observability/phoenix/SKILL.md Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production AI systems with real-time insights. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 773a529 | |
langsmith-observability 17-observability/langsmith/SKILL.md LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systematic testing pipelines for AI applications. | 68 68 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 773a529 | |
outlines 16-prompt-engineering/outlines/SKILL.md Guarantee valid JSON/XML/code structure during generation, use Pydantic models for type-safe outputs, support local models (Transformers, vLLM), and maximize inference speed with Outlines - dottxt.ai's structured generation library | 56 56 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Reviewed: Version: 773a529 | |
instructor 16-prompt-engineering/instructor/SKILL.md Extract structured data from LLM responses with Pydantic validation, retry failed extractions automatically, parse complex JSON with type safety, and stream partial results with Instructor - battle-tested structured output library | 61 61 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 773a529 | |
guidance 16-prompt-engineering/guidance/SKILL.md Control LLM output with regex and grammars, guarantee valid JSON/XML/code generation, enforce structured formats, and build multi-step workflows with Guidance - Microsoft Research's constrained generation framework | 56 56 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 773a529 | |
dspy 16-prompt-engineering/dspy/SKILL.md Build complex AI systems with declarative programming, optimize prompts automatically, create modular RAG systems and agents with DSPy - Stanford NLP's framework for systematic LM programming | 56 56 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 773a529 | |
sentence-transformers 15-rag/sentence-transformers/SKILL.md Framework for state-of-the-art sentence, text, and image embeddings. Provides 5000+ pre-trained models for semantic similarity, clustering, and retrieval. Supports multilingual, domain-specific, and multimodal models. Use for generating embeddings for RAG, semantic search, or similarity tasks. Best for production embedding generation. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 773a529 | |
qdrant-vector-search 15-rag/qdrant/SKILL.md High-performance vector similarity search engine for RAG and semantic search. Use when building production RAG systems requiring fast nearest neighbor search, hybrid search with filtering, or scalable vector storage with Rust-powered performance. | 70 70 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Reviewed: Version: 773a529 | |
pinecone 15-rag/pinecone/SKILL.md Managed vector database for production AI applications. Fully managed, auto-scaling, with hybrid search (dense + sparse), metadata filtering, and namespaces. Low latency (<100ms p95). Use for production RAG, recommendation systems, or semantic search at scale. Best for serverless, managed infrastructure. | 68 68 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Reviewed: Version: 773a529 | |
faiss 15-rag/faiss/SKILL.md Facebook's library for efficient similarity search and clustering of dense vectors. Supports billions of vectors, GPU acceleration, and various index types (Flat, IVF, HNSW). Use for fast k-NN search, large-scale vector retrieval, or when you need pure similarity search without metadata. Best for high-performance applications. | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 773a529 | |
chroma 15-rag/chroma/SKILL.md Open-source embedding database for AI applications. Store embeddings and metadata, perform vector and full-text search, filter by metadata. Simple 4-function API. Scales from notebooks to production clusters. Use for semantic search, RAG applications, or document retrieval. Best for local development and open-source projects. | 68 68 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Reviewed: Version: 773a529 | |
llamaindex 14-agents/llamaindex/SKILL.md Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal support. Use for document Q&A, chatbots, knowledge retrieval, or building RAG pipelines. Best for data-centric LLM applications. | 64 64 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 773a529 | |
langchain 14-agents/langchain/SKILL.md Framework for building LLM-powered applications with agents, chains, and RAG. Supports multiple providers (OpenAI, Anthropic, Google), 500+ integrations, ReAct agents, tool calling, memory management, and vector store retrieval. Use for building chatbots, question-answering systems, autonomous agents, or RAG applications. Best for rapid prototyping and production deployments. | 60 60 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 773a529 | |
crewai-multi-agent 14-agents/crewai/SKILL.md Multi-agent orchestration framework for autonomous AI collaboration. Use when building teams of specialized agents working together on complex tasks, when you need role-based agent collaboration with memory, or for production workflows requiring sequential/hierarchical execution. Built without LangChain dependencies for lean, fast execution. | 70 70 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 773a529 | |
autogpt-agents 14-agents/autogpt/SKILL.md Autonomous AI agent platform for building and deploying continuous agents. Use when creating visual workflow agents, deploying persistent autonomous agents, or building complex multi-step AI automation systems. | 56 56 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 773a529 | |
evolving-ai-agents 14-agents/a-evolve/SKILL.md Provides guidance for automatically evolving and optimizing AI agents across any domain using LLM-driven evolution algorithms. Use when building self-improving agents, optimizing agent prompts and skills against benchmarks, or implementing automated agent evaluation loops. | 69 69 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 773a529 | |
weights-and-biases 13-mlops/weights-and-biases/SKILL.md Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B - collaborative MLOps platform | 61 61 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 773a529 | |
tensorboard 13-mlops/tensorboard/SKILL.md Visualize training metrics, debug models with histograms, compare experiments, visualize model graphs, and profile performance with TensorBoard - Google's ML visualization toolkit | 56 56 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 773a529 | |
experiment-tracking-swanlab 13-mlops/swanlab/SKILL.md Provides guidance for experiment tracking with SwanLab. Use when you need open-source run tracking, local or self-hosted dashboards, and lightweight media logging for ML workflows. | 71 71 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 773a529 | |
mlflow 13-mlops/mlflow/SKILL.md Track ML experiments, manage model registry with versioning, deploy models to production, and reproduce experiments with MLflow - framework-agnostic ML lifecycle platform | 61 61 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Reviewed: Version: 773a529 | |
serving-llms-vllm 12-inference-serving/vllm/SKILL.md Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism. | 65 65 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 773a529 | |
tensorrt-llm 12-inference-serving/tensorrt-llm/SKILL.md Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling. | 70 70 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 773a529 | |
sglang 12-inference-serving/sglang/SKILL.md Fast structured generation and serving for LLMs with RadixAttention prefix caching. Use for JSON/regex outputs, constrained decoding, agentic workflows with tool calls, or when you need 5× faster inference than vLLM with prefix sharing. Powers 300,000+ GPUs at xAI, AMD, NVIDIA, and LinkedIn. | 68 68 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 773a529 | |
llama-cpp 12-inference-serving/llama-cpp/SKILL.md Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU. | 68 68 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 773a529 | |
nemo-evaluator-sdk 11-evaluation/nemo-evaluator/SKILL.md Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking. | 68 68 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 773a529 |