github.com/Orchestra-Research/AI-Research-SKILLs
| Skill | Added | Review |
|---|---|---|
academic-plotting 20-ml-paper-writing/academic-plotting/SKILL.md Generates publication-quality figures for ML papers from research context. Given a paper section or description, extracts system components and relationships to generate architecture diagrams via Gemini. Given experiment results or data, auto-selects chart type and generates data-driven figures via matplotlib/seaborn. Use when creating any figure for a conference paper. | 74 74 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
ara-compiler 22-agent-native-research-artifact/compiler/SKILL.md Compiles any research input — PDF papers, GitHub repositories, experiment logs, code directories, or raw notes — into a complete Agent-Native Research Artifact (ARA) with cognitive layer (claims, concepts, heuristics), physical layer (configs, code stubs), exploration graph, and grounded evidence. Use when ingesting a paper or codebase into a structured, machine-executable knowledge package, building an ARA from scratch, or converting research outputs into a falsifiable, agent-traversable form. | 74 74 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Version: 773a529 | |
ara-research-manager 22-agent-native-research-artifact/research-manager/SKILL.md Records research provenance as a post-task epilogue, scanning conversation history at the end of a coding or research session to extract decisions, experiments, dead ends, claims, heuristics, and pivots, and writing them into the ara/ directory with user-vs-AI provenance tags. Use as a session epilogue — never during execution — to maintain a faithful, auditable trace of how a research project actually evolved. | 70 70 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Version: 773a529 | |
ara-rigor-reviewer 22-agent-native-research-artifact/rigor-reviewer/SKILL.md Performs ARA Seal Level 2 semantic epistemic review on Agent-Native Research Artifacts, scoring six dimensions (evidence relevance, falsifiability, scope calibration, argument coherence, exploration integrity, methodological rigor) and producing a constructive, severity-ranked report with a Strong Accept-to-Reject recommendation. Use after Level 1 structural validation passes, when an ARA needs an objective epistemic critique before publication or release. | 67 67 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Version: 773a529 | |
audiocraft-audio-generation 18-multimodal/audiocraft/SKILL.md PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation. | 64 64 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
autogpt-agents 14-agents/autogpt/SKILL.md Autonomous AI agent platform for building and deploying continuous agents. Use when creating visual workflow agents, deploying persistent autonomous agents, or building complex multi-step AI automation systems. | 56 56 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
autoresearch 0-autoresearch-skill/SKILL.md Orchestrates end-to-end autonomous AI research projects using a two-loop architecture. The inner loop runs rapid experiment iterations with clear optimization targets. The outer loop synthesizes results, identifies patterns, and steers research direction. Routes to domain-specific skills for execution, supports continuous agent operation via Claude Code /loop and OpenClaw heartbeat, and produces research presentations and papers. Use when starting a research project, running autonomous experiments, or managing a multi-hypothesis research effort. | 66 66 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Version: 773a529 | |
awq-quantization 10-optimization/awq/SKILL.md Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
axolotl 03-fine-tuning/axolotl/SKILL.md Expert guidance for fine-tuning LLMs with Axolotl - YAML configs, 100+ models, LoRA/QLoRA, DPO/KTO/ORPO/GRPO, multimodal support | 57 57 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
blip-2-vision-language 18-multimodal/blip-2/SKILL.md Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance. | 64 64 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
brainstorming-research-ideas 21-research-ideation/brainstorming-research-ideas/SKILL.md Guides researchers through structured ideation frameworks to discover high-impact research directions. Use when exploring new problem spaces, pivoting between projects, or seeking novel angles on existing work. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
chroma 15-rag/chroma/SKILL.md Open-source embedding database for AI applications. Store embeddings and metadata, perform vector and full-text search, filter by metadata. Simple 4-function API. Scales from notebooks to production clusters. Use for semantic search, RAG applications, or document retrieval. Best for local development and open-source projects. | 68 68 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Version: 773a529 | |
clip 18-multimodal/clip/SKILL.md OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding. | 63 63 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
constitutional-ai 07-safety-alignment/constitutional-ai/SKILL.md Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system. | 52 52 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
creative-thinking-for-research 21-research-ideation/creative-thinking-for-research/SKILL.md Applies cognitive science frameworks for creative thinking to CS and AI research ideation. Use when seeking genuinely novel research directions by leveraging combinatorial creativity, analogical reasoning, constraint manipulation, and other empirically grounded creative strategies. | 67 67 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
crewai-multi-agent 14-agents/crewai/SKILL.md Multi-agent orchestration framework for autonomous AI collaboration. Use when building teams of specialized agents working together on complex tasks, when you need role-based agent collaboration with memory, or for production workflows requiring sequential/hierarchical execution. Built without LangChain dependencies for lean, fast execution. | 70 70 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 | |
deepspeed 08-distributed-training/deepspeed/SKILL.md Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention | 41 41 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 773a529 | |
distributed-llm-pretraining-torchtitan 01-model-architecture/torchtitan/SKILL.md Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, torch.compile, and distributed checkpointing. | 69 69 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Version: 773a529 | |
dspy 16-prompt-engineering/dspy/SKILL.md Build complex AI systems with declarative programming, optimize prompts automatically, create modular RAG systems and agents with DSPy - Stanford NLP's framework for systematic LM programming | 56 56 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Version: 773a529 | |
evaluating-code-models 11-evaluation/bigcode-evaluation-harness/SKILL.md Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards. | 68 68 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 773a529 |