github.com/synthetic-sciences/openscience
Skill | Added | Review |
|---|---|---|
grpo-rl-training backend/cli/skills/ml-training/grpo-rl-training/SKILL.md Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training | 57 57 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 52845c3 | |
gptq backend/cli/skills/ml-training/gptq/SKILL.md Post-training 4-bit quantization for LLMs with minimal accuracy loss. Use for deploying large models (70B, 405B) on consumer GPUs, when you need 4× memory reduction with <2% perplexity degradation, or for faster inference (3-4× speedup) vs FP16. Integrates with transformers and PEFT for QLoRA fine-tuning. | 70 70 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 52845c3 | |
geniml backend/cli/skills/ml-training/geniml/SKILL.md This skill should be used when working with genomic interval data (BED files) for machine learning tasks. Use for training region embeddings (Region2Vec, BEDspace), single-cell ATAC-seq analysis (scEmbed), building consensus peaks (universes), or any ML-based analysis of genomic regions. Applies to BED file collections, scATAC-seq data, chromatin accessibility datasets, and region-based genomic feature learning. | 69 69 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 52845c3 | |
deepspeed backend/cli/skills/ml-training/deepspeed/SKILL.md Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention | 40 40 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 52845c3 | |
colab-finetuning backend/cli/skills/ml-training/colab-finetuning/SKILL.md Fine-tune LLMs on Google Colab GPUs directly from openscience. Connects to Colab runtimes via WebSocket bridge for remote training with Unsloth. Supports SFT, GRPO, DPO, vision, and TTS workflows on free T4 to Pro A100 GPUs. | 57 57 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 52845c3 | |
evaluating-code-models backend/cli/skills/ml-training/bigcode-evaluation-harness/SKILL.md Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards. | 68 68 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 52845c3 | |
axolotl backend/cli/skills/ml-training/axolotl/SKILL.md Expert guidance for fine-tuning LLMs with Axolotl - YAML configs, 100+ models, LoRA/QLoRA, DPO/KTO/ORPO/GRPO, multimodal support | 52 52 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 52845c3 | |
awq-quantization backend/cli/skills/ml-training/awq/SKILL.md Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 52845c3 | |
adaptyv backend/cli/skills/ml-training/adaptyv/SKILL.md Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 52845c3 | |
huggingface-accelerate backend/cli/skills/ml-training/accelerate/SKILL.md Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mixed precision (FP16/BF16/FP8). Interactive config, single launch command. HuggingFace ecosystem standard. | 64 64 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 52845c3 | |
groq-inference backend/cli/skills/ml-inference/groq/SKILL.md Ultra-fast LLM inference on custom LPU hardware. OpenAI-compatible API at api.groq.com. Lowest latency in the industry (500-1000+ tok/s). Supports chat completions, vision, audio (Whisper STT + TTS), tool calling, JSON mode, and streaming. Free tier available. Inference only — no training. | 62 62 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 52845c3 | |
gguf-quantization backend/cli/skills/ml-inference/gguf/SKILL.md GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements. | 70 70 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 52845c3 | |
hugging-face-cli backend/cli/skills/llm-tools/hugging-face-cli/SKILL.md Execute Hugging Face Hub operations using the `hf` CLI. Use when the user needs to download models/datasets/spaces, upload files to Hub repositories, create repos, manage local cache, or run compute jobs on HF infrastructure. Covers authentication, file transfers, repository creation, cache operations, and cloud compute. | 70 70 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 52845c3 | |
guidance backend/cli/skills/llm-tools/guidance/SKILL.md Control LLM output with regex and grammars, guarantee valid JSON/XML/code generation, enforce structured formats, and build multi-step workflows with Guidance - Microsoft Research's constrained generation framework | 61 61 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 52845c3 | |
generate-image backend/cli/skills/llm-tools/generate-image/SKILL.md Generate or edit images using AI models (FLUX, Gemini). Use for general-purpose image generation including photos, illustrations, artwork, visual assets, concept art, and any image that isn't a technical diagram or schematic. For flowcharts, circuits, pathways, and technical diagrams, use the scientific-schematics skill instead. | 72 72 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Reviewed: Version: 52845c3 | |
faiss backend/cli/skills/llm-tools/faiss/SKILL.md Facebook's library for efficient similarity search and clustering of dense vectors. Supports billions of vectors, GPU acceleration, and various index types (Flat, IVF, HNSW). Use for fast k-NN search, large-scale vector retrieval, or when you need pure similarity search without metadata. Best for high-performance applications. | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 52845c3 | |
dspy backend/cli/skills/llm-tools/dspy/SKILL.md Build complex AI systems with declarative programming, optimize prompts automatically, create modular RAG systems and agents with DSPy - Stanford NLP's framework for systematic LM programming | 54 54 Impact — No eval scenarios have been run Securityby Critical Do not install without reviewing Reviewed: Version: 52845c3 | |
crewai-multi-agent backend/cli/skills/llm-tools/crewai/SKILL.md Multi-agent orchestration framework for autonomous AI collaboration. Use when building teams of specialized agents working together on complex tasks, when you need role-based agent collaboration with memory, or for production workflows requiring sequential/hierarchical execution. Built without LangChain dependencies for lean, fast execution. | 70 70 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 52845c3 | |
constitutional-ai backend/cli/skills/llm-tools/constitutional-ai/SKILL.md Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system. | 55 55 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 52845c3 | |
chroma backend/cli/skills/llm-tools/chroma/SKILL.md Open-source embedding database for AI applications. Store embeddings and metadata, perform vector and full-text search, filter by metadata. Simple 4-function API. Scales from notebooks to production clusters. Use for semantic search, RAG applications, or document retrieval. Best for local development and open-source projects. | 63 63 Impact — No eval scenarios have been run Securityby High Do not use without reviewing Reviewed: Version: 52845c3 | |
hmdb-database backend/cli/skills/databases/hmdb-database/SKILL.md Access Human Metabolome Database (220K+ metabolites). Search by name/ID/structure, retrieve chemical properties, biomarker data, NMR/MS spectra, pathways, for metabolomics and identification. | 56 56 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 52845c3 | |
gwas-database backend/cli/skills/databases/gwas-database/SKILL.md Query NHGRI-EBI GWAS Catalog for SNP-trait associations. Search variants by rs ID, disease/trait, gene, retrieve p-values and summary statistics, for genetic epidemiology and polygenic risk scores. | 61 61 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 52845c3 | |
geo-database backend/cli/skills/databases/geo-database/SKILL.md Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis. | 61 61 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Reviewed: Version: 52845c3 | |
gene-database backend/cli/skills/databases/gene-database/SKILL.md Query NCBI Gene via E-utilities/Datasets API. Search by symbol/ID, retrieve gene info (RefSeqs, GO, locations, phenotypes), batch lookups, for gene annotation and functional analysis. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 52845c3 | |
fred-economic-data backend/cli/skills/databases/fred-economic-data/SKILL.md Query FRED (Federal Reserve Economic Data) API for 800,000+ economic time series from 100+ sources. Access GDP, unemployment, inflation, interest rates, exchange rates, housing, and regional data. Use for macroeconomic analysis, financial research, policy studies, economic forecasting, and academic research requiring U.S. and international economic indicators. | 73 73 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Reviewed: Version: 52845c3 |