github.com/synthetic-sciences/openscience
| Skill | Added | Review |
|---|---|---|
torchdrug backend/cli/skills/chemistry/torchdrug/SKILL.md PyTorch-native graph neural networks for molecules and proteins. Use when building custom GNN architectures for drug discovery, protein modeling, or knowledge graph reasoning. Best for custom model development, protein property prediction, retrosynthesis. For pre-trained models and diverse featurizers use deepchem; for benchmark datasets use pytdc. | 65 65 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
torchforge-rl-training backend/cli/skills/ml-training/torchforge/SKILL.md Provides guidance for PyTorch-native agentic RL using torchforge, Meta's library separating infra from algorithms. Use when you want clean RL abstractions, easy algorithm experimentation, or scalable training with Monarch and TorchTitan. | 52 52 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
torch-geometric backend/cli/skills/coding/torch_geometric/SKILL.md Graph Neural Networks (PyG). Node/graph classification, link prediction, GCN, GAT, GraphSAGE, heterogeneous graphs, molecular property prediction, for geometric deep learning. | 55 55 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
training-data-pipeline backend/cli/skills/ml-training/training-data-pipeline/SKILL.md Build training datasets for LLM specialization from production data, frontier model distillation, and synthetic bootstrapping. Use when formatting production logs into SFT data, distilling from frontier APIs, or preparing data for fine-tuning. Covers JSONL formatting, data quality validation, deduplication, and train/eval splitting. | 62 62 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
training-llms-megatron backend/cli/skills/ml-training/megatron-core/SKILL.md Trains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies. Use when training models >1B parameters, need maximum GPU efficiency (47% MFU on H100), or require tensor/pipeline/sequence/context/expert parallelism. Production-ready framework used for Nemotron, LLaMA, DeepSeek. | 65 65 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
transformer-lens-interpretability backend/cli/skills/ml-training/transformer-lens/SKILL.md Provides guidance for mechanistic interpretability research using TransformerLens to inspect and manipulate transformer internals via HookPoints and activation caching. Use when reverse-engineering model algorithms, studying attention patterns, or performing activation patching experiments. | 66 66 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 4082a2e | |
transformers backend/cli/skills/llm-tools/transformers/SKILL.md This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets. | 63 63 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
treatment-plans backend/cli/skills/biology/treatment-plans/SKILL.md Generate concise (3-4 page), focused medical treatment plans in LaTeX/PDF format for all clinical specialties. Supports general medical treatment, rehabilitation therapy, mental health care, chronic disease management, perioperative care, and pain management. Includes SMART goal frameworks, evidence-based interventions with minimal text citations, regulatory compliance (HIPAA), and professional formatting. Prioritizes brevity and clinical actionability. | — | |
umap-learn backend/cli/skills/coding/umap-learn/SKILL.md UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data. | 60 60 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
uniprot-database backend/cli/skills/databases/uniprot-database/SKILL.md Direct REST API access to UniProt. Protein searches, FASTA retrieval, ID mapping, Swiss-Prot/TrEMBL. For Python workflows with multiple databases, prefer bioservices (unified interface to 40+ services). Use this for direct HTTP/REST work or UniProt-specific control. | 61 61 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
unsloth-fine-tuning backend/cli/skills/ml-training/unsloth/SKILL.md Fast LLM fine-tuning with Unsloth - 2-5x faster training, 50-80% less VRAM. Use for single-GPU LoRA/QLoRA SFT, GRPO/RL reasoning training, vision/TTS fine-tuning, and GGUF export to Ollama/vLLM/llama.cpp. Supports 300+ models including Llama, Qwen, Gemma, DeepSeek, Mistral, Phi, and gpt-oss. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
uspto-database backend/cli/skills/databases/uspto-database/SKILL.md Access USPTO APIs for patent/trademark searches, examination history (PEDS), assignments, citations, office actions, TSDR, for IP analysis and prior art searches. | 60 60 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
vaex backend/cli/skills/data-engineering/vaex/SKILL.md Use this skill for processing and analyzing large tabular datasets (billions of rows) that exceed available RAM. Vaex excels at out-of-core DataFrame operations, lazy evaluation, fast aggregations, efficient visualization of big data, and machine learning on large datasets. Apply when users need to work with large CSV/HDF5/Arrow/Parquet files, perform fast statistics on massive datasets, create visualizations of big data, or build ML pipelines that do not fit in memory. | 66 66 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
vast-ai-gpu-cloud backend/cli/skills/cloud-compute/vast-ai/SKILL.md Safely inspect and operate Vast.ai marketplace instances with the vastai CLI, live offer data, and explicit approval before paid or destructive actions. | 67 67 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 4082a2e | |
verl-rl-training backend/cli/skills/ml-training/verl/SKILL.md Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends. | 60 60 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
wave-propagation backend/cli/skills/physics/wave-propagation/SKILL.md Simulate wave propagation — acoustic, electromagnetic, elastic, and quantum waves. FDTD, spectral methods, and absorbing boundary conditions for 1D/2D/3D wave equations with sources, scattering, and dispersion. | 66 66 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
weights-and-biases backend/cli/skills/ml-training/weights-and-biases/SKILL.md Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B - collaborative MLOps platform | 52 52 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
whisper backend/cli/skills/llm-tools/whisper/SKILL.md OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR. | 56 56 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
zarr-python backend/cli/skills/data-engineering/zarr-python/SKILL.md Chunked N-D arrays for cloud storage. Compressed arrays, parallel I/O, S3/GCS integration, NumPy/Dask/Xarray compatible, for large-scale scientific computing pipelines. | 53 53 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e | |
zinc-database backend/cli/skills/databases/zinc-database/SKILL.md Access ZINC (230M+ purchasable compounds). Search by ZINC ID/SMILES, similarity searches, 3D-ready structures for docking, analog discovery, for virtual screening and drug discovery. | 55 55 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 4082a2e |