github.com/calesthio/OpenMontage
| Skill | Added | Review |
|---|---|---|
3d-asset-generation .agents/skills/3d-asset-generation/SKILL.md Generate, reconstruct, inspect, and route production 3D assets for OpenMontage worlds using Atlas Cloud, fal.ai, licensed catalogs, and Blender. | 65 65 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
acestep .agents/skills/acestep/SKILL.md AI music generation with ACE-Step 1.5 — background music, vocal tracks, covers, stem extraction for video production. Use when generating music, soundtracks, jingles, or working with audio stems. Triggers include background music, soundtrack, jingle, music generation, stem extraction, cover, style transfer, or musical composition tasks. | 74 74 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
agents .agents/skills/agents/SKILL.md Build voice AI agents with ElevenLabs. Use when creating voice assistants, customer service bots, interactive voice characters, or any real-time voice conversation experience. | 84 84 1.73x Agent success vs baseline Impact 92% 1.73xAverage score across 3 eval scenarios Securityby Low Low-risk findings worth noting Reviewed: Version: 08e2151 | |
ai-video-gen .agents/skills/ai-video-gen/SKILL.md Generate AI videos from text prompts using multiple provider gateways. Use when: (1) Generating videos from text descriptions, (2) Creating AI-generated video clips for content production, (3) Image-to-video generation with a reference image, (4) Choosing between video generation providers (VEO, Kling, Sora, Runway, Seedance, MiniMax, Gemini Omni). Supports gateways: HeyGen API, fal.ai API, Kling official direct API, and the Gemini API (Gemini Omni Flash). | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
ai-video-gen .claude/skills/ai-video-gen/SKILL.md Generate AI videos from text prompts using multiple provider gateways. Use when: (1) Generating videos from text descriptions, (2) Creating AI-generated video clips for content production, (3) Image-to-video generation with a reference image, (4) Choosing between video generation providers (VEO, Kling, Sora, Runway, Seedance, MiniMax, Gemini Omni). Supports gateways: HeyGen API, fal.ai API, and the Gemini API (Gemini Omni Flash). | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
atlas-cloud .agents/skills/atlas-cloud/SKILL.md Generate or edit images and videos through the Atlas Cloud gateway. Use for Atlas-hosted Seedance 2.5/2.0, Gemini Omni Flash, MiniMax H3, Seedream 5.0, GPT Image 2, Nano Banana 2, or when one ATLASCLOUD_API_KEY should access multiple media model families. | 70 70 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
avatar-video .agents/skills/avatar-video/SKILL.md Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API. Use when: (1) Choosing a specific avatar and voice for a video, (2) Writing exact scripts for an avatar to speak, (3) Building multi-scene videos with different backgrounds per scene, (4) Creating transparent WebM videos for compositing, (5) Using talking photos as video presenters, (6) Integrating HeyGen avatars with Remotion, (7) Batch video generation with exact specs, (8) Brand-consistent production videos with precise control. | 66 66 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
azure-speech-to-text .agents/skills/azure-speech-to-text/SKILL.md Transcribe audio to text using Azure AI Speech (Fast Transcription REST API). Use when converting audio/video to text, generating subtitles, or processing spoken content in OpenMontage. Optional cloud STT provider — preferred when AZURE_SPEECH_KEY is configured; the local faster-whisper `transcriber` is the default offline path. | 68 68 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 08e2151 | |
azure-text-to-speech .agents/skills/azure-text-to-speech/SKILL.md Generate neural narration audio using Azure AI Speech (REST text-to-speech). Use when synthesizing voiceovers or narration in OpenMontage. Optional cloud TTS provider — preferred when AZURE_SPEECH_KEY is configured; the local piper_tts remains the default offline path. Shares one Speech resource with azure_stt. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
beautiful-mermaid .agents/skills/beautiful-mermaid/SKILL.md Render Mermaid diagrams as SVG and PNG using the Beautiful Mermaid library. Use when the user asks to render a Mermaid diagram. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
bfl-api .agents/skills/bfl-api/SKILL.md BFL FLUX API integration guide covering endpoints, async polling patterns, rate limiting, error handling, webhooks, and regional endpoints with Python and TypeScript code examples. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
character-animation-qa .agents/skills/character-animation-qa/SKILL.md Review local character animation with schema checks, Playwright browser previews, frame sampling, and FFmpeg/ffprobe final output checks. | 60 60 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
comfyui .agents/skills/comfyui/SKILL.md Use when working with ComfyUI workflows in OpenMontage, including comfyui_image/comfyui_video/comfyui_music, custom workflow_json/workflow_path inputs, output_node selection, missing model setup, LoRAs, low-VRAM workflow choices, and community workflow imports. | 69 69 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
create-video .agents/skills/create-video/SKILL.md Create videos from a text prompt using HeyGen's Video Agent. Use when: (1) Creating a video from a description or idea, (2) Generating explainer, demo, or marketing videos from a prompt, (3) Making a video without specifying exact avatars, voices, or scenes, (4) Quick video prototyping or drafts, (5) One-shot prompt-to-video generation, (6) User says "make me a video" or "create a video about X". | 70 70 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
d3-viz .agents/skills/d3-viz/SKILL.md Creating interactive data visualisations using d3.js. This skill should be used when creating custom charts, graphs, network diagrams, geographic visualisations, or any complex SVG-based data visualisation that requires fine-grained control over visual elements, transitions, or interactions. Use this for bespoke visualisations beyond standard charting libraries, whether in React, Vue, Svelte, vanilla JavaScript, or any other environment. | 67 67 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
dashscope .agents/skills/dashscope/SKILL.md DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). Use when generating images via Qwen-Image, narrating via Qwen-TTS, or transcribing with word-level timestamps via Qwen-ASR. | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
doubao-tts .agents/skills/doubao-tts/SKILL.md Generate Mandarin and multilingual narration with Volcengine Doubao Speech 2.0. Use when creating Chinese voiceovers, when the user prefers Doubao/Volcengine/火山引擎/豆包 TTS, or when narration needs character-level timestamp metadata for subtitles. | 71 71 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
elevenlabs .agents/skills/elevenlabs/SKILL.md Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio synthesis task. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
faceswap .agents/skills/faceswap/SKILL.md Swap faces in a video using AI via the HeyGen API. Use when: (1) Replacing a face in a video with another face, (2) Face swapping from a source image onto a target video, (3) Creating personalized videos by swapping in a person's face, (4) Working with HeyGen's /v1/workflows/executions endpoint for face swap processing. | 67 67 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
ffmpeg .agents/skills/ffmpeg/SKILL.md Video and audio processing with FFmpeg. Use for format conversion, resizing, compression, audio extraction, and preparing assets for Remotion. Triggers include converting GIF to MP4, resizing video, extracting audio, compressing files, or any media transformation task. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
fish-audio-tts .agents/skills/fish-audio-tts/SKILL.md Generate expressive, multilingual narration with fish.audio (S1 / S2-generation models) and reuse cloned voices via reference_id. Use when the user prefers fish.audio/Fish Audio TTS, wants a specific playground voice model, or needs high-emotion voice-clone narration. | 70 70 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
flux-best-practices .agents/skills/flux-best-practices/SKILL.md Comprehensive guide for BFL FLUX image generation models. Covers prompting, T2I, I2I, structured JSON, hex colors, typography, multi-reference editing, and model-specific best practices for FLUX.2 and FLUX.1 families. | 63 63 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
framer-motion .agents/skills/framer-motion/SKILL.md Use when implementing Disney's 12 animation principles with Framer Motion in React applications | 59 59 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
gemini-omni .agents/skills/gemini-omni/SKILL.md Generate and conversationally edit short videos with Google Gemini Omni Flash (`gemini-omni-flash-preview`). Use when: (1) iterating on a clip with natural-language edits instead of regenerating ("make the phone invisible, keep everything else the same"), (2) generating 3-10s 720p clips with synthesized audio, rendered on-screen text, or timecoded beats, (3) binding reference images to roles with <FIRST_FRAME>/<IMAGE_REF_N> prompt tags, (4) editing an existing uploaded video. Accessed via the `gemini_omni_video` tool using the project's GEMINI_API_KEY/GOOGLE_API_KEY — the same key as Imagen and Google TTS. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 | |
grok-media .agents/skills/grok-media/SKILL.md xAI Grok image and video generation guide covering authentication, endpoints, prompt structure, image editing, reference-image video, and async polling. | 54 54 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 08e2151 |