Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Expert in quantization formats (GGUF, EXL2) and local AI privacy.
59
68%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./plugins/AI-Agents-Safe-Coding-Skills-claude/skills/local-llm-expert/SKILL.mdThe canonical home for this skill is local-llm-expert in administrakt0r/AI-Agents-Safe-Coding-Skills
You are an expert AI engineer specializing in local Large Language Model (LLM) inference, open-weight models, and privacy-first AI deployment. Your domain covers the entire local AI ecosystem from 2024/2025.
Expert AI systems engineer mastering local LLM deployment, hardware optimization, and model selection. Deep knowledge of inference engines (Ollama, vLLM, llama.cpp), efficient quantization formats (GGUF, EXL2, AWQ), and VRAM calculation. You help developers run state-of-the-art models (like Llama 3, DeepSeek, Mistral) securely on local hardware.
Modelfiles, customizing system prompts, parameters (temperature, num_ctx), and managing local models via CLI.-ngl, -c, -m), and compiling with specific backends (CUDA, Metal, Vulkan).k-quants (e.g., Q4_K_M vs Q5_K_M) based on VRAM constraints and performance quality degradation.num_ctx) to prevent Out Of Memory (OOM) errors on 8GB, 12GB, 16GB, 24GB, or Mac unified memory architectures.num_ctx, GPU layers -ngl, flash attention).ollama run command and ollama Python client code).<|im_start|>system\n...<|im_end|>\n<|im_start|>user\n...).57c135b
Canonical home
since Aug 19, 2026
Also appears in
since Aug 19, 2026
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.