github.com/muratcankoylan/Agent-Skills-for-Context-Engineering
| Skill | Added | Review |
|---|---|---|
advanced-evaluation skills/advanced-evaluation/SKILL.md This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated quality assessment. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 6dbe1a1 | |
bdi-mental-states skills/bdi-mental-states/SKILL.md This skill should be used when modeling agent mental states with BDI concepts: beliefs, desires, intentions, RDF-to-belief transformations, rational agency traces, cognitive agents, BDI ontologies, and neuro-symbolic AI integration. | — | |
book-sft-pipeline examples/book-sft-pipeline/SKILL.md This skill should be used for book-to-SFT pipelines: ePub extraction, literary segmentation, author-voice dataset construction, style-transfer training, LoRA workflows, and model evaluation for voice replication. | — | |
comprehensive-research-agent examples/interleaved-thinking/generated_skills/comprehensive-research-agent/SKILL.md Ensure thorough validation, error recovery, and transparent reasoning in research tasks with multiple tool calls | 44 44 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 6dbe1a1 | |
context-compression skills/context-compression/SKILL.md This skill should be used when long-running agent sessions need context compression, structured summarization, compaction, token-per-task optimization, or durable handoff summaries that preserve decisions, files, risks, and next actions. | — | |
context-degradation skills/context-degradation/SKILL.md This skill should be used for diagnosing and mitigating context degradation: lost-in-middle failures, context poisoning, context clash, context confusion, attention-pattern issues, and agent performance degradation caused by accumulated or conflicting context. | — | |
context-engineering-collection SKILL.md A comprehensive collection of Agent Skills for context engineering, harness engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, evaluating, or debugging agent systems that require effective context management and reliable operating loops. | 52 52 Impact — No eval scenarios have been run Securityby — The risk profile of this skill Version: 6dbe1a1 | |
context-fundamentals skills/context-fundamentals/SKILL.md This skill should be used to explain or reason about the foundational concepts of context engineering: what context is, the anatomy of a context window, how attention mechanics work, the U-shaped attention curve, why context quality matters more than quantity, and the mental models needed to interpret every other context-engineering decision. Use this for conceptual explanation, onboarding, and background reading. Route operational work to the specialized skills: debugging attention failures goes to context-degradation, token-efficiency work goes to context-optimization, conversation summarization goes to context-compression, and project-shape decisions go to project-development. | — | |
context-optimization skills/context-optimization/SKILL.md This skill should be used for improving context efficiency: context budgeting, observation masking, prefix or KV-cache strategy, partitioning, token-cost reduction, retrieval scoping, and extending effective context capacity without lowering answer quality. | — | |
digital-brain examples/digital-brain-skill/SKILL.md This skill should be used for personal operating-system workflows: content creation, voice consistency, relationship lookup, meeting preparation, weekly review, goal tracking, personal brand management, and network management. | — | |
evaluation skills/evaluation/SKILL.md This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and outcome measurement for agent pipelines. | — | |
filesystem-context skills/filesystem-context/SKILL.md This skill should be used when agent work needs file-backed context: durable scratchpads, tool-output offloading, just-in-time discovery, cross-agent handoff files, filesystem memory, or cleanup policies for context stored outside the prompt. | — | |
harness-engineering skills/harness-engineering/SKILL.md This skill should be used when designing autonomous agent harnesses: research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR preparation, and human approval boundaries. | 71 71 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 6dbe1a1 | |
hosted-agents skills/hosted-agents/SKILL.md This skill should be used when designing hosted or background agent infrastructure: sandboxed execution, remote coding environments, warm pools, session persistence, multiplayer collaboration, self-spawning agents, or Modal-style sandboxes. | — | |
latent-briefing skills/latent-briefing/SKILL.md This skill should be used when the user asks to "share memory between agents", "KV cache compaction for multi-agent", "orchestrator worker context", "latent briefing", "reduce worker tokens", "cross-agent memory without summarization", or discusses Attention Matching compaction, recursive language models with workers, or token explosion in hierarchical agents. | 60 60 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 6dbe1a1 | |
long-horizon-prompting skills/long-horizon-prompting/SKILL.md This skill should be used when writing, enhancing, or evaluating the launch prompt for a long-running autonomous agent or a parallel multi-agent orchestration attacking a hard problem: pseudo-formal task briefs that define terms and an exact success predicate linguistically, enumerate non-counting outcomes, set persistence rules with explicit stop and return conditions and effort floors, manage a diverse portfolio of parallel approaches with an approach registry and blocked-route bookkeeping, and gate the return on adversarial audit. Route agent topology and coordination protocols to multi-agent-patterns, runtime control surfaces and loop governance to harness-engineering, evaluator and quality-gate construction to evaluation, judge design to advanced-evaluation, and compaction or memory mechanics to context-compression and memory-systems. | 70 70 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 6dbe1a1 | |
memory-systems skills/memory-systems/SKILL.md This skill should be used for persistent semantic memory in agent systems: cross-session knowledge retention, entity tracking, temporal validity, graph or vector retrieval, memory consolidation, and memory benchmark selection. Route file-backed scratchpads to filesystem-context, handoff summaries to context-compression, and token-efficiency tactics to context-optimization. | — | |
multi-agent-patterns skills/multi-agent-patterns/SKILL.md This skill should be used when designing multi-agent systems that need context isolation, supervisor or swarm coordination, explicit handoffs, parallel execution, or a decision on whether multiple agents are justified. | — | |
project-development skills/project-development/SKILL.md This skill should be used for project-level decisions about LLM-powered systems: whether an LLM is the right primitive for the task at hand, the shape of a multi-stage batch or agent pipeline, token and cost estimation, choosing between single-agent and multi-agent at the project level, structured output design for downstream parsing, and structuring agent-assisted iteration. Use this when the unit of work is a whole project or a multi-stage pipeline. Route individual tool design to tool-design and individual skill-loading or context-budget tactics to context-optimization. | — | |
reasoning-trace-optimizer examples/interleaved-thinking/SKILL.md Debug and optimize AI agents by analyzing reasoning traces, context degradation, tool confusion, instruction drift, repeated task failures, and performance regressions. | 55 55 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 6dbe1a1 |