According to AlphaMoat, as of 2026-07, Agent Evaluation has 5,763 downloads and 8 stars — ranked #1,155 of 106,927 Claude skills overall, and #217 of 17,222 in AI Agent.
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
| Month | Downloads | MoM | Stars | Installs |
|---|---|---|---|---|
| 2026-07 | 5.8K | +5.7% | 8 | 372 |
| 2026-06 | 5.5K | +5.2% | 8 | 372 |
| 2026-05 | 5.2K | +23.4% | 8 | 372 |
| 2026-03 | 4.2K | — | 6 | 329 |
Captures learnings, errors, and corrections to enable continuous improvement for the agent.
Browser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking...
Self-reflection + Self-criticism + Self-learning + Self-organizing memory. Agent evaluates its own work, catches mistakes, and improves permanently. Use when...
Security-first skill vetting for AI agents. Use before installing any skill from ClawdHub, GitHub, or other sources. Checks for red flags, permission scope, and suspicious patterns.
Typed knowledge graph for structured agent memory and composable skills. Use when creating/querying entities (Person, Project, Task, Event, Document), linkin...
Transform AI agents from task-followers into proactive partners that anticipate needs and continuously improve. Now with WAL Protocol, Working Buffer, Autonomous Crons, and battle-tested patterns. Part of the Hal Stack 🦞
Headless browser automation CLI optimized for AI agents with accessibility tree snapshots and ref-based element selection
Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.
Optimize multi-agent systems with coordinated profiling, workload distribution, and cost-aware orchestration. Use when improving agent performance, throughput, or reliability.
Humanize AI-generated text by removing telltale AI writing patterns. Use when text needs to sound natural and human-written — removing em-dashes, AI filler p...
Prevent context loss during LLM compaction via Write-Ahead Logging (WAL), Working Buffer, and automatic recovery. Three mechanisms that ensure critical state...
AI/LLM red team testing skill. Point at any LLM API endpoint and run automated security assessments. 160+ attack payloads across prompt injection, jailbreak,...
Interact with MoltX (Twitter for AI agents). Post, reply, like, follow, check notifications, and engage on moltx.io. Use when doing MoltX social engagement,...
Data month: 2026-07 · Downloads, stars and installs are aggregated monthly from public skill registries (ClawHub, SkillHub). See methodology.