Skip to content
agentscamp

Review & QA — AI Agents, Skills & Tools

Agents, skills, guides, tools, and commands for review & qa — 56 curated resources for building with AI coding agents.

Agent

Refactoring Specialist

Use this agent to safely restructure code without changing behavior — extracting, renaming, decoupling. Examples — breaking up a god object, removing duplication, improving testability.

sonnet
Agent

Accessibility Auditor

Use this agent to audit web UI against WCAG 2.2 AA — semantics, keyboard, ARIA, contrast, forms, and motion. Examples — auditing a new component for keyboard traps, checking a form for accessible errors, running a pre-ship a11y pass on a page.

sonnet4
Agent

Code Reviewer

Use this agent to review code changes for correctness, security, and maintainability before merging. Examples — reviewing a PR diff, auditing a new module, checking a refactor for regressions.

sonnet4
Agent

Debugger

Use this agent to diagnose failing tests, runtime errors, or unexpected behavior by forming and testing hypotheses. Examples — a stack trace to root-cause, a flaky test, a "works locally but not in CI" bug.

sonnet
Agent

Performance Engineer

Use this agent to profile and optimize performance — latency, throughput, memory, bundle size. Examples — a slow endpoint, an N+1 query, a heavy render, a large JS bundle.

opus
Agent

QA Automation Engineer

Use this agent for end-to-end and UI test automation — building flake-resistant Playwright/Cypress suites, stabilizing flaky browser tests, structuring page objects and fixtures, and reviewing E2E suites. Examples — adding E2E coverage for a checkout or signup flow, killing a test that fails 1-in-5 in CI, choosing a framework and folder structure, replacing sleeps with web-first waits, or auditing a suite that's slow and brittle.

sonnet5
Agent

Security Auditor

Use this agent to find security vulnerabilities — injection, auth flaws, secrets, unsafe deserialization, dependency risks. Examples — auditing an API surface, reviewing auth code, pre-release security pass.

opus4
Agent

Test Engineer

Use this agent to write and improve automated tests — unit, integration, and edge cases. Examples — adding coverage to an untested module, writing regression tests for a bug, designing a test plan.

sonnet6
Skill

Commit Splitter

Split one big, mixed-up change into a series of small, atomic commits — each a single logical change that builds and passes tests on its own — by grouping hunks by intent and staging them piecemeal. Use when a working tree or a fat commit mixes a feature, a refactor, a bug fix, and formatting, or before opening a PR you want reviewers to actually read.

invocablev1.0.0
Skill

Git Blame Investigator

Reconstruct why a line of code exists from Git history — find the originating commit, read its message and full diff for intent, and see through reformatting/rename commits with ignore-revs and the pickaxe — before you change or delete it. Use when a line looks wrong or pointless and you want to remove it, when tracing a regression to its commit, or when onboarding to unfamiliar code.

invocablev1.0.0
Skill

PR Description

Draft a clear pull request description from the branch diff against its base. Use when you have a finished branch and want a reviewer-ready PR body before opening the PR.

invocablev1.0.0
Skill

Bundle Analyzer

Analyze a JS/TS production bundle and surface the biggest size wins — heavy dependencies, duplicate packages, missing code-splitting, oversized polyfills, and dev/server code leaking into the client. Use when a bundle is too large and you need a ranked, actionable reduction plan.

invocablev1.0.0
Skill

Flamegraph Analyzer

Turn a CPU profile or flamegraph into a concrete optimization instead of guessing where the time goes: capture under a realistic workload with a sampling profiler, read the graph correctly (width = time, depth ≠ time), find the widest self-time leaves, ask if that work is necessary/redundant/algorithmically wrong, fix the biggest contributor, then re-profile. Use when code is CPU-bound and slow, a function is hot but you don't know which part, or you have a profile you can't interpret.

invocablev1.0.0
Skill

Memory Leak Hunter

Find and fix a memory leak in a running app: confirm it's a real leak under steady load, diff two heap snapshots to name the growing object and its retention path, cut the root reference that blocks collection, and re-run to confirm memory plateaus. Use when RSS climbs until OOM/restart, heap grows unbounded across a steady workload, or GC pauses worsen the longer the process runs.

invocablev1.0.0
Skill

React Render Profiler

Find and fix wasteful React re-renders by classifying the cause — unstable prop/callback/object identities, context value churn, state lifted too high, expensive work in render, or unvirtualized lists — confirming it with a measurement, then applying the one targeted fix and re-measuring. Use when a React UI is janky, slow to type in, or re-renders far more than the data actually changed.

invocablev1.0.0
Skill

Dead Code Finder

Find genuinely unused code — unreferenced exports, unreachable files, and unused dependencies — and remove it safely with build/test verification. Use when trimming a codebase or untangling years of accreted cruft.

invocablev1.0.0
Skill

Auth Flow Reviewer

Read-only review of authentication AND authorization flows — session/token model, cookie flags, CSRF, token rotation, password-reset/email-verification, OAuth redirect/state, and per-route object-level access checks — for exploitable gaps. Use before shipping login/session/token code, when adding a protected route or sharing-by-URL feature, or during a security pass. Reports findings by severity with location, impact, and the concrete fix; never edits code.

invocablev1.0.0
Skill

Secret Scanner

Scan a repo or a diff for committed secrets — API keys, tokens, private keys, .env files, and high-entropy strings — then triage real leaks from fixtures. Use before pushing, in review, or when a credential may have leaked.

invocablev1.0.0
Skill

Security Headers Hardener

Audit and harden a web app's or API's HTTP security headers — Content-Security-Policy, HSTS, X-Content-Type-Options, frame-ancestors, Referrer-Policy, Permissions-Policy, and CORS — and produce a staged rollout that won't break the site. Use before a launch, during a security pass, or when a scanner (Mozilla Observatory, securityheaders.com, a pentest) flags missing or weak headers. Audits and edits header config; rolls CSP out Report-Only first.

invocablev1.0.0
Skill

Threat Model Builder

Build a practical threat model for a feature or system using STRIDE — diagram the data flow, mark trust boundaries, enumerate concrete threats where data crosses them, and prioritize by likelihood × impact so security is reasoned about before shipping instead of bolted on after. Use when designing a feature that touches auth, money, or sensitive data, running a security design review, or hardening before a launch.

invocablev1.0.0
Skill

Contract Test Designer

Design consumer-driven contract tests between services so an API provider can't break its consumers unnoticed — without slow, flaky full end-to-end environments. Use when independent services or teams integrate over an API, when integration bugs only surface in staging or prod, or when E2E suites are too slow and brittle to catch breaking API changes.

invocablev1.0.0
Skill

Coverage Gap Finder

Run the project's coverage tool and identify the highest-value untested paths — error branches, edge cases, and critical modules — then propose specific test cases for each gap. Use when you have a coverage report but don't know where new tests will pay off most.

invocablev1.0.0
Skill

Integration Test Designer

Design integration tests that exercise components against REAL collaborators — actual database, queue, HTTP boundary — at a deliberately chosen seam, instead of a unit suite that mocks everything or a slow flaky full E2E. Use when bugs slip past green unit tests, when wiring or contracts between layers break in production, or when a mocked DB test passes but the real query/migration/serialization fails.

invocablev1.0.0
Skill

Mock Data Factory

Generate a typed mock/fixture factory for a given type, interface, or schema, inferring believable values from field names and types. Use when tests or local dev need realistic, type-safe sample data with per-field overrides.

invocablev1.0.0
Skill

Mutation Test Runner

Measure whether a test suite actually catches bugs by running mutation testing — introduce small faults into the code and check which ones a test kills versus which slip through silently. Use when line coverage is high but bugs still ship, when you suspect tests assert weakly, or to find the exact assertions a suite is missing.

invocablev1.0.0
Skill

Property Test Designer

Design property-based tests — generate hundreds of random inputs and assert invariants that must hold for ALL of them — to surface the edge cases hand-picked examples never reach. Use when code has a large input space (parsers, serializers, encoders, math, data transforms), when a bug keeps slipping through despite green example tests, or when you can't enumerate every case worth checking.

invocablev1.0.0
Skill

Test Scaffolder

Scaffold a test file with sensible cases for a given module or function. Use when adding tests to untested code and you want a fast, structured starting point.

invocablev1.0.0
Guide

Best AI Code Review Tools in 2026

The AI code reviewers worth running in 2026 — CodeRabbit, Greptile, and Qodo compared, plus open-source PR-Agent and when Copilot's review is enough.

3m read· AgentsCamp
Guide

Testing and Debugging Claude Code Skills

Verify a Claude Code skill triggers on the right prompts, check its output, and fix the five common failures — from vague triggers to broken paths.

7m read· AgentsCamp
Guide

TDD with AI Agents: Red-Green as an Agent Loop

Test-driven development found its killer app: agents. How write-the-test-first turns AI coding into a verifiable loop, and the workflow that makes it stick.

2m read· AgentsCamp
Guide

How to Test AI-Generated Code

AI writes the code; tests decide whether to trust it. The verification stack for agent-written changes — contracts, generated tests, and the review that's left.

2m read· AgentsCamp
Guide

Testing LLM Applications: How to Test Non-Deterministic Software

How to test software that calls LLMs when outputs are non-deterministic — the testing pyramid, assertion strategies, golden datasets, and CI gating.

6m read· AgentsCamp
Guide

An AI Code Review Workflow That Actually Catches Bugs

Layer the review stack — self-review, AI reviewers, tests, and a human pass focused on what machines miss — into a workflow tuned for AI-written code.

3m read· AgentsCamp
Tool

Coderabbit

An AI code reviewer that posts line-by-line feedback and summaries on every pull request.

freemiumreview
Tool

Greptile

An AI code review agent that reviews pull requests with full-codebase context — catching multi-file logical bugs and learning your team's standards.

paidreview
Tool

Playwright MCP

Microsoft's open-source MCP server that gives AI agents structured browser automation via Playwright's accessibility tree.

open sourcemcp
Tool

Qodo

A quality-first AI code review platform (ex-CodiumAI) — multi-agent PR review with your team's rules, plus IDE, CLI, and codebase-intelligence products.

freemiumreview
Command

Audit Accessibility

Audit a component or page for accessibility against WCAG — semantics, names, keyboard, ARIA, contrast, forms, motion.

/audit-accessibility<file, component, or page to audit>
Command

Explain Error

Diagnose an error message or stack trace and propose a fix.

/explain-error<error message or stack trace>
Command

Clean Branches

Safely prune merged and stale Git branches: drop dead remote-tracking refs, list merged candidates for review, then delete with the safe -d variant.

/clean-branches
Command

Create PR

Push the current branch and open a GitHub pull request with a generated title and body.

/create-pr[base branch or notes]
Command

Git Bisect

Drive git bisect to find the exact commit that introduced a regression.

/git-bisect<bug description; optional good and bad refs>
Command

Git Undo

Safely reverse the last Git operation in the current repo — pick the right tool (restore, reset --soft/--mixed, revert, or reflog recovery) based on what happened and whether it was already pushed, and confirm before anything destructive.

/git-undo<what to undo — e.g. 'last commit', 'unstage file X', 'the bad merge'; empty to detect>
Command

Resolve Merge Conflicts

Walk through resolving the in-progress merge, rebase, or cherry-pick conflict in the current repo by understanding both sides, then verify before continuing.

/resolve-conflict
Command

Find N+1 Queries

Scan code read-only for N+1 query patterns — loops that query per iteration and handlers that fan out per-row — and report each with a location, why it is N+1, and the concrete eager-load/batch/set-based fix.

/find-n-plus-one<path or area to scan (optional)>
Command

Extract Function

Extract a code region into a well-named function and update the call site.

/extract-function<file:lines or description>
Command

Optimize Imports

Remove unused imports and organize the rest in a file or directory per the project's conventions — preferring the project's own import tool where one is configured — without changing behavior.

/optimize-imports<file, directory, or empty for the current changed files>
Command

Refactor

Refactor the target for readability and structure without changing behavior.

/refactor[file or function]
Command

Find Bug

Investigate a reported symptom, form hypotheses, and locate the root cause.

/find-bug[symptom]
Command

Review PR

Review a pull request for correctness, security, and style, and summarize findings.

/review-pr[PR number]
Command

Review Tests

Review the quality of a test suite, not just whether it passes — find weak assertions, missing edge cases, and tests coupled to implementation.

/review-tests<test file or area to review>
Command

Security Scan

Scan the current diff or given paths for security vulnerabilities.

/security-scan[paths]
Command

Fix Failing Test

Diagnose and fix a failing test by finding the real root cause.

/fix-failing-test[test name or path]
Command

Hunt Flaky Tests

Reproduce a flaky test, find the real source of nondeterminism, and fix the cause.

/flaky-test-hunt<suspected test or area (optional)>
Command

Generate E2E Test

Scaffold a resilient end-to-end test for a user flow grounded in the real UI.

/generate-e2e-test<user flow to test>
Command

Write Tests

Generate tests covering the happy path and edge cases for the given target.

/write-tests[file or function]