Skip to content
agentscamp

AI Agents & Systems — AI Agents, Skills & Tools

Agents, skills, guides, tools, and commands for ai agents & systems — 79 curated resources for building with AI coding agents.

Agent

Agent Tool Integration Engineer

Use this agent to wire tools and function-calling into an agent loop reliably — clean tool schemas, errors fed back as observations, retries with limits, idempotency, and parallel calls. Examples — "connect our APIs as agent tools", "our agent calls tools wrong / ignores tool errors", "add function-calling with proper error recovery to our agent".

sonnet6
Agent

Browser Agent Engineer

Use this agent to build, harden, or debug browser-automation agents — web tasks via Browser Use, Stagehand, Skyvern, or Playwright-based stacks. Examples: automate a portal workflow, make a flaky browser agent reliable, add verification and guardrails to web automation, choose between vision and DOM grounding.

sonnet
Agent

Agent Reliability Reviewer

Use this agent to make an AI agent production-ready — reviewing its loops, cost controls, error handling, tool use, human-in-the-loop gates, checkpointing, and observability, then reporting concrete failure modes and fixes. Examples — "is our agent safe to ship?", "our agent loops forever / burns tokens, harden it", "add guardrails and recovery before we put this agent in front of users".

sonnet6
Agent

Context Engineer

Use this agent to engineer what an LLM agent carries in its context window — deciding what to include vs exclude vs retrieve on demand, designing project/agent memory (CLAUDE.md), compacting growing history, and allocating the token budget across system prompt, memory, retrieved docs, tool results, and conversation. Examples — "my agent forgets the schema we agreed on three turns ago", "the agent gets dumber and more inconsistent as the chat grows", "we're burning 60k tokens of tool output every turn", "what should this support agent always know vs look up?".

opus3
Agent

Eval Driven Developer

Use this agent to drive AI feature development with evals the way TDD drives code with tests — define success criteria and a representative eval set BEFORE iterating on prompts/models, then optimize against measured scores instead of vibes. Examples — "make the summarizer better" (turn it into measurable criteria first), "our prompt change keeps regressing quality, set up a loop that catches it", "add an eval gate to CI so a model swap can't silently degrade output", "we tweak prompts and pray — give us a baseline and a change-by-change scoreboard".

opus5
Skill

Tool Definition Generator

Generate clean function/tool schemas for an LLM agent from existing code or a spec — accurate JSON Schema, model-facing descriptions, honest required fields, and enums that make invalid calls impossible. Use when wiring functions into an agent's tool-calling loop.

invocablev1.0.0
Skill

Agent Trajectory Evaluator

Evaluate a multi-step AI agent's whole run — tool calls, intermediate steps, and final result — not just final-answer correctness, so you can pinpoint WHERE it went wrong. Use when building or debugging a tool-using or multi-step agent, when final-answer-only evals can't explain failures, or when a prompt/model change quietly makes the agent less efficient or more error-prone even though the answer still looks right.

invocablev1.0.0
Skill

Web Research Pipeline

Run a structured web-research pass on a question: plan the searches, find sources via search APIs, fetch and read the best ones, cross-check claims, and synthesize a cited answer — with source quality and disagreements surfaced honestly. Use for 'research X and tell me what's actually true' tasks that need more than one search and less than a day.

invocablev1.0.0
Skill

Agent Memory Designer

Design a project's CLAUDE.md and memory hierarchy by exploring the repo to learn its real build/test/lint commands, architecture, and non-obvious gotchas, then writing a concise, skimmable memory that keeps only what belongs in context. Use when onboarding a repo to Claude Code with no CLAUDE.md, or when an existing one is bloated, stale, or being ignored.

invocablev1.0.0
Skill

Human In The Loop Gate

Add a human approval checkpoint to an agent so it pauses before a risky or irreversible action (spending money, deleting data, sending messages, merging code) and resumes only after a human approves. Use when an agent acts autonomously on consequential operations.

invocablev1.0.0
Guide

Sandboxing AI-Generated Code: E2B vs Modal vs Daytona vs Vercel Sandbox

Where should agent-written code run? E2B vs Modal vs Daytona vs Vercel Sandbox compared on isolation, persistence, and cost, plus rules for safe execution.

2m read· AgentsCamp
Guide

Aider vs Claude Code: Open-Source vs Anthropic's Agent (2026)

Aider vs Claude Code — model-agnostic open-source pair-programmer vs Anthropic's tuned terminal harness. Which terminal coding agent fits your stack.

3m read· AgentsCamp
Guide

Browser Agents in 2026: Browser Use vs Stagehand vs Skyvern vs Playwright MCP

Four ways to give AI a browser — Browser Use, Stagehand, Skyvern, and Playwright MCP compared honestly on control, cost, and reliability for 2026.

2m read· AgentsCamp
Guide

Claude Code vs Gemini CLI: Which Terminal Agent (2026)

Claude Code vs Gemini CLI: first-party stability and a deep programmable harness vs open-source TypeScript, a big free tier, and the Antigravity cutover.

3m read· AgentsCamp
Guide

Claude Skills vs Custom GPTs: Different Answers to Reuse

Claude Skills are portable procedures; Custom GPTs are packaged chatbots inside ChatGPT. How they differ on portability, API access, sharing, and where each wins.

3m read· AgentsCamp
Guide

LangChain vs LlamaIndex in 2026: Agents or Data?

The classic framework confusion resolved — LangChain's agent loop and ecosystem vs LlamaIndex's data-and-documents depth — and when you'd genuinely use both.

2m read· AgentsCamp
Guide

LangGraph vs CrewAI: Agent Frameworks Compared (2026)

LangGraph vs CrewAI — explicit state-machine control vs role-based crew abstractions. Which agent framework fits your reliability bar and team.

2m read· AgentsCamp
Guide

Mem0 vs Zep vs Letta: Agent Memory Compared (2026)

Three philosophies of agent memory — Mem0's drop-in layer, Zep's temporal knowledge graphs, Letta's self-managing agents — and which fits your architecture.

3m read· AgentsCamp
Guide

n8n vs Dify: Which AI Workflow Platform? (2026)

Automation-first vs AI-native: n8n's 400+ integrations and agent nodes vs Dify's LLM-app platform with built-in RAG. Licenses, pricing, and the fit test.

2m read· AgentsCamp
Guide

OpenAI Agents SDK vs LangGraph: Minimal vs Controllable (2026)

OpenAI Agents SDK's three-primitive minimalism vs LangGraph's explicit graph and durable state — which agent framework matches your reliability bar in 2026.

3m read· AgentsCamp
Guide

Which Agent Framework in 2026? LangGraph vs CrewAI vs AutoGen vs OpenAI Agents SDK vs Claude Agent SDK

A decision guide to the major AI agent frameworks — control vs. abstraction, multi-agent models, state and durability, and which fits your project.

2m read· AgentsCamp
Guide

Agent Memory Architecture: Short-Term, Long-Term, and When to Use Each

How AI agents remember — working memory vs. persistent long-term memory, what to store, how to retrieve it, and how to keep context small.

2m read· AgentsCamp
Guide

Agentic RAG: When Retrieval Needs an Agent in the Loop

What agentic RAG is — retrieval as a tool an agent uses iteratively, with query planning, self-correction, and multi-source routing — and when the upgrade pays.

3m read· AgentsCamp
Guide

AI Coding Statistics 2026: The Numbers That Are Actually Sourced

How much code AI writes, who uses the tools, and what it does to quality — every statistic dated and traced to its primary source, updated on a cadence.

3m read· AgentsCamp
Guide

How Computer-Use Agents Work

Inside the perception-action loop that lets AI operate real software — screenshots in, clicks out — plus grounding, reliability, and when to use APIs instead.

3m read· AgentsCamp
Guide

Production Tool & Function Calling: Feed Errors Back as Observations

How agents use tools — the call/observe/retry loop, why errors must return to the model, and the schemas, idempotency, and limits that keep it reliable.

3m read· AgentsCamp
Guide

Getting Web Data into AI Agents: Search & Scraping APIs Compared

The agent web-data layer — Exa for semantic search, Firecrawl for extraction at scale, Tavily for all-in-one, Jina Reader for zero-setup — and how they compose.

2m read· AgentsCamp
Guide

The AI Engineer Roadmap for 2026

A staged path from API calls to production agents — the skills that matter in 2026, what to skip, and the guides and tools for each stage, in order.

2m read· AgentsCamp
Guide

The Agent Skills Standard: One SKILL.md for Every AI Tool

Agent Skills became an open standard in December 2025. Which tools read SKILL.md today — Copilot, Cursor, VS Code, Gemini CLI, Codex — and how to write portable skills.

4m read· AgentsCamp
Guide

Claude Skills on claude.ai and the API

How Agent Skills work beyond Claude Code: uploading to claude.ai, the /v1/skills API with code execution, Managed Agents, and the Agent SDK.

4m read· AgentsCamp
Guide

What Are Claude Skills? The Complete Guide

Claude Skills explained: what a SKILL.md is, how progressive disclosure keeps skills cheap, where they run, and how to install or write your own.

5m read· AgentsCamp
Guide

Why Your Agent Loops: Debugging AI Agents

The recurring agent failure modes — loops, premature victory, tool misuse, context poisoning, scope creep — diagnosed by their signatures, with fixes.

2m read· AgentsCamp
Guide

Realtime Voice Agents: Build on LiveKit, Buy Vapi, or Pipeline with Pipecat

The three ways to ship a realtime voice agent in 2026 — open infrastructure, managed platform, or OSS pipeline — and how speech-to-speech models fit in.

2m read· AgentsCamp
Guide

Human-in-the-Loop AI Workflows: Approval Gates That Keep Agents Safe and Trusted

How to design human-in-the-loop into agent workflows — when to require approval, gate patterns, confidence escalation, review UX, and feedback loops.

6m read· AgentsCamp
Tool

AgentOps

Observability for AI agents — session replay, cost and latency tracking, and debugging for multi-step runs.

freemiumobservability
Tool

Agno

A Python framework for building multi-agent systems with memory, knowledge, and tools — formerly Phidata, now with the AgentOS runtime.

open sourcesdk
Tool

Augment Code

AI coding assistant built for large, real-world codebases — a Context Engine that indexes the whole repo, with agents, chat, and completions in IDEs and a CLI.

freemiumextension
Tool

AutoGen (AG2)

A multi-agent conversation framework where agents collaborate via message-passing, with group chat and code execution.

open sourcesdk
Tool

Browser Use

The most-adopted open-source browser-agent framework — point an LLM at a task and it drives a real browser: navigating, clicking, typing, extracting.

open sourcesdk
Tool

Browserbase

Managed headless-browser infrastructure for AI agents and web automation — serverless cloud browsers with stealth, proxies, live view, and Playwright/Stagehand.

freemiumplatform
Tool

CrewAI

A Python framework for orchestrating role-playing AI agents as collaborating 'crews', plus event-driven flows.

open sourcesdk
Tool

Daytona

Sub-90ms agent sandboxes — isolated computers with snapshots, volumes, Git and LSP tools, on Linux, Windows, or Android; AGPL self-host or managed cloud.

freemiumplatform
Tool

Dify

The visual platform for LLM apps and agentic workflows — canvas-built chatflows, RAG pipeline, agent nodes with 50+ tools, and LLMOps, self-hosted via Docker.

freemiumplatform
Tool

E2b

Open-source Firecracker-microVM sandboxes where AI agents safely run untrusted code — stateful interpreters, full Linux, pause/resume, desktop VMs.

freemiumplatform
Tool

Factory

Factory is an agent-native software development platform whose Droids plan, write, test, and ship code from the terminal, IDE, and web with org context.

paidagent
Tool

Flowise

Open-source, low-code visual builder for LLM apps and AI agents — drag-and-drop to assemble chains, agents, and RAG; self-host or use Flowise Cloud.

open sourceplatform
Tool

Google ADK

Google's open-source, code-first framework to build, evaluate, and deploy AI agents. Python and Java, model-agnostic, deploys to Vertex AI Agent Engine.

open sourcesdk
Tool

Jina Reader

Prepend r.jina.ai/ to any URL and get LLM-ready markdown — JS rendering, PDFs and Office docs, image captioning, and s.jina.ai for read-the-results search.

freemiumplatform
Tool

Kilo Code

Open-source AI coding agent extension for VS Code and JetBrains, built as a superset of Roo Code and Cline, with bring-your-own-key and zero model markup.

freemiumextension
Tool

Langchain

The provider-agnostic agent framework, post-1.0: a standard create_agent loop on the LangGraph runtime, middleware hooks, and the largest integration ecosystem.

open sourcesdk
Tool

LangGraph

A low-level library for building stateful, controllable agents as graphs, with checkpointing and human-in-the-loop.

open sourcesdk
Tool

Letta

Stateful agents from the MemGPT creators — an Apache-2.0 server with self-editing memory, and Letta Code, the memory-first model-agnostic coding harness.

freemiumplatform
Tool

Mastra

An open-source TypeScript framework for building AI agents, workflows, RAG, and tool-calling, with memory, model routing, and built-in observability.

open sourcesdk
Tool

Mem0

A memory layer for AI agents and apps — persistent, personalized long-term memory across sessions.

open sourcesdk
Tool

Modal

Serverless AI infrastructure in pure Python — GPU functions with sub-second cold starts, secure sandboxes for agent code, batch jobs, and per-second billing.

freemiumplatform
Tool

N8n

Fair-code workflow automation with native AI: a visual canvas plus code, 400+ integrations, and LangChain agent nodes; self-host free or cloud per-run.

freemiumplatform
Tool

OpenAI Agents SDK

OpenAI's lightweight, open-source framework for agents — handoffs, guardrails, sessions, and built-in tracing.

open sourcesdk
Tool

OpenHands

Open-source autonomous AI software-development agent (formerly OpenDevin) — writes code, runs commands, and browses the web in a sandbox.

open sourceagent
Tool

Pydantic AI

Type-safe agent framework from the Pydantic team: validated structured outputs, dependency injection, durable execution, 'that FastAPI feeling.'

open sourcesdk
Tool

Skyvern

Open-source vision + LLM browser automation aimed at replacing brittle RPA — workflow builder, CAPTCHA/2FA handling, and self-host or cloud.

open sourceplatform
Tool

Stagehand

Browserbase's open-source SDK for browser agents — act, extract, observe, and agent primitives that mix natural language with code-level control.

open sourcesdk
Tool

Swe Agent

Open-source autonomous coding agent from Princeton/Stanford that turns an LLM into a software engineer to fix real GitHub issues.

open sourceagent
Tool

Tavily

The web-access layer for agents — Search, Extract, Crawl, Map, and Research APIs purpose-built for LLMs, behind one key, with a hosted MCP server.

freemiumplatform
Tool

Trae

Trae is an AI-native IDE from ByteDance — a VS Code-style editor with a built-in Builder agent and an autonomous SOLO mode that writes code across a project.

freemiumide
Tool

Vercel Sandbox

Ephemeral Firecracker microVMs on Vercel for untrusted and AI-generated code — millisecond startup, Node and Python runtimes, persistent by default.

freemiumplatform
Tool

Zep

Agent memory on temporal knowledge graphs — Zep Cloud for sub-200ms context retrieval at enterprise scale, with Graphiti as its open-source graph engine.

freemiumplatform
Command

Add Human Approval Step

Scaffold a human-in-the-loop approval gate into an agent so it pauses before a consequential action and resumes after approval.

/add-human-approval<the action/tool to gate, or the agent file>
Command

Create Subagent

Scaffold a new Claude Code subagent definition file into .claude/agents/ with a routing-ready description, scoped tools, and a system prompt.

/create-subagent<what the subagent should do>
Term

A2A (Agent2Agent Protocol)

A2A is an open protocol that lets AI agents discover each other's capabilities and delegate tasks across vendors, complementing MCP's tool connections.

Term

Agent Engineering

Agent engineering is the discipline of building reliable AI agents — designing the tools, context, guardrails, evals, and recovery paths around the model.

Term

Agent Harness

An agent harness is the system around the model that makes it an agent — the loop, tools, context management, permissions, and recovery machinery.

Term

Agent Memory

Agent memory is how an AI agent retains information beyond its context window — working state during a task and persistent knowledge across sessions.

Term

Agent Skills

Agent Skills are reusable procedures packaged as folders with a SKILL.md file — loaded by an AI agent on demand when a task matches, now an open standard.

Term

Agentic AI

Agentic AI is the class of AI systems that act toward goals — planning, calling tools, and iterating on results — rather than only generating content.

Term

AI Agent

An AI agent is an LLM-driven system that pursues a goal in a loop — calling tools, observing results, iterating — instead of returning one answer.

Term

Computer Use

Computer use is an AI agent operating software through its real interface — reading the screen, moving the cursor, clicking, and typing like a person would.

Term

Function Calling (Tool Calling)

Function calling lets an LLM request structured invocations of your code: describe tools with schemas, the model emits typed calls, your app executes them.

Term

Human-in-the-Loop (HITL)

Human-in-the-loop design inserts human judgment at decisive points in an AI workflow — approving actions, resolving ambiguity, owning the irreversible steps.

Term

ReAct (Reasoning + Acting)

ReAct is an agent loop that interleaves reasoning with tool actions — Thought, Action, Observation, repeat — so the model plans, calls a tool, and revises.