Observability and evaluation tools for agents
Trace agent runs, evaluate outcomes, and diagnose failures in production.
Compare providers
44 agent-ready tools mapped to this capability.
AgentHaus
Visual agent builder that exports production-ready TypeScript
AgentOpsMCP
Session replay, monitoring, and evaluation for AI agents
AgentuityMCPCLI
Full-stack cloud platform built for AI agents
Arga LabsMCPCLI
Stateful service twins for testing and training AI agents
Arize PhoenixMCPCLI
Open-source tracing and evaluation for LLM applications
AxiomCLI
Cloud-native logs, traces, and event data for agent systems
Baserun
Observability, testing, and evaluations for LLM applications
BentoLabs AI
Monitoring and learning for long-running AI agents
BraintrustMCPCLI
Evaluation, tracing, and prompt management for AI products
Confident AIMCPCLI
Evaluation and observability platform for reliable AI agents
ContextForgeMCPCLI
Open-source MCP, A2A, and API gateway for AI agents
CovalMCPCLI
Simulation, evaluation, and monitoring for voice and chat agents
Credal.aiMCPCLI
Security and governance control plane for enterprise agents
DAGWorksCLI
Open-source frameworks for reliable dataflows and agents
dari.devCLI
Stateful model routing for applications and coding agents
Galileo
AI observability and evaluations for reliable production agents
GolfMCPCLI
Security and governance infrastructure for MCP tools
Guardrails AICLI
Open-source validation framework for safer, structured LLM applications
HatchetCLI
Durable task orchestration for agents and background jobs
HeliconeMCP
Open-source AI gateway and observability platform
HUDMCPCLI
RL environments and scalable evaluations for AI agents
HumanLayerCLI
Multiplayer IDE and cloud control plane for coding agents
InngestMCPCLI
Durable workflows and background jobs for AI agents
Kestrel AIMCPCLI
Deterministic platform-engineering workflows built and operated by AI agents
LaminarMCPCLI
Tracing, evaluations, and debugging for AI agents
LangfuseMCPCLI
Open-source LLM engineering platform: traces, evals, and prompts
LangSmith
Framework-agnostic tracing, evaluation, and monitoring for agents
LemmaMCP
Production tracing and debugging for AI agents
LiteLLMMCPCLI
Open-source gateway for calling 100+ LLM APIs
MastraMCPCLI
TypeScript framework for building and operating AI agents
OmnaraCLI
Control plane for supervising coding agents from anywhere
OpenlayerCLI
Evaluation, observability, and governance for production AI
OpenPipe
Fine-tuning and model optimization for production agents
Parea
Evaluation and observability for LLM applications and agents
PortkeyMCPCLI
AI Gateway for 1600+ LLMs — observability, guardrails, and governance in one platform
PostHogMCPCLI
Product analytics and observability for agent experiences
PromptfooCLI
Open-source evaluations and red teaming for LLM applications and agents
Respan
AI gateway, evaluations, and observability for agents
Runloop
Persistent development sandboxes and benchmarks for coding agents
SentryMCPCLI
Full-stack error, performance, and AI agent monitoring
Superagent
Security scans, guardrails, and red teaming for AI-native software
TensorZeroCLI
Open-source gateway, observability, and optimization stack for LLM applications
Traceloop
OpenTelemetry observability for LLM and agent applications
Trigger.devMCPCLI
Durable background jobs and workflows for production AI agents