Agent OBSERVABILITY and Evaluation

Production AI starts with observability.

HoneyHive is the agent observability platform behind mission-critical AI — how enterprises observe, evaluate, and trust the agents they run in production.

TRUSTED BY
A new primitive

Agents drift, regress, and fail silently. HoneyHive is how you observe agent behavior, evaluate every output, and improve quality with every release.

See what your agent actually did.

  • Inspect every trace

    Replay any run step by step: every tool call, prompt, and decision in the order it happened. Jump straight to the span where behavior broke.

  • Debug long running agents

    Follow trajectories that spans hours or days: retries, loops, and handoffs between sub-agents. Find the step where it started going wrong.

  • Monitor quality in production

    Evaluate live traffic and connect scores with real user outcomes, so quality is measured against production, not your test suites.

  • Alert before users notice

    Set thresholds on any score, detect drift, and get notified when one breaks, by email or webhook.

Support Agent Trace
42 spans · 3.24s
100%
Support Agent Trace 3.24s
Process Query 3.21s
Analyze Intent 0.44s
Retrieve Context 1.12s
embeddings.create 0.20s
pinecone.query 0.28s
Generate Answer 1.63s
ChatCompletion 1.61s
guardrail.check 0.02s
Thread Raw JSON gpt-4o · 1,284 → 96 tok
SYSTEM
You are Acme Support. Always confirm order status with the lookup tool before you answer. Cite the knowledge-base article you used. Never disclose internal pricing or refund thresholds.
USER
Where is order #48213? It still says processing after four days.
ASSISTANT
order_lookup({ id: "48213" })
Your order left our Reno warehouse Tuesday evening and is out for delivery Thursday. Tracking is 1Z9Y4482 — I've emailed the link to the address on file.
OpenTelemetry
Agent Frameworks
Model Providers
Coding Agents
No-Code Platforms
research_pipeline
agent_run_7f3 · 24 of 214 steps
on track elevated anomaly
Model Reasoning quality higher is better
claude-sonnet-5
Tool Tool error rate lower is better
Order Lookup
Inventory Check
Ticket Creation
Chain Goal completion higher is better
Multi-Turn Session
Support Agent
User Turn
AGENT STEPS →
2 4 6 8 10 12 14 16 18 20 22 24
Trajectory View
Agent Graphs
Threads
Handoff Detection
Groundedness
research_agent v3
LIVE
0.9124h avg10% sampled
rolling avg · threshold 0.80 · 12.4k scored today
LIVE SCORES
session_9f4a0.94
session_9f490.88
session_9f480.61
session_9f470.90
session_9f460.93
session_9f450.85
session_9f440.92
session_9f430.89
session_9f420.96
session_9f410.87
session_9f400.91
session_9f48 scored 0.61 — below threshold
LLM Evaluators
Code Evaluators
Human Feedback
Intelligent Sampling
Quality anomaly detected
finance-copilot · 4m ago
P1
Groundedness
0.61↓ from 0.90
threshold 0.80
breached 4m ago
18:0020:0022:00now
Notified Email Webhook
Aggregate Thresholds
Drift Detection
Custom Webhooks

Turn production behavior into better releases.

  • Run experiments, spot regressions

    Compare agent versions on the same dataset and see exactly what improved, what regressed, and whether a change is ready to ship.

  • Build a dataset from real traces

    Automatically turn production failures into test suites. Test every future version against the cases that actually broke production.

  • Evaluate using LLMs or code

    Score LLM outputs, tool calls, A2A interactions, and full trajectories using LLM judges, custom code, or composite evaluators.

  • Bring experts into the loop

    Automatically route flagged traces to annotation queues where domain experts can judge what happened.

support-agent-v3 vs support-agent-v4 run 85d2da76 Completed
Resolution0.710.96
Checks2/55/5
Tone0.800.93
Latency2.1s3.4s
support-agent-v33/5 failed
support-agent-v45/5 passed
No resolution conf 0.62
Resolved + tracking conf 0.94
Sorry for the delay on order #48213.
Sorry for the delay on order #48213.
It looks like it's still processing on our end.
+It shipped Tuesday 7:40pm from our Reno warehouse.
No tracking number, no delivery estimate.
+Tracking 1Z9Y4482 — arriving Thursday by 8pm.
Answered without calling order_lookup.
+Cites KB §4.2 — Shipping delays.
Anything else I can help with?
Anything else I can help with?
Human review
3.0
Human review
5.0
Regression Detection
Drill-Down Analysis
Human Review Queues
Add to dataset
Sessions
5 selected
session_8f2a "refund eligibility?" 0.42
session_3b71 "cancel plan mid-cycle" 0.55
session_9c04 "export invoices to csv" 0.38
session_1d55 "upgrade seat count" 0.91
session_6e20 "refund window policy" 0.47
session_4a19 "duplicate charge dispute" 0.51
session_7f88 "change billing email" 0.88
coding-agents-v3
Evaluator library
LLM · Code · Human
12 active
Groundedness v3 AM
Toxicity v2 JD
schema_valid draft v1 SL
Refund policy adherence v2 TW
Helpfulness v5 RK
Citation coverage v1 MB
latency_budget v4 DP
Tone & brand voice draft v2 EN
LLM Evaluators
Code Evaluators
Human Criteria
Custom Evaluators
Low confidence escalations / queue_Qd93d 38 / 120
trace_8f2a41 low confidence
USER
Can I appeal a denied claim after 60 days?
AGENT
Appeals must be filed within 45 days of the denial notice, so a claim denied 60 days ago is outside the standard window. You can still request a late review if the delay was caused by a documented hardship — submit form A-12 with supporting records and a reviewer will respond within 10 business days.
REVIEW SL
Answer faithfulness
Every claim is supported by the retrieved policy documents.
4 / 5
Failure mode
Pick the category that best describes the miss.
Missing citation
Reviewer notes
What should a correct answer have said?
Cite the 45-day clause directly and link the hardship form
3 of 3 complete Submit
SL AM RK 3 reviewers κ 0.82
Numeric Rating
Binary Rating
Categorial Scores
Free-Form Text

Find the failure. Then fix it.

  • Root-cause the failure

    Investigate alerts with your coding agent using HoneyHive MCP, CLI, and purpose-built Skills.

  • Fix it at the source

    Make the fix, version the prompt, re-run the evals that caught it, and redeploy in the same pass.

  • Gate the release

    Run automated evals in CI on every release and block anything that regresses from reaching users.

  • Let agents run the loop

    Let your coding agents run the optimization loop with purpose-built Skills for every step.

Claude Codev2.1.97
Welcome back Mohak!
Opus 4.8 · Claude Pro
~/repos/finance-copilot
Tips for getting started
Ask Claude to root-cause an alert or triage failing traces…
Recent activity
honeyhive-docs MCP · connected
Root-cause what's causing this alert: https://app.us.honeyhive.ai/p/jh6vrIdH9w3TaZ_U_AH-lpyh/alerts/01KRF4E3Q2Q8PBRG9WJYV1WF3E
honeyhive · get_alert (MCP)
⌙ fetched 20 failing sessions · faithfulness < 0.6
Clustering spans across 20 traces
18/20 skip the refund-policy lookup before answering
* Writing root-cause report…
esc to interrupt
HoneyHive MCP
HoneyHive CLI
Purpose-Built Skills
system_prompt.md v4 v5
You are the refund support agent.
Answer refund questions for the user.
Decide eligibility yourself.
Assume annual plans are non-refundable.
+First call lookup_policy("refunds").
+Decide only from the returned policy text.
+Cite the clause; if none applies, say you're unsure.
+Escalate to a human when the policy is silent.
Keep the reply under 120 words.
re-runs eval suite
Prompt Management
Prompt Versioning
Playground
honeyhive Bot commented 1 hour ago
Support agent (HEAD-1714341466)
Eval suite passed — 64 improvements, 11 regressions across 5 scorers.
Score
Average
Improved
Regressed
Policy compliance
94%
(+4%)
12
2
Resolution quality
87%
(+6%)
18
5
Escalation accuracy
91%
(+3%)
9
1
Tool-call accuracy
89%
(+5%)
11
3
Groundedness
96%
(+2%)
14
0
GitHub Actions
PyTest
Vite
Alert
Root-cause
Fix
Re-eval
Optimization
iteration 3 · +18% quality
Agent Skills
HoneyHive CLI
HoneyHive MCP
Built for Platform Teams

Make reliability the default.

Give every team self-serve observability and evaluation, with the shared tooling, standards, and workflows to build and operate production agents across the enterprise.

Apply policy… ORG DEFAULT
EVALUATORS
Groundedness ≥ 0.90 RAG
Answer Faithfulness RAG
Tool Correctness Multi-Agent
Toxicity Safety
MONITORS
Token consumption per run
Faithfulness 7d trend
Cost $ / 1k runs
Auto-applied to every new workspace 14 workspaces
Org-wide Policies

Define evaluation and monitoring standards once and apply them consistently across every agent you trace in HoneyHive.

node < skills add honeyhiveai/skills
$ npx skills add honeyhiveai/skills --skill
Need to install the following packages:
skills@1.5.21
Ok to proceed? (y) y
skills
Source: https://github.com/honeyhiveai/skills.git
Repository cloned
Found 5 skills
Select skills to install
Search:
↑↓ move, space select, enter confirm
❯ ◉honeyhive-alert-root-cause
  ○honeyhive-cli
  ○honeyhive-evaluate
  ○honeyhive-improve
  ○honeyhive-instrument
Description
Decode a HoneyHive Discover URL, query the flagged sessions and their trace trees, classify each finding as a true or false positive with evidence, and recommend specific guardrails (hooks, evaluator changes, prompt additions) to prevent recurrence.
Agent Skills

Give developers reusable agent skills to instrument applications, set up evaluations, investigate failures, and more using coding agents.

relevance_llm_judge.yaml
evaluators/ · main
synced
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
name: relevance-llm
type: LLM
return_type: float
scale: 5
model_provider: openai
model_name: gpt-4o
sampling_percentage: 25
description: Rates how well the answer addresses the question.
criteria: |
  [Instruction]
  Rate the assistant's answer for relevance to the
  question on a scale of 1 to 5.
  [Question]
  {{ inputs.question }}
  [Answer]
  {{ outputs.content }}
$ honeyhive evaluators push Validated in CI
CLI and Config-as-Code

Manage evaluators, datasets, prompts, and alerts as code—version them in Git, automate workflows in CI, and make them accessible to coding agents.

Custom roles
org · workspace · project scopes
roles.yaml
Org Admin org
Workspace Admin ws
Evaluator Author proj
Data Labeler proj
Billing Viewer org
New role
DataLabeler:
  label: Data Labeler
  actor: user
  scope_type: project
  permissions:
    allow:
      - project.annotation_queue.update
      - project.events.get
      - project.alert.post
      - project.dataset.delete
      - project.templates.set
      - project.membership.remove
Validated · applied to 12 labelers
GROUPS ps_annotations_write ps_traces_read ps_queues_read ps_datasets_read
Granular RBAC

Isolate teams and define custom roles across dozens of fine-grained permissions, scoped to organization, workspace, or project.

Snowflake
Databricks
BigQuery
Amazon S3
Redshift
Kafka
HONEYHIVE
Streaming every 5 min OTLP · Parquet · JSONL
Data Exports

Export traces, evaluations, and metrics to your warehouse, lakehouse, or streaming platform in open formats.

One source of truth for everyone building agents.

HoneyHive is designed for the full cross-functional team: from the engineer instrumenting the first agent to the risk officer signing off the last deployment.

Get startedGet started
AI Platform Teams
Set the foundation.

Give every team a proven way to instrument, evaluate, and operate agents across frameworks, clouds, and business units.

AI Engineers
Debug, evaluate, ship.

Debug complete trajectories, test changes on real production cases, and catch regressions before customers do.

Domain experts
Turn expertise into evaluation.

Review real outputs, define rubrics, and make expert judgment reusable across every agent and release.

Risk & Governance
Produce audit-ready evidence.

See traces, evaluation results, and human reviews in one defensible record for internal and regulatory review.

Customer Spotlight

Scaling AI agents responsibly at Australia's largest bank

Learn MoreLearn More
17M
Retail consumers served by agents in production
55K
Internal users served by agents in production
HoneyHive powers observability and evaluation across dozens of mission-critical AI applications at CBA, enabling safe and responsible deployment of AI agents serving 17M+ consumers.
Financial Services
#4 on Evident AI Index
INTEGRATIONS

See every agent, wherever it runs.

OpenTelemetry-native and fully vendor-agnostic. Trace the agents your teams build in-house, the coding agents your developers use, and the agents you build in 3rd-party platforms.

  • 01

    Custom Agents

    MORE
    LESS
    +100

    Automatically instrument 100+ models and frameworks with our SDKs. Trace LLM outputs, tool calls and handoffs, score outcomes, compare releases, and monitor production behavior without changing how teams build.

  • 02

    Coding Agents

    MORE
    LESS

    Bring coding-agent sessions into the same observability standard. Track cost, latency, tool use, and outcomes across repositories, teams, models, and vendors.

  • 03

    No-Code Platforms

    MORE
    LESS

    Observe agents inside ServiceNow, Microsoft Copilot, Salesforce, and other enterprise systems. Centralize quality signals, approvals, and audit evidence across vendors.

IntegrationsIntegrations
RESOURCES
Start building with HoneyHive
Security

Your data stays under your control

Agent traces carry rich, highly sensitive I/O that traditional observability can’t handle. HoneyHive is designed specifically for AI traces and centralizes visibility without centralizing sensitive data, keeping every deployment isolated by design.

SAAS
Fully managed SaaS. Isolated by default.

HoneyHive operates both planes, with a virtual data plane isolated for each tenant.

HYBRID
Sensitive data never leaves your environment.

Store traces and run evals in your cloud. HoneyHive manages the rest.

SELF HOSTED
Fully self-hosted. Nothing leaves your network.

Both planes run in your infrastructure, deployed and managed through Kubernetes.

The same controls, every deployment
Granular RBAC

Define custom roles across dozens of fine-grained permissions, scoped from organization down to project.

SOC 2 · GDPR · HIPAA

Audited to SOC 2 Type II. GDPR-compliant with EU data residency. HIPAA BAA available for healthcare.

SSO & SAML

Okta, Azure AD, Google, PingSSO. JIT provisioning, enforced MFA, and session policies managed by your IdP.

Audit Logging

Stream audit logs to Splunk, Datadog, or any SIEM. Every access, change, and export is auditable upstream.

BLOG & CHANGELOG

What's new at HoneyHive

Guides
July 22, 2026
Responsible AI Playbook for Enterprise Agents
Responsible AI Playbook for Enterprise Agents

From input guardrails, trajectory evaluation to session outcomes: playbook for regulated teams buil...

READ
Insights
July 8, 2026
Standardizing AI Observability Before It Breaks: A Case Study on 73,000 Agent Schemas

Why AI observability needs a standard — and why we're betting on OTel GenAI.

READ
START FOR FREE

Production AI starts with observability.