Skip to content

Catch the failures your evals can’t see

Agents don't crash, but they will degrade. Bad context, hallucinated outputs, and silent behavioral drift all return a clean 200.

Find the root cause in minutes

Go from a bad output to the exact span that produced it — and the source data behind it.

Scale agents while keeping MTTR low

See token spend, latency, and error patterns in aggregate and with automated tooling, so adding agents doesn't mean more engineering load.

Works with your entire agent stack

Your agents don't live in one place, and neither does the risk. Monte Carlo plugs into every layer your agents depend on — the models they call, the frameworks they run on, the telemetry they emit, and the data and context they pull from — with native integrations rather than glue code you have to maintain.

That coverage is what makes the rest possible: you get lineage from raw data in the warehouse all the way to the action an agent takes, so when something breaks you can trace it back to the exact pipeline failure, model change, or context drift that caused it.

Monte Carlo Integrations
Critical features

Everything you need to trust the agents you put in front of people

Evals get you to launch. Monte Carlo gets you to scale.

Add monitor

Browse Ask AI

Metric

Detect unexpected changes in metrics

Custom SQL

Monitor your data with flexible SQL

Query performance

Set expectations for run time of queries

Agent observability

Agent metric Performance

Track latency, token usage, duration and error rates

Agent trajectory Behavior

Track expected agent path and steps

Agent evaluation Outputs

Monitor agent output quality

Agent validation

Identify individual traces or spans with quality issues

Not sure where to start?

Describe what you want to monitor — AI will set it up.

Ask AI

Agent metric monitor examples

Max of total_tokens

is > 1,200

Mean of duration

is > 5s

Null (%) of email_address

is > 0.5%

Agents/ai-agent/a726b62e…12983
Explain this trace + Add monitor
Status OK Duration 46m16s Spans 419 Error spans 8 Tokens 554,216 Models Claude Sonnet 4.6
grab_write_queries.task674ms
genai.write_queries597ms
run_data_correlation_rca.task30,846ms
genai.sampling_result8,671ms
RunnableSequence.task5,836ms
ChatBedrockWithFallback.chat5,833ms
run_data_exploration_agent.task2,714,078ms
data_exploration_agent.task2,713,621ms
model.task7,985ms
ChatBedrockWithFallback.chat2,164ms
tools.task10,001ms
get_table_schema.tool9,968ms
genai.warehouse_query9,959ms
model.task4,475ms
run_warehouse_query.tool9,959ms
Span data_exploration_agent.task

Error

botocore.exceptions.ReadTimeoutError: Read timeout on endpoint URL: “https://bedrock-runtime.us-east-1.amazonaws.com/model/…/converse”

Details Metadata
StatusError
Duration2,713,621ms
Agentai-agent
WorkflowTroubleshooting Agent
Taskrun_data_exploration_agent
Start timeJul 16, 2026 · 19:23:56 PDT
Span IDc72012959428a6d8
Monitors/Total tokens 80th percentile is anomalously high
Details Results Alerts Run history Activity

Field: total_tokens 80th percentile

warehouse_main:telemetry.agent_otel_traces

Date

Past 3 weeks

total_tokens

80th percentile · itinerary-agent

80th percentile Threshold
01k2k3k4k5k6k 6.539k Jul 28Aug 02Aug 05Aug 09Aug 13Aug 17 ingest_ts
Customer story

How Axios ships reliable AI agents in production

The challenge

Axios's AI team set out to remove friction from the newsroom's most tedious work, freeing journalists to focus on reporting. As they scaled past a dozen LLM-powered agents, they wanted the same confidence in their agents that they already had in their data.

The solution

Agent Observability, instrumented in just a few lines of code via Monte Carlo's OpenTelemetry SDK. Axios now sees prompt and response traces, token usage, latency, and evaluation scores over time — right alongside the data and ML monitoring they already rely on, in one familiar interface.

Axios

“We were able to get started with two lines of code.”
Read the full case study

Ship agents. Not fire drills.

Connect in minutes, start monitoring out of the box, and scale coverage as your agent fleet grows.

Trusted by 400+ enterprises

T. Rowe Price PepsiCo Cisco Comcast Nasdaq Disney Gap Highmark Target Salesforce

G2 names Monte Carlo as #1 leader for the 13th consecutive quarter

X