Skip to content
AI Observability Updated Jul 15 2026

How Smartsheet’s Data, AI, & Platform Engineering teams use Monte Carlo to catch issues before they reach the business

How Smartsheet’s Data, AI, & Platform Engineering teams use Monte Carlo to catch issues before they reach the business
AUTHOR | Virna Sekuj

Overview

Smartsheet, the enterprise platform for modern work management, serves over 85% of Fortune 500 enterprises. It also runs a data and AI platform that powers internal analytics, ML models, and product intelligence at scale. As the volume and complexity of Smartsheet’s data estate grew, so did the cost of data downtime: silent pipeline failures, freshness gaps, and schema drift that reached downstream dashboards before anyone noticed.

Monte Carlo’s agent trust platform became a central pillar of Smartsheet’s reliability strategy by helping the team implement proactive, automated observability across data and AI.

The challenge

Before Monte Carlo, Smartsheet’s data engineering teams faced a set of compounding reliability problems common to any organization scaling its data estate aggressively:

  • Data issues reaching dashboards and business decisions before engineers could catch them
  • Manual, custom unit tests that were expensive to write and slow to catch schema changes or volume drops
  • Incident ownership spread across Slack threads with no structured workflow for assignment or resolution tracking
  • A growing blind spot in agentic and ML workflows, where LLM interactions were a “black box” for validation teams
  • Alert fatigue from misconfigured or overly sensitive monitors during initial setup phases

As Dharmendra D., Senior Software Engineer on Smartsheet’s Data & AI Platform team, described it: teams were handling incident coordination “in a much messier way across Slack threads” — with no clear ownership or resolution tracking to prevent the same issues from resurfacing.

The solution: Monte Carlo’s Agent Trust Platform

Smartsheet deployed Monte Carlo as an end-to-end observability layer across their data and AI stack, integrating tightly with Databricks, Snowflake, Looker, PagerDuty, and Slack. The platform’s automated ML-driven monitoring, field-level lineage, and incident management workflows replaced the manual, fragmented approach that had previously left teams reactive rather than proactive.

Proactive anomaly detection before issues reach the business

Monte Carlo’s ML models learn baseline behavior for each data asset and flag deviations — freshness drops, volume anomalies, schema changes — before downstream consumers notice.

Rather than flooding teams with noise, the system surfaces only meaningful anomalies, “significantly reducing alert fatigue and helping our team focus on real issues rather than chasing false positives.”

Seamless integration with the modern data and AI stack

Smartsheet’s engineers operate across a multi-tool environment spanning Databricks, Snowflake, Looker, PagerDuty, and Slack. Monte Carlo’s broad integration surface made centralized observability possible without disrupting existing workflows.

“I love how easily Monte Carlo integrates with Databricks to automatically catch anomalies in our pipelines. Instead of writing endless custom unit tests for schema changes or volume drops, the automated ML alerts catch data downtime instantly, saving our engineering team hours of manual troubleshooting every week.”

Ruchir K., Software Engineer-2, Enterprise, Smartsheet

End-to-end lineage that cuts debugging time

Lineage visualization across data and AI assets has been a recurring need at Smartsheets, with engineers describing it as the feature that most directly translated to hours saved. The ability to trace an issue from a Looker dashboard all the way back to a Snowflake warehouse, or from a pipeline failure to its upstream source, eliminated the manual investigation that previously consumed debugging cycles. Being able to trace data from source to consumption in a clean, interactive graph saved Smartsheet engineers hours of investigation during incidents. 

“The platform’s machine learning-driven alerting is incredibly smart; it quickly learns our data’s baseline behavior and catches anomalies, freshness issues, or volume drops before our downstream users even notice. The user interface is highly intuitive, making it easy to trace an issue from a Looker dashboard all the way back to our Snowflake warehouse. It has saved our data engineering team countless hours of manual debugging.”

Vandan T., Associate Software Engineer, Smartsheet

Structured Incident Management Replacing Ad-Hoc Slack Coordination

One of the most operationally significant improvements across Smartsheet’s teams was the shift from informal incident handling to structured, ownership-driven workflows. Monte Carlo’s incident management module brought clear assignment, severity classification, and resolution tracking to what had previously been a coordination problem.

“The incident management workflow is a highlight as well,” noted Dharmendra D. “It keeps the team aligned on data quality issues with clear ownership and resolution tracking — something we previously handled in a much messier way across Slack threads.” The result: fewer escalations and faster resolution across Smartsheet’s Data & AI Platform.

On ROI: “For a platform team, the ROI shows up as fewer escalations and faster incident resolution.” The time saved debugging incidents, the reduction in manual monitoring effort, and improved organizational trust in data all compounded into measurable returns.

Agent observability for emerging AI workloads

Smartsheet engineers are seeing the value of Monte Carlo beyond traditional data pipelines and into the agentic layer — a particularly resonant point given Smartsheet’s active AI and ML development.

Ruchir K. described the team’s challenge: “In terms of Agent Observability, LLM interactions can be a bit of a black box for validation teams. We implemented an internal judge system for LLM-based projects, but Monte Carlo has also helped us get the big picture on how well our models are performing.” Monte Carlo provided the visibility layer where internal tooling fell short.

For teams managing multiple data squads under a larger analytics function — like Smartsheet’s — Monte Carlo also simplified coverage tracking: “One of our ongoing challenges has been making sure all the different teams have proper coverage for our IP. We have a lot of squads under Analytics, and this has helped us keep the process moving so we can consistently ensure our products are covered appropriately.”

“As Smartsheet scales its AI initiatives, the reliability of our agents and their underlying data is a first-order engineering concern. Monte Carlo gives our teams end-to-end visibility into the entire system, from the behavior, output, and performance of the agents we are deploying, all the way through to the data infrastructure that powers them  — so we can move fast and still trust what we’re shipping.”

Kapil Ashar, VP, Engineering at Smartsheet

Results

Across Smartsheet deployments, Monte Carlo delivered on three fronts: engineering efficiency (eliminating manual unit test writing, cutting debugging time through field-level lineage, and freeing teams for higher-order work), data and AI reliability (catching issues before they reached dashboards, reducing alert fatigue through ML-calibrated detection, and building cross-organizational trust in data and agents), and operational maturity (replacing ad-hoc Slack coordination with structured incident ownership, clear resolution tracking, and scalable multi-team coverage).

About Smartsheet

Smartsheet (NYSE: SMAR) is the enterprise platform for modern work management, helping organizations plan, capture, manage, automate, and report on work. Over 85% of Fortune 500 companies trust Smartsheet. Headquartered in Bellevue, WA.

Our promise: we will show you the product.

Recommended for you