Skip to content

How Monte Carlo’s Reinforcement Loop Caught a Silent Issue in Our Own Troubleshooting Agent

By Lior Gavish

Inside the Monte Carlo platform, our Troubleshooting Agent (TSA) works behind the scenes to analyze data and AI incidents, pinpointing root causes and suggesting fixes in real time. To keep TSA—and our other production agents—running efficiently, we rely on the Reinforcement Loop, an automated monitoring system designed to continuously evaluate agent performance, catch subtle inefficiencies, …

The Open vs. Closed AI Debate Misses the Point: Most Orgs Cannot Measure Either

By Lior Gavish

The discussion around enterprise AI often focuses on the choice between open-weight models, like Meta’s Llama and models from Mistral, vs. closed frontier models, such as Anthropic’s Claude and OpenAI’s GPT models.  While closed frontier models promise state-of-the-art reasoning without managing infrastructure,  open-weight models promise sovereignty, portability, and freedom from vendor lock-in. Public commentary has …

Self-improvement as Infrastructure: the Blueprints and the Labor of Agentic RL

By Lior Gavish

Most people imagine self-improving AI like a switch. You ship an agent, a smarter foundation model drops, and suddenly the system starts fixing its own code. Self-improvement arrives fully formed, baked into the weights. It’s unfortunately far more labor intensive than that. Self-improvement in agentic systems is a loop you build that requires a certain …

How to build an AI native engineering org: what we actually did

By Lior Gavish

In March we restructured Monte Carlo’s engineering organization. As I’ve thought about sharing our decision-making process, I’ve wanted to be far enough past the restructure to say something honest, objective, and which other engineering leaders can hopefully find useful. The short version is that our teams and operating principles weren’t broken, which made it particularly …

Stop Cleaning Up Your Clickstream Data. Let Claude Ship It Clean.

By Lior Gavish

Every product team has the same recurring nightmare. A PM opens Mixpanel to answer a simple question — how many users completed onboarding last week — and finds three events that could plausibly mean “completed onboarding”: Onboarding Complete, onboarding_finished, and Completed Setup. None of them is documented. Two of them stopped firing in March. The …

G2 names Monte Carlo as #1 leader for the 13th consecutive quarter

X