Advancing the Ball on AI: How the New York Jets Are Building Trust Into Every Deployment
Every organization says it’s building AI agents. Almost none are scaling them reliably in production. Barr Moses opened our recent webinar with research bearing that out: across hundreds of data and AI leaders, 46% have agents in full production, 64% say they deployed faster than their teams were ready to support, and 54% expect to significantly rebuild what they’ve already shipped. The technology is moving faster than most organizations’ ability to trust it — and that gap is exactly what’s standing between the agentic use cases everyone’s building and the ROI everyone’s on the hook to deliver.
Monte Carlo’s answer is to treat agent trust as a systems problem covering four layers: context (is the underlying data fresh, complete, and correct?), performance (is the agent fast and reliable enough to be useful?), behavior (what is the agent actually doing, and why?), and output (is the final answer correct and grounded?). As Barr put it, dashboards have been wrong for decades, but agents are wrong confidently — they’ll deliver a bad answer with total conviction. That framing set up the real subject of the conversation: how the Jets are living it.
Iwao Fusillo is the Jets’ first Chief Data and Analytics Officer, and he’s been given a mandate that would make most data leaders sweat: build the best AI decision-support infrastructure in all of sports. No peer to call, no playbook to copy — just figuring it out in real time.
Inside the Jets’ AI operation
Fusillo’s CDAO office is unusual on purpose. It’s the only role in the league that spans both football and business analytics, includes application development, and carries an explicit AI mandate under one leader. The Jets’ relationship with Monte Carlo predates his arrival — the team started with data observability back in 2024. When Fusillo joined earlier this year, he reframed that foundation as something bigger: not a hygiene project, but the trust layer underneath the entire AI program.
He runs the organization through a three-horizon operating model:
- Horizon 1 — AI Fluency. Copilot, org-wide, with 97% front-office adoption. Fusillo’s advice to every team he talks to: don’t skip this step. The people using AI daily surface the highest-value use cases for Horizon 2.
- Horizon 2 — Innovation & Workflow Automation. Repeatable, double-digit productivity gains, plus entirely new use cases that weren’t possible before AI.
- Horizon 3 — Hybrid Workforce. Teams of humans and AI agents delivering 2-3x productivity out of the gate today, with the Jets engineering toward 20-30x.
The headline number: at the end of Q1, the Jets had 7 AI deployments. By the webinar, that had grown to 51 — a sevenfold increase in about three months, across all eight major football and business functions. “Deployments” is a deliberately broad term here: a mix of autonomous agents, AI models, and AI-built applications, with the agentic share growing quickly.
A few he highlighted:
- An AI sales coach that scores hundreds of thousands of annual ticket-sales calls against a rubric built by the Jets’ own sales managers — producing team-wide insights for Monday meetings and rep-level coaching at a scale no human manager could match.
- A fully autonomous media intelligence agent, stood up by the communications team itself, delivering a daily report to Jets executives before any news cycle hits — and feeding a shared repository other marketing agents mine for campaign moments.
- A Combine medical-dictation agent that turned four weeks of manual analysis of unstructured player medical transcripts into two days — built under real time pressure, three weeks before the event.
Every one of the 51 deployments has a named owner, defined performance metrics, and an explicit autonomy level — full autonomy or human-in-the-loop — decided deliberately rather than by default.
What actually breaks when you scale
Asked what broke first as the Jets scaled from 7 deployments to dozens, Fusillo didn’t hesitate: quality — and it broke silently. Early on, with humans tightly in the loop, complaints were mostly about speed and latency. As more agents came online, failures shifted to things like an unrefreshed data pipeline or a schema change nobody flagged — problems an agent will never volunteer on its own.
The Jets’ deployment discipline is staged on purpose:
- Start as a prompt, not an agent — sometimes for days, sometimes for months — building in context, checking for bias, defining constraints, and requiring the model to self-evaluate before anyone trusts it with autonomy.
- Run it manually first, with a human reviewing every output, to catch unstable data sources or drift before automating on top of them.
- Deploy the agent, keeping a human in the loop for anything high-visibility or high-impact.
- Selectively move to “loop engineering” — agents that iterate the way a person would, trading a little completeness for far lower latency and cost.
Even the executive-facing media agent — fully autonomous by design — still has humans reviewing its output before it reaches leadership, with a second and third model layered in to check the first.
Data trust and agent trust are the same problem
Asked whether data trust and agent trust are separate problems, sequential, or the same thing, Fusillo was unambiguous: they’re the same problem, amplifying each other in both directions. Bad data quality gets amplified by the agent built on top of it — and a well-designed agent can just as easily be turned around to find and fix data quality issues no human had time to chase down. His example: an agent tracked down months of archived call transcripts sitting with a third-party vendor, a job his data engineering team had estimated at five months of manual work, solved in minutes.
Governance is the thing keeping him up at night
With 51 deployments and climbing, Fusillo, his head of software engineering, his football data science lead, and his business analytics lead currently certify every deployment by hand. It works today. He’s confident it won’t work at 100, and certain it won’t at 300 or 400. A model registry with some form of self-certification is the likely next step — though what’s commonplace at a tech company isn’t necessarily the right structure for a sports franchise, and governance patterns from further-along industries like retail, CPG, and financial services are more useful reference points.
Trust, in Fusillo’s framing, isn’t a fourth pillar bolted onto the three-horizon model — it’s underneath all three of them, from the front-office employee trusting Copilot to take meeting notes, to the multi-agent architectures the Jets are just beginning to build in Horizon 3.
Watch the full conversation here:
Our promise: we will show you the product.