We made the case that AI agents need their own observability and governance layer, separate from the platforms that build and run them. Infrastructure monitoring watches machines. Platform admin consoles watch their own agents and grade their own homework. Yet still, neither one can tell you how many agents your company runs, who owns them, what they cost, or whether they still do the job they were built for.
So the central question shifts. If agents need a supervision layer of their own, what should that layer actually do?
The answer comes down to three capabilities enterprises should expect from any serious approach:
1. discover what's running
2. monitor how each agent behaves, and
3. manage the risk each one carries
Dataiku Agent Management, which reaches general availability this October, is built to that shape. Here's what each capability means in practice.

Enterprises are good at managing software they buy. There's a contract, an owner, a renewal date. AI agents work differently. A team spins one up inside a platform it already licenses, in an afternoon, with no purchase order and nothing in the normal machinery of IT to notice. Multiply that across Copilot Studio, Agentforce, Bedrock, Vertex, Databricks, Snowflake, and n8n, and the definitive list stops existing. Only about 18% of organizations keep a current, complete AI inventory, according to IBM research this year.
A supervision layer should close that gap by connecting to the platforms where teams actually build, then scanning them into one inventory on a schedule. Each agent arrives with an owner, a stated purpose, and metadata that maps it to the business. The first scan tends to be the moment that lands, because agents the central team never knew existed show up on the list. From there, you can facet by platform, team, or business process. The board asks how many agents we run, and the answer becomes a filter rather than a three-week project.
Discovery counts agents, but it doesn't tell you whether they work. That takes monitoring built around the agent itself versus the server it happens to run on.
An agent rarely crashes when it goes wrong. Instead, it keeps answering, slightly off, while the error compounds and nobody reads the ten thousand conversation logs that would reveal it. Quality drifts when a model updates or when users start asking an agent things it was never built for. So the monitoring layer tracks operational health hourly, counts real adoption from actual usage traces, turns token consumption into a cost figure, and scores quality against each agent's purpose. It also watches for behavioral drift in the tools an agent calls and the topics people bring it.
Picture an HR policy agent that starts fielding payroll questions. Its success rate sags. Topic modeling surfaces the new cluster and explains the drop, and an alert reaches the owner before the first complaint does. The owner decides what to do next. The AI agent with three users and the agent with three thousand finally look different from each other.
Most agents are small, useful, and low-stakes. A handful are customer-facing, touch sensitive data, or act on real transactions. And those few need the full treatment.
Managing agent risk means certifying each agent against a registry of named risks you can extend, things like over-privileged access, shared credentials, or no human override. Each risk carries the mitigations it requires: monitoring, alerting, and recurring tests against a golden dataset. An agent that passed in March gets re-checked in April. When an auditor calls, certification status, test history, and alert history export per agent instead of triggering a scramble across five teams.
This layer observes, measures, and records. It does not run, filter, or guardrail agents. That means work stays with the platforms agents run on, and for Dataiku-orchestrated agents, with the LLM Mesh. The control a CIO is accountable for lives at the company level. It decides which agents may run, knows what each is entitled to do, verifies they still behave as intended, and keeps the record when they don't.
Discover, monitor, manage risk – three capabilities, one portfolio view, and drill-down from the whole population to any single agent. That's what a supervision layer should deliver. It's why we built Dataiku Agent Management to sit outside the platforms it watches. Every agent is treated the same, whichever platform it came from.
Learn more about agent management
Let's go