The AI Demo Always Works. Your Production Data Is the Real Test
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
20m ago - last edited 17m ago
Over the years I've worked with a lot of different clients on ServiceNow, and these days AI comes up in nearly every first conversation.
One engagement has stuck with me. It was a large enterprise that wanted AI in their ServiceNow environment. They had the enthusiasm and the budget. What they didn't have was any appetite for fixing the underlying data first. The thinking was that the AI would cope.
It didn't, and it wasn't the AI's fault. It used exactly what we gave it. A month or two in, it was clear this wasn't going to work the way anyone had hoped. The answers were only as good as the records behind them, and we had to go back and fix the foundations before the AI could be trusted.
I've been thinking about that project ever since, and it changed the first question I ask. It used to be "what can we turn on?" Now it's "what is this thing going to trust, and who owns that?"
The demo is never the problem
Demos are great. Now Assist summarises an incident in seconds, an agent suggests a resolver group, and everyone starts talking about go-live dates. I've sat in those rooms and enjoyed them.
Then the pilot meets real records. A suggested assignment group belongs to a CI that was retired ages ago. A summary repeats a workaround nobody has used for a year. Search returns three similar articles that disagree with each other. Nobody can quite say who should approve the next step.
None of that is new. Every service desk already has someone who knows the assignment group on that CI is wrong, or which of the three articles is the real one. That knowledge lives in people's heads and quietly covers for the data. AI doesn't have that person. It only has the records.
A filled-in field isn't a good field
This is a trap I've seen a lot. An incident has a CI, an assignment group and a service, all populated, so it looks fine. Then you dig in. The CI owner changed six months ago, the team no longer supports the application, and the linked article describes an older version.
A person who knows the environment spots that quickly. AI treats it as input and can still write a very confident answer, which is what makes it dangerous. A wrong answer that sounds sure of itself is harder to catch than one that looks wrong.
Dashboards won't save you here either. You can have, say, 98% of CIs with an owner and still have bad data for the one service your pilot depends on. I care a lot more about the quality of the records behind the use case than the overall CMDB score.
What I'd do before switching anything on
Pick a use case small enough to test. "Deploy Now Assist" is too broad. "Summarise incidents for one support team" is something you can measure. I like to think in the order Read, Summarise, Recommend, Act. Plenty of organisations are ready for the first three well before they're ready for the last.
Trace real records. Take 20 or 30 real incidents or questions from your pilot scope and follow every CI, group, relationship and article involved. For each one, ask whether you'd be comfortable letting an AI answer or action depend on it. That's a much tougher test than "is it populated?"
Look at knowledge honestly. Having thousands of articles doesn't mean having a good knowledge base. Who owns them? Do review dates mean anything? Are duplicates and conflicting articles sorted out? Once AI is answering from that content, an outdated article becomes an outdated answer delivered with confidence.
Trust a slice of the CMDB, not all of it. I wouldn't wait for a perfect CMDB, because in a large enterprise that day never comes. If the first use case covers a set of business applications, validate the CIs, owners, support groups and relationships for those, with the teams who run them. Know what you can trust, and keep the AI inside that boundary.
Check access early. ACLs and roles matter more with AI, because it may be retrieving information on a user's behalf. The question I always ask is what the user can see, and what the AI can see on their behalf. If the answer isn't the same, that's a design problem to fix before go-live.
Be strict with agents. An agent that answers a question is very different from one that updates records. Before it touches a production workflow, I want to know its objective, what information it can use, which tools it can call, what it's allowed to change, where a person has to approve, and what it does when something doesn't make sense. A good prompt is only one piece of that.
Sometimes the best answer is "not yet"
I'd be comfortable saying it if nobody owns the data, if knowledge contradicts itself, if the CMDB relationships the use case needs can't be trusted, if access rules aren't understood, if there's no way to test against realistic scenarios, or if an agent is being given more authority than anyone can control.
Looking back at that project, "not yet" at the start would have saved us a lot of time. It's hard to say when the demo looks impressive and everyone is excited. I'd still rather delay a rollout than put something live that people trust only because it sounds sure of itself.
Where I'd start
ServiceNow's AI will keep getting better at context, search, evaluation and action. But it can't know that an old CI owner is wrong, or which of two articles reflects the current process, unless the organisation gives it that context.
So before you switch on the next AI feature, spend a week with the data it will read. Not the demo data and not the tidy example. The real records.
If this response helped, please mark it as correct and close the thread ✅ — it helps future readers find the solution faster.
Thanks
