- Post History
- Subscribe to RSS Feed
- Mark as New
- Mark as Read
- Bookmark
- Subscribe
- Printer Friendly Page
- Report Inappropriate Content
2 hours ago
Read the instance
before you design the agent.
Harvest, filter, score, visualise — a method for turning an instance into a decision, in that order and no other.
← Part 1: One Flow, Three Layers
In Part 1 we described a pattern: three layers, three decisions, every request passing through all of them. We also admitted its limitation. A pattern tells you what to look for. It does not tell you what to fund.
This part is about closing that gap, and it is the part we get asked about most often — usually phrased as “fine, but how did you get those numbers?”
The short answer: a method, not a tool. There are tools that accelerate it, and we use them where a customer has them. But every step below can be reproduced with reports and queries that already exist on your own instance, by your own team, and we have deliberately written it that way. If you finish this article and cannot run it yourself on Monday, we have failed.
Why an overview has to come first
Skip this step and one of two things happens. We have watched both, more than once.
The organisation builds the idea with the loudest sponsor. It works, technically. It touches four hundred records a year. Nobody can explain in the next steering committee why that was the right first move, so the second wave never gets funded.
Or it builds the technically easiest idea. That one ships fast and then sits in a review where somebody asks what it saved, and the honest answer is “we do not have a baseline.”
Most AI initiatives we see do not fail on capability. They fail because nobody could state, in one page, what the organisation’s work actually looks like before the first agent was built. Without that page, every result is unfalsifiable — and unfalsifiable results do not survive a budget cycle.
An overview is cheap. It is a few days of querying and thinking. It costs no build effort, no licence, no change window — and it is the only artefact that makes the layers in Part 1 arguable instead of assertable.
The method: four steps, strictly in order
The order matters more than the individual steps. Every time we have been tempted to score before filtering, we have produced a ranking that had to be thrown away.
Harvest — breadth before depth
Pull wide, shallow slices across several angles rather than one deep slice on the topic you already suspect. We use six: incident volume and channel distribution, request and catalog usage, knowledge base and search behaviour, routing behaviour, resolution behaviour, and intent clustering over short descriptions.
Two practical rules. Use a window long enough to survive seasonality — three months for incidents, six for requests, 180 days for catalog usage. And record the window next to every number, because you will compare figures across pulls later and the windows will not match.
Filter — where most of the judgement sits
This is the step everyone underestimates and the step that changes conclusions. The raw volume in a mature instance is not addressable work.
A substantial share of what a mature instance produces is machine-generated — monitoring integrations, scheduled jobs, technical accounts. Those records are already structured at source. They are not a deflection target at all, because there is nobody to deflect; they belong to a separate automation track. Leave them in the denominator and you silently deflate every deflection figure you go on to calculate, including the ones you will quote in a steering committee.
So before scoring anything: separate machine-generated from human-originated, note the split, and carry it forward as context rather than deleting it. The same applies to duplicate catalog variants and to records created by bulk migration. What that split looked like on the instances we examined — and what follows from it — is Part 3.
Score — two axes, four verdicts
Once you have clusters of comparable work, score each on two independent axes: value (how much is it worth if this stops being manual) and feasibility (how likely is it that an agent can actually do it). Crossing them gives four verdicts, and the vocabulary alone changes how the conversation goes.
Home Run
High value, high feasibility. Build these first. There will be fewer than you hope.
Enablement
High feasibility, lower value. Cheap wins that build credibility and a delivery muscle.
Moon Shot
High value, low feasibility. Worth a roadmap slot, not a first wave.
Danger Zone
Low on both. The ideas that sound impressive in a workshop and quietly consume a quarter.
What goes into each axis is where you should argue with us. Ours: value takes volume, elapsed resolution time, routing churn and business impact; feasibility takes channel structure, how focused the topic is, how consistent the wording of the requests is, and whether any automation history exists. The weights are debatable — the discipline of scoring on two independent axes is not.
High wording variety in short descriptions lowers feasibility, but it does not lower it as much as people assume. In several clusters we scored, descriptions were 100% unique — and the underlying process was still completely uniform. The variety was in how people described the problem, not in the problem. That is an intent-mapping job, not a process-redesign job, and it is much cheaper than it looks.
Visualise — one page per layer
Do not present a ranked list of clusters. Present the three layers from Part 1, and under each one put what your instance says about that stage. That single formatting choice is the difference between “here are twenty-four ideas” and “here is what your organisation looks like, and here is where the work is.”
The layer view also makes gaps visible in a way a ranked list hides. If Layer 1 has no measurement at all, that fact belongs on the page as prominently as any number — because it means every deflection figure you might quote later is currently unknowable.
The two questions that produce different answers
Now the part we would most like you to take away, because it caught us out.
There are two entirely different questions you can ask of a cluster of work, and they are routinely conflated:
| Question A | Question B |
|---|---|
| “Which existing agent fits this?” A coverage question. Matches your work against what already ships. |
“Could an agent be built for this?” A feasibility question. Ignores what exists and asks whether the steps are expressible at all. |
We have run both against the same dataset. Question A returned almost no usable coverage — of several hundred evaluations, only a small fraction found an existing agent that could actually cover the case. Question B, on the same data, returned a buildable path for essentially every cluster, with candidate steps generated for each.
Both answers are correct. They are answers to different questions. Read only A and you conclude the platform cannot help. Read only B and you conclude everything is easy and skip the build effort in your plan. The honest conclusion is the one you can only reach by running both: little of this comes for free, and almost all of it is possible.
The granularity you cluster at determines what you can see. In one analysis, a coarse pass produced two dozen clusters and a fine pass produced ninety-three — and an entire domain worth roughly 15% of the sample appeared in the fine pass and in none of the coarse clusters. Not because it was small, but because it was fragmented across access, locks, single sign-on, master data and interface errors, each too small to surface on its own. If you only ever cluster coarsely, your biggest blind spot is invisible by construction.
Measurement or sample — say which
This is a discipline, not a technique, and it is the single habit we would most like to see spread.
Some numbers are measured across the full population: channel distribution across all incidents in the window, reassignments per incident, the share of records closing without notes. Others come from a sample — intent clustering typically runs on a bounded number of records — and turning those into annual figures means extrapolating from a sample that is almost certainly not stratified.
Extrapolate anyway. It is the only way to size a business case. But label it. Write “directional, extrapolated from an n-record sample” under the number, every time. Two reasons: an executive who later discovers the caveat you did not mention discounts everything else you said, and a cluster that is thin in the sample will be wildly wrong when multiplied.
Never mix a measurement and an extrapolation in the same sentence without saying which is which. It costs one clause. It buys you the right to be believed on everything else.
Run it yourself: the starter set
Six things to pull. None of them requires a specialist tool, and together they populate the three-layer page.
| What to measure | What it tells you | Layer |
|---|---|---|
| Volume by contact channel, plus the top record-creating accounts | Where work actually enters — and how much of it is machine-generated and therefore not deflectable | 1 |
| Catalog item usage over 180 days: created versus active versus actually used | How much of your largest deflection surface is live, and how much is carrying no volume at all | 1 |
| Search queries: total, and the share that are a single word versus a full question | Whether your search surface is being asked questions or keywords — this decides what deflection can even look like | 1 |
| Reassignments and field modifications per record before close | What a missing qualification layer costs today, in churn rather than in opinion | 2 |
| Share of records closing with no close notes | Whether resolutions become reusable knowledge or evaporate — nothing downstream can reuse what was never written down | 2 |
| Assignment groups touched per topic cluster | Routing spread. A topic scattered across many groups has no owner, and that is a taxonomy problem before it is an AI problem | 3 |
Two of these — the churn figures and the close-notes share — are the ones we would run first if we only had an hour. They are single queries, they need no interpretation, and they are almost always high enough to start a conversation on their own.
The platform side of this exists as a product: AI Agent Advisor performs the clustering, matching and scoring described above against your instance data. Use it if you have it — it is faster and more consistent than hand-rolled queries. But the method above is what it is doing, and understanding the method is what lets you defend the output when someone challenges a number. Worth being explicit about the distinction, since this series contains both: that is a supported product, whereas anything we hand you in these articles is a demo-grade MVP. Product documentation.
What comes out the other end
One page, three columns, one per layer. Under each: the volume that reaches it, three or four measured figures about how it behaves today, and one honest verdict sentence.
What made the biggest difference for us was allowing that verdict to be negative — and allowing it, in one case, not to be a number at all. A layer whose honest verdict is “we cannot measure this yet” looks like a weak finding next to three columns of percentages. It is usually the most valuable line on the page, because it is the one thing that has to happen before any claim about that layer means anything, and it costs no build effort to fix.
Resist the urge to fill that column with something more impressive. A page with one honest gap on it is defensible. A page where every column looks equally solved is the one that falls apart in the second review.
This is the artefact both people from Part 1 were missing. It is what Miriam takes into the board meeting instead of forty-one ideas, and it is what tells Jonas which one to open on Monday. Neither of them needed a better answer — they needed one page that made the question decidable.
Figure 1. The shape of the page, not our numbers. Each column gets filled from your own instance.
From that page, sequencing follows almost mechanically. Layers 1 and 2 are configuration and taxonomy work, and the build lands in weeks — what paces them is agreeing the category set, not writing anything. Layer 3 rolls out in waves, starting with whatever your Home Run quadrant actually contains — not with whatever was most impressive in the idea list.
Part 3 takes Layer 1 apart: why deflection has to be designed per channel, why catalog cleanup is not housekeeping, and what happens when you move deflection into the record producer itself. That part includes a working build and an update set you can import.
The series
| Part | Topic | Status |
|---|---|---|
| 01 | One flow, three layers — the pattern | Read now |
| 02 | Read the instance before you design the agent — harvest, filter, score, visualise | You are here |
| 03 | Layer 1 — deflect where the request is born, including a working build to import | Read now |
| 04 | Layer 2 — the gatekeeper you can build today | Read now |
| 05 | Layer 3 — specialists scoped to a category, not a team | Next week |
| 06 | Six things that bite you when you build this | Next week |
Tell us where we are wrong
If you weight the two scoring axes differently, if your filter step throws away something ours keeps, or if you have found a seventh measurement that earns its place in the starter set — the comments are the right place. We answer them and we update the article.
Views are our own and do not represent our team, employer, partners, or customers. Anything we build and share in this series is a demo-grade MVP — not a ServiceNow product, not part of any roadmap, and not supported.