Interested in a ServiceNow event built for developers? Registration for now[dev]26 is officially open!

Luis Estéfano
ServiceNow Employee
Now Assist · AI Voice Agents · Assistant Designer · AI Agent Studio

AI Voice Agents in practice: configuring the Voice Assistant and the AI Agent for the voice channel

A voice agent that works well in a demo and a voice agent that works well in a store at 8:55 in the morning, with a queue at the till and a printer that refuses to print, are two different things. The difference rarely comes from the model. It comes from dozens of small configuration and design decisions: which channel callers use, how the assistant greets them, what happens when the agent gets stuck, and how the instructions and knowledge are written for the ear instead of the eye.

There is already a great overview article on AI Voice Agents covering what they are, the architecture, and the out-of-the-box use cases. This article picks up where that one stops. It walks through the configuration screen by screen and shares the lessons learned while taking a voice agent live for a European retailer's store operations.

 

Author
Luis Estéfano
Applies to
Now Assist · AI Voice Agents · Mobile · Telephony
Release
Zurich · Australia
 
01 — Before you start

One plugin suite, two roles, a few prerequisites

AI Voice Agents ship as part of the Now Assist suite. Everything described in this article runs on the application below.

Now Assist suite · Plugin

ServiceNow Otto for Voice Agents

App ID: sn_voice_aia

Latest: Now Assist Suite version 29.5.20260918 (App version 5.2.4)

Configuration surfaces

Where you will work

  • Assistant Designer — the Voice Assistant: language, voice, channels, caller verification, safeguards.
  • AI Agent Studio — the AI voice agent: instructions, tools, data access.
  • Mobile App Builder — the voice launcher, if you use the mobile channel.
Task Role required Where
Create and configure a voice assistant virtual_agent_admin or admin Assistant Designer
Create and configure AI voice agents sn_voice_aia.admin AI Agent Studio

A few things need to be in place before the first test call:

  • Now Assist entitlement for the workflows you plan to cover. The out-of-the-box voice agents for ITSM, HR, and CSM come as separate Store applications — install them only if you want to start from those templates.
  • Authentication factors configured at platform level (Soft PIN, SMS OTP, TOTP, Okta Verify, email OTP, or security questions) if you plan to use the telephony channel. The Voice Assistant can only select factors that already exist.
  • Clean user phone numbers. Caller identification by phone number matches the incoming caller ID against the user record. Store numbers in international E.164 format (for example +34600000000) or matching will silently fail.
  • Voice-ready knowledge. If the agent answers from the knowledge base, the articles need to work when read aloud. Section 07 covers what that means.
Heads-up — Assist consumption

Voice consumption is tiered by the number of tools executed during a call, and caller identification or authentication counts as a tool. Lean agents with a focused tool list are cheaper as well as faster. Confirm the details that apply to your entitlement with your account team.

02 — The mental model

Two building blocks, two very different jobs

Every configuration conversation gets easier once this split is clear. The Voice Assistant and the AI voice agent are configured in different places and answer different questions. Most early confusion — "where do I make sure callers only get their own data?" — comes from looking for a setting in the wrong one.

Assistant Designer

Voice Assistant — the front door

How it sounds, who it hears, where it lives, and what happens when things go wrong.

  • Language, voice persona, greetings, pronunciation
  • Communication channels (telephony, mobile, web)
  • Caller identification and authentication
  • Fallback, call duration, inactivity timeout
  • Noise cancellation
AI Agent Studio

AI voice agent — the brain and the hands

What it knows, what it can do, and what data it is allowed to touch.

  • Role, description, and step-by-step instructions
  • Tools: knowledge search, flows, record operations, scripts, MCP
  • Who can use it (ACLs) and whose identity it runs as
  • Voice channel activation and testing

One assistant can host several voice agents, and the orchestrator picks the right one based on the caller's intent. Keep one topic per agent, rather than one agent that tries to do everything. Shorter, focused instructions and tool lists are followed more reliably, and they keep latency down.

LuisEstfano_1-1791296662826.png
Screenshot — Assistant Designer — voice assistant Overview tab showing language, channel, welcome message, and the linked AI voice agents.
03 — Voice Assistant configuration

Eight screens, in order

Go to All > Conversational Interfaces > Assistant Designer > Assistants, select Create assistant, and choose Voice-only. The wizard walks you through the screens below. You can revisit any of them later from the assistant's Settings tab.

STEP 1

Basic details and assistant instructions

Give the assistant a name tied to the business outcome (Store Support — DE, HR Service Desk) and a meaningful description so it does not overlap with other assistants. The Assistant instructions apply to every voice agent behind the assistant, so this is the place for rules that hold across topics: the assistant's role, pacing, and what to do in specific situations.

// Assistant instructions — example for a store support line
You are the store support voice assistant for [Company]. Your job is to help
store employees with hardware, point-of-sale, and store system issues.

Speak in short, clear sentences. Ask one question at a time and wait for the
answer. Never read out long lists; offer the most relevant option first.

If a caller asks about payroll or other personal HR topics, explain that this
line handles store support only and offer to create a ticket instead.
LuisEstfano_2-1791296697431.png
Screenshot — Basic details step with the Name and Assistant instructions fields.
STEP 2

Add the AI voice agents

Select Add from library to attach existing voice agents, or Create to build a new one. This step is optional during creation, but the assistant can't be activated without at least one agent. Activating the assistant also activates any inactive agents linked to it.

STEP 3

Language and voice

This screen shapes how the assistant sounds and how well it hears. It deserves more time than it usually gets.

Setting What it does Recommendation
Opening message Always plays first, before any language selection. Single language: end it with an intent question (“How can I help you today?”). Several languages: keep it general, for example a recording notice.
Welcome message Plays after the caller selects a language. One per language. Only used when secondary languages exist. End each one with an intent question.
Language input How callers pick a language: DTMF keypad or by voice. DTMF is the recommended option. Only one can be enabled.
Voice persona The voice shared by every agent in the assistant. Preview several. A warmer voice suits HR; a crisp, faster one suits store operations.
Pronunciation dictionary Controls how text-to-speech says a word, per language. Add brand names, product names, and acronyms the voice gets wrong.
Key term dictionary Helps speech-to-text recognize words it might mishear. Up to 30 terms per language. Add internal system names and product codes that sound like everyday words.
Tip — hearing versus speaking

The two dictionaries solve opposite problems. The key term dictionary fixes what the agent hears (speech-to-text). The pronunciation dictionary fixes what the agent says (text-to-speech). A store system name that is misheard as a common word needs a key term; a product name the voice mispronounces needs a phoneme. Build both lists from real test transcripts, not from guesses.

LuisEstfano_3-1791296763569.png
Screenshot — Language and voice step — primary and secondary languages, opening message, and the Voice persona tab with audio preview.
LuisEstfano_4-1791296815659.pngLuisEstfano_5-1791296842908.png
Screenshot — Pronunciation dictionary and Key term dictionary tabs with sample entries.
STEP 4

Communication channels

Select the Provider application first; it is required for every channel type. Then configure at least one of the two channel families:

  • Telephony provider — callers dial a phone number. Pick a channel type (SIP, PSTN, or WebSocket, depending on the provider) and a CCaaS provider such as Genesys Cloud, Twilio, Amazon Connect, Five9, or 3CLogic. Check the documentation for the providers supported on your release. The assistant generates the connection details you paste into your CCaaS configuration.
  • Web Real-Time Communication (WebRTC) — voice inside an application. Mobile applications covers the ServiceNow mobile apps (voice launcher functions, chat launcher, prominent action button). Web applications covers the voice call widget on a Service Portal or Engagement Messenger.
Field note — start with mobile

In most of the deployments I've worked on, the mobile channel is the better first step. There's no CCaaS integration to set up or pay for, and the user experience is almost identical: one tap on a launcher instead of dialing a number. Employees already signed in to the app are already known to the platform. Telephony is the right choice when callers don't have the app, or when the voice agent has to sit behind an existing contact center number.

LuisEstfano_6-1791296934683.png
Screenshot — Communication channels step — WebRTC tab with the mobile voice launcher function selected.
STEP 5

Caller verification (telephony only)

Authentication settings apply only to the telephony channel. If you only use the mobile channel, skip this step: the user is already signed in to the app. For telephony, ServiceNow separates two ideas that are easy to mix up. Identification works out who the caller claims to be, for example by matching the incoming number to a user record. Authentication proves it.

Factor What the caller does Input
Knowledge-based (security questions) Answers questions configured at platform level Voice or DTMF
Okta Verify push Approves a push notification on their device Approve on device
SMS verification code Reads back a one-time code sent by SMS Voice or DTMF
Authenticator app (TOTP) Reads the time-based code from their authenticator app Voice or DTMF
Soft PIN Provides a numeric PIN enrolled through ServiceNow Voice or DTMF
Email one-time password Reads back a code sent to their email Voice or DTMF
  • Configure a fallback identification method for callers who can't be uniquely identified by the first one, for example two users sharing a store landline.
  • Multi-factor authentication is the default. To allow single factor, set the glide.voice.authenticate.mfa_mandatory system property to false.
  • Retry sets the number of attempts before the caller is routed to a live agent. The default is 3.
  • Authenticate at the start of the call prompts every caller before any request is handled. Leave it off if some agents answer non-sensitive questions that don't need authentication.
  • Changed a factor at platform level? Select Refresh configuration on this page; changes outside Assistant Designer are not picked up automatically.
LuisEstfano_7-1791296974270.png
Screenshot — Caller verification step — identification method cards, first and second factor, and advanced options.
STEP 6

Safeguards — design the fallback on purpose

The fallback kicks in when the agent can't finish the job: a tool fails, the agent has no information to move forward, the caller insists on escalating, or the call reaches its maximum duration. Treat it as part of the experience, not as an error path. The available options depend on the channel:

Channel Connect to live agent Generate a ticket with record producer
Telephony Available. Transfer must also be set up in your telephony provider. Optionally capture details first so the call reaches the right team. Available. One of the two options is required.
Mobile / WebRTC Not available Required

When the record producer is used, the agent fills it conversationally: it asks the caller for each field and submits the ticket. That makes the record producer design critical. Keep it to a short description, a description, and optionally a callback number. Every extra mandatory field is another question the caller has to answer after the agent has already failed to help. The ticket can be an incident, a case, or a simple log record for a support team to pick up later. Enable Require authentication if only authenticated employees may create tickets.

Call constraint Limit Recommendation
Max call duration Up to 10 minutes Triggers the fallback when reached. Size it to your longest successful test call plus a buffer, not to the maximum.
Inactivity timeout Up to 300 seconds (default 30) The caller is reprompted, then disconnected. Hands-on troubleshooting needs longer than a status check — a store employee opening a cabinet with a key needs time.
LuisEstfano_8-1791297006632.png
Screenshot — Safeguards step — fallback behavior with record producer selected, and call constraints.
STEP 7

Advanced settings

Noise cancellation reduces background noise from the caller's side. Low picks up even quiet sounds and suits quiet rooms. Medium (the default) suits public places and normal chatter. High only reacts to loud noise and suits shop floors, warehouses, or construction sites. Test it from the real environment — a store near the tills at opening time is not an office. If you use secondary languages, the Language and voice tab here also sets the retry and goodbye messages for language selection.

LuisEstfano_9-1791297039636.png
Screenshot — Advanced settings — noise cancellation levels.
STEP 8

Review and activate

Select Save and activate. Activation only succeeds when all of the following are true:

  • At least one communication channel is configured (telephony provider or mobile channel).
  • At least one AI voice agent is linked to the assistant.
  • If telephony is enabled, authentication is set up and the fallback is either live agent or record producer.
  • If the mobile channel is enabled, the fallback is a record producer.
04 — Channel design

The mobile voice launcher and multilingual patterns

On the mobile channel, a voice launcher function in Mobile App Builder is what puts the voice agent one tap away. The function is linked to the assistant on the Communication channels step, and its visibility conditions decide who sees it. That makes the launcher a useful routing tool: show it only to users with a given role, in a given location, or with a given language.

The same experience can be embedded outside the app: a voice call widget on the employee portal or Engagement Messenger, opened from a laptop browser or a phone. Which device the user is on doesn't matter; where they already work does.

Two ways to go multilingual

Voice assistants now support a primary language plus secondary languages, with an in-call language selection step. That opens two valid patterns:

Pattern How it works Best when
A — One assistant, many languages Primary and secondary languages on one assistant. Callers pick their language by DTMF or voice at the start of the call, then hear that language's welcome message and voice persona. A shared phone number serves callers from several countries, and you don't know the caller's language in advance.
B — One assistant per language A dedicated assistant per language, each with its own persona and greeting, all linked to the same voice agents. A launcher per language, with conditions on the user's location or language, opens the right one. The channel already knows who the user is (mobile, portal). The user skips the language question and gets straight to the point.
Tip — keep the agent shared

In both patterns, the voice agents behind the assistant stay the same. Write their instructions so they answer in the caller's language. For the knowledge the Agent has access to, maintain the logic once, and only the front door changes per country.

05 — AI voice agent configuration

Same studio, a few voice-specific decisions

In AI Agent Studio > Create and manage > AI agents, open the Add drop-down list and select Voice. You can also start from the Asset Library in Assistant Designer (Create asset > AI Agent). The flow is the same as for chat agents, with a handful of voice-specific choices.

STEP 1

Define the specialty

Name, description, role, and the list of steps work exactly as for chat agents, and Generate details can draft them from a short description. The difference is in how they are written: they will end up spoken, not read. Section 06 covers that in detail.

STEP 2

Add tools and information

At least one tool is required. Recommend Tools proposes tools based on the description and instructions. For voice, the documentation calls out that tool inputs and outputs should be strings for the best experience.

Tool Voice-specific advice
Search retrieval Create a dedicated search profile containing only the knowledge articles meant for voice. A smaller search scope means lower latency and fewer off-topic answers.
Record operation, flow action, subflow Return only the fields the agent needs to say out loud. A short, speakable result beats a full record the model has to filter.
Script Useful for calls to external systems, such as resetting a device through its API. Keep them fast, and have the agent tell the caller when a call might take a few seconds.
MCP server tool Reaches data and agents outside ServiceNow. Apply the same rules: fast, focused, least privilege.
File upload Static reference material, such as the technical manual nobody reads end to end.
STEP 3

Define security controls

This is where data protection actually lives — not in the prompt, and not in the Voice Assistant.

  • Who can access the agent: any authenticated user, users with specific roles, or public.
  • Public with Require caller identification: the agent identifies the caller using the assistant's identification methods before the conversation starts. Identification doesn't authenticate the caller or grant more access — the session keeps public permissions. Don't use this to protect personal data.
  • Dynamic user (the default): the agent runs as the user who invokes it, so that user's ACLs decide what data it can reach. This is what makes "only my data" work.
Field note — scope data in the tools, not the prompt

The most common question in HR voice discussions is "how do we make sure an employee only hears their own data?" The answer is always the same: run as the dynamic user, respect ACLs, and scope each tool. A get_user_benefits tool should return the benefits for the caller's own country and contract, not all of them with an instruction to filter. Rules in a prompt are guidance; ACLs and tool filters are guarantees.

STEP 4

Channels, activation, and first test

Select Allow to enable phone calls, pick the voice assistants that host the agent, and toggle the status to active. Test in assistant opens the voice testing panel directly when one assistant is linked. It is disabled when no assistant is linked yet.

06 — Instructions for voice

Write for the ear, not the eye

You can reuse the same agent logic you already run in chat. What changes is the delivery. A chat user skims a bulleted list in two seconds; a caller has to hold it in their head. Research from teams building voice agents points in the same direction: people are comfortable with responses that start within about 1.5 seconds, and every tool call adds an extra model round trip before the agent can speak (Voice AI and Voice Agents: An Illustrated Primer). Silence on a phone line feels like a dropped call.

A short block of voice rules at the top of the agent's steps makes a visible difference. This is the pattern I use as a starting point:

// Voice rules — add at the top of the AI agent's list of steps
Your answers are spoken aloud to the caller.
- Use plain spoken sentences. No lists, bullet points, markdown, links, or emojis.
- Keep each turn to one to three short sentences. Ask one question at a time.
- Give one instruction per turn, then wait until the caller confirms it is done.
- Before using a tool that may take a few seconds, say so first:
  "Give me a moment while I check that."
- Never go silent. If you need information to continue, ask for it.
- Read numbers and codes digit by digit and confirm them back to the caller.
- If an answer sounds unexpected, it may be a transcription error. Confirm it:
  "Did you say self-checkout?"
- If the caller asks you to repeat, repeat the last step in simpler words.
- Ask for confirmation before changing or closing any record.
- If the caller asks for a person, or the issue is not solved after the last
  step, explain what happens next and hand over.
- End by summarizing what was done and the outcome of the ticket.

And the steps themselves, for a store hardware troubleshooting agent:

// List of steps — store hardware troubleshooting (example)
1. Ask the caller to describe the issue in their own words.
2. Identify the affected device and its identifier, for example the register
   number. Ask which type it is (conveyor belt or self-checkout).
3. Search the knowledge base for the troubleshooting article that matches the
   device and type. Use only that article to guide the caller.
4. Guide the caller through the steps one at a time. After each step, ask what
   they see, for example which indicator light is on.
5. Answer side questions from the FAQ section of the same article.
6. If the issue is resolved, write a short summary of the conversation on the
   ticket and close it as resolved.
7. If it is not resolved after the last step, document the steps tried and
   hand over using the fallback.
Tip — test the prompt by listening to it

Read your instructions and a sample answer aloud. If you run out of breath, the answer is too long. If you had to look at the screen to remember the second option, the caller won't remember it either.

07 — Knowledge for voice

Knowledge articles that work when read aloud

In the deployment described below, the agent's intelligence was mostly in the knowledge base. The troubleshooting articles decided how good the conversation felt. Articles written for technicians with screenshots rarely survive the trip to voice unchanged. What worked:

  • Start with identification. Put the questions that tell variants apart at the top (“Is it a conveyor belt or a self-checkout register?”), so the agent asks them before giving any instruction.
  • One physical action per step, with the expected result. “Switch the printer off with the toggle on the left, wait 10 seconds, and switch it back on. The green light should turn on.”
  • Describe locations in words. The caller can't see an image. “The lock cylinder is on the outside frame of the door” works; “see picture 3” doesn't. Instructions that only live in images or wide tables don't translate to speech.
  • Add a short FAQ block at the end for side questions (“which button resets it?”), so the agent can answer them without losing its place in the procedure.
  • Define the exit. State when to stop troubleshooting and escalate, so the agent doesn't improvise a step that isn't in the article.
  • Keep voice articles together. A dedicated knowledge base or category feeds the dedicated search profile from section 05.
  • Maintain one version per language you serve, following the multilingual pattern you picked in section 04.
08 — Field notes

What a live retail deployment taught us

The first use case we took live for a European retailer was deliberately narrow: store employees solving hardware issues with a voice agent in the retail mobile app, grounded in troubleshooting knowledge articles. Here is how a real conversation about a receipt printer flows:

01
Tap to talk
The store employee taps the voice launcher in the mobile app and describes the issue: the receipt printer isn't printing.
02
Identify
The agent asks for the register number and whether it's a conveyor belt or a self-checkout register, and picks the matching article.
03
Guide
It explains how to open the register door with the special key. When the employee asks it to repeat, it rephrases the step more simply.
04
Restart and check
Switch the printer off, wait 10 seconds, switch it on. The agent asks which light is on and confirms the printer works.
05
Document and close
The agent writes the conversation summary on the ticket and closes it as resolved. No human touched it.


Lessons we took forward

  • Start where trust is cheap. Pick a frequent, non-sensitive, knowledge-grounded topic. Store hardware issues were ideal. Sensitive topics such as payslips come later, once identification and data scoping are proven. Region-specific benefits questions answered from a knowledge article per country are a good bridge.
  • Earn the right to expand. Once the first use case was live and the feedback was good, the same pattern was onboarded for other streams (IT and HR). Trust grows one working use case at a time.
  • Voice is a delivery layer on top of the same agent. The logic was the same as in other channels. What made it work on the phone was the voice rules, the knowledge structure, and a conversational fallback.
  • The agent reads what nobody reads. Technical manuals are impossible for people to memorize. Give them to the agent and it can answer the obscure question in the middle of a busy shift.


The road from guided to proactive

01
Guided (live)
The agent reads the knowledge article and the employee performs each action. Fast to build, immediate value.
02
Remediated (tested)
A script tool calls the device through its API and resets it directly. The employee just confirms it works.
03
Proactive (the goal)
An event flags the failing device, CMDB and CSDM relationships show which store and service it affects, and it's restarted before anyone notices. This depends on CMDB maturity.
09 — Testing and going live

Test the whole stack, in the real environment

Voice agents are non-deterministic, and their inputs are multi-turn conversations spoken with background noise, accents, and interruptions. Test them like that: many runs of each scenario, not one happy-path call.

  • Test in the platform. The testing experience in Assistant Designer runs the full stack — assistant, agents, and tools — without an external call, in voice mode or in a quicker chat mode. It needs Zurich Patch 7 or later, Conversational Studio v7, and microphone access in the browser.
  • Test from the real device and place. Once it works in the browser, call from the store device on the shop floor. This is where noise cancellation, inactivity timeouts, and dictionaries get tuned.
  • Read the transcripts. Transcripts and logs are stored in dedicated tables and are the fastest way to see where the agent misheard, rambled, or picked the wrong tool.
  • Watch the analytics. The Analytics tab in Assistant Designer shows conversation volumes, outcomes, and agent performance once real users arrive.
  • Evaluate at scale. Use voice agent evaluations to score performance and outcomes across many conversations before and after every change to instructions or tools.

Scenarios worth adding to every test plan:

Scenario What to check
Numbers and codes Register numbers, ticket numbers, and PINs are understood, read back digit by digit, and confirmed.
Names and email addresses Spelling-heavy input is confirmed rather than assumed.
Interruptions The caller cuts in mid-sentence; the agent stops and answers the new question.
Silence The caller goes quiet during a hands-on step; the reprompt timing feels right.
Tool failure A tool errors or times out; the agent says so and the fallback fires cleanly.
Out-of-scope request The caller asks about something the assistant doesn't cover; the agent says so politely.
Noisy environment Background talk or machines don't trigger false interruptions or transcription errors.
Wrong language A caller speaking another language gets a sensible experience.
10 — Checklist

Voice-specific do's and don'ts

The overview article has a great general list for AI projects. These are the ones specific to the voice channel:

Do
  • Start with a frequent, non-sensitive, knowledge-grounded topic.
  • Consider the mobile channel first when users already have the app.
  • Keep one topic per voice agent and a focused tool list.
  • Add a voice rules block to every agent's instructions.
  • Build pronunciation and key term entries from real transcripts.
  • Use a dedicated search profile for voice knowledge.
  • Design a short, conversational record producer for the fallback.
  • Scope data with dynamic user, ACLs, and tool filters.
  • Test from the real device, in the real noise.
Don't
  • Paste chat-optimized instructions unchanged.
  • Let the agent read lists, links, or long identifiers.
  • Leave the caller in silence while a tool runs.
  • Start with payslips or other sensitive data.
  • Rely on the prompt to protect personal data.
  • Treat caller identification as authentication.
  • Set the maximum call duration to 10 minutes by default.
  • Forget to store phone numbers in E.164 format.
  • Serve a language without knowledge in that language.

That's the full picture: a Voice Assistant that decides how the conversation sounds and where it lives, an AI voice agent that decides what it knows and what it can do, and a handful of voice-specific habits that turn a good demo into something store employees actually reach for. Start small, keep it spoken, design the fallback on purpose, and let the first working use case earn the next one.

 

#servicenow #nowassist #aivoiceagents #voiceai #agenticai #assistantdesigner #aiagentstudio #mobile #retail

 

If this article was useful, please consider marking it as helpful. Feedback is always welcome.

Kind regards — Luis Estéfano
Version history
Last update:
13m ago
Updated by:
Contributors