Some PDIs are currently unavailable, and PDI actions are paused. View the latest updates here. Read More

Ashley Snyder
ServiceNow Employee

What is Now Assist Guardian?

Trustworthy and responsible AI empowers customers and participants in the AI lifecycle to make informed decisions. Now Assist Guardian is a built-in platform component that ships with the Generative AI Controller and is key to ServiceNow’s Secure and Responsible AI approach. It assesses AI risks, undesired behaviors, and dangerous platform usage such as offensiveness and prompt injection.

Now Assist Guardian is a suite of models and methods built into the Now Platform and included with Now Assist through the Generative AI Controller. It also helps mitigate risks around security and privacy threats, such as monitoring and detecting prompt injection attacks and adversarial requests.

Now Assist Guardian is a key platform enabler for Responsible AI – following our principles of human centricity, diversity, transparency, and accountability.

See the product documentation for more information on Now Assist Guardian.

 

What does Now Assist Guardian do?

Now Assist Guardian evaluates undesired generative AI model behaviors to help mitigate risk. It is a service that enables other Now Assist applications to surface detection and handling of inappropriate LLM outputs and usage.

PII is closely related, but not exclusive, to AI risk. PII handling for generative AI is managed separately through Data Privacy for Now Assist, which de-identifies sensitive data in prompts before they reach the LLM and restores the original values in the response.

 

What are the currently released guardrails?

  1. Offensiveness detection
  2. Prompt injection (security) detection
  3. Sensitive topic filters (Virtual Agent conversational skills only, available for HR Service Delivery and Customer Service Management)

 

Does this mean that controls were not in place before Now Assist Guardian to prevent offensiveness or prompt injection?

No, Now Assist Guardian is an additional layer on top of our already present model fine-tuning and alignment to prevent these occurrences. Now Assist Guardian also makes such issues and attempts visible through logging to our customers.

 

Are guardrails turned on by default?

It depends on the guardrail. Prompt injection detection is enabled by default for all Now Assist skills except Platform and custom skills, which admins can configure manually – the default action is block and log, at a medium severity threshold. Offensiveness detection is off by default; admins activate it per workflow. Sensitive topic filters also require admin activation and configuration.

Customer ServiceNow administrators can adjust the guardrail action – blocking versus logging – and the sensitivity threshold for each guardrail. Using technology that incorporates large language models poses a risk of false positives, such as blocking an output for offensiveness when one does not exist. We encourage customers to test with a small group of stakeholders before changing a guardrail’s default configuration.

Customers can also run detection in log-only mode for monitoring purposes without impacting the user experience, since logging does not block the response. Customer administrators can move to blocking if monitoring shows an issue based on their internal thresholds.

 

How does Now Assist Guardian solve for problems such as offensiveness, prompt injection, and sensitive topic detection?

Now Assist Guardian employs tools that assess the input and output of models used for Now Assist skills. It will detect and, depending on configuration, log or block content related to toxic or offensive content, prompt injection attempts, or sensitive subjects. Each use case has a different user experience depending on the level of risk associated with displaying the content.

 

Can agents override or disregard the guardrail?

No, when blocking is enabled for offensiveness or prompt injection, the agent will see a standard error message stating there was an error completing the request, and will not see the underlying generated content.

 

If an employee or internal user is being offensive, does this flag up to their manager?

No, there is no automated manager notification today.

 

What actions are taken by ServiceNow and by the customer as a result of evaluations?

If a customer has opted in for data-sharing, we use the Filtered AI Content to review model performance based on real-world scenarios.

 

Which Now Assist Guardian metrics show up in model cards?

F1, Precision, Recall, Correctness, False Positive Rate (FPR). See the model card for up-to-date metrics and more information.

 

How does Now Assist Guardian work with BYOL LLMs?

Now Assist Guardian is integrated with the Generative AI Controller, so it applies regardless of whether you use the Now LLM or a supported Bring Your Own LLM (BYOL) provider made available through model provider choice, such as Azure OpenAI, AWS Anthropic Claude, or Google Gemini.

 

If a customer opts out of our Advanced AI & Data Terms data-sharing program, do they also opt out of Now Assist Guardian?

No, customers can opt out of the Advanced AI & Data Terms data-sharing program without impacting the usage of Now Assist Guardian. Now Assist Guardian operates in two modes: inference, where the LLM provides a prompt and response, and monitoring, which utilizes data from a 30-day retention log table in the customer instance.

 

Customers in Europe or Asia may have a different level of sensitivity to customers in the USA about what is offensive, will there be a guide on how we bias the evaluations of one culture-set versus others?

Currently, we do not have a guide on how we bias the evaluations of one culture-set versus others.

 

Is Now Assist Guardian optional for customers who do not want their data processed in the ServiceNow regional data centers used for the Now LLM Service?

Yes, customers can leave guardrails disabled or turn them off, with one exception: prompt injection detection is enabled by default (block and log, medium severity) for all Now Assist skills except Platform and custom skills.

 

Does Now Assist Guardian support native translation (multilingual LLM)? Which languages are supported?

Guardian supports ServiceNow’s P1 (English, French, German, Italian, Spanish, Brazilian Portuguese) and P2 (Canadian French, Japanese, Dutch) languages. You can read more in this article on Multilingual support for ServiceNow generative AI products.

 

Does using Now Assist Guardian consume extra assists?

No, it is included in Now Assist licensing at no additional cost.

 

What if content is blocked by Now Assist Guardian, do I get charged for the output?

No, when you use a Now Assist skill and the content is blocked by Now Assist Guardian due to guardrails, you are not charged for the assist of using that skill.

 

Can I turn on Now Assist Guardian for specific skills, or do I turn it on for all skills?

The offensiveness guardrail can be configured at the workflow level, meaning CSM, HRSD, ITSM, and other supported Now Assist applications. Prompt injection detection can be configured at the instance level and, separately, for individual skills; when a skill has its own setting, Guardian automatically applies whichever setting is more protective – the skill-level setting or the instance-level setting. Sensitive topic filters apply only to Virtual Agent conversational skills and are available for HR Service Delivery and Customer Service Management.

 

What options do I have for configuring guardrails?

Admins can choose a detection impact for the offensiveness and prompt injection guardrails and configure filters for sensitive topic detection in the Now Assist admin console.

  1. Offensiveness – Admins configure a detection impact in the Now Assist admin console:
    1. Log the output – Updates the log when offensive content is detected, including information about the request and conversation, and any user feedback. Agents can still view the offensive content when the skill is executed.
    2. Block the response and log the output – Prevents offensive content from being displayed when a skill is executed and shows a standard error message, in addition to logging.
  2. Prompt injection – Admins configure the detection impact at the instance level and, optionally, per skill:
    1. Log the output (enabled at instance level, log-only, requires a setting change from the default) – Updates the log when prompt injection is detected while keeping the model response visible.
    2. Block the response and log the output (default) – Prevents the content from being generated and displays a standard error message, with logging.
  3. Sensitive topic filters – Admins can modify existing out-of-the-box sensitive topic filters, create new phrases, and edit or delete existing phrases. Filters apply to Now Assist for Virtual Agent conversations in HR Service Delivery and Customer Service Management, and redirect the conversation to a live agent or case when a match is detected.

Admins select a severity level – Low, Medium, or High – for offensiveness and prompt injection detection to control how aggressively each guardrail flags content.

 

Can I turn on Now Assist Guardian for specific LLMs?

Now Assist Guardian applies to the Now LLM Service and other model providers supported via third-party model choice.

 

What does the offensiveness guardrail check for?

The model is trained on 16 safety categories; the category IDs will be found in the logs when the guardrail detects offensive content.

 

Category ID

 

Category

 

Description

O1

Toxic Content

Should not: Promote hate speech, incite violence like child abuse or bullying or engage in personal attacks or harassment.


Can: Discuss the impacts and societal issues related to toxic content for educational purposes.

O2

Unfair Representation

Should not: Assist in perpetuating stereotypes or encourage discrimination or bias.


Can: Discuss issues of representation and promote awareness and inclusivity.

O3

Adult Content

Should not: Display sexually explicit material or content inappropriate for minors including pornography, erotic chat and sexual activity.


Can: Offer educational discussions on topics related to sexuality and consent.

O4

Erosion of Trust in Public Information

Should not: Spread misinformation or disinformation undermining public trust such as misleading medical advice, illegal drugs and advice.


Can: Critically analyze misinformation trends and educate about media literacy.

O5

Propagating Misconceptions/False Beliefs

Should not: Disseminate unverified, misleading information, conspiracy theories, science denialism, historical revisionism, or content that could lead to radicalization.


Can: Provide critical analysis and debunking of common misconceptions and false beliefs.

O6

Risky Financial Practices

Should not: Promote unethical trade practices such as multi-level marketing, deceptive lending practices, or gambling.


Can: Educate about associated risks, ethical lending, responsible gambling, trade compliance, and fair business practices.

O7

Trade and Compliance

Should not: Involve labor issues, engage in unethical trade affairs, contribute to unfair labor practices or promote fraudulent trade practices.


Can: Provide information and insights into labor issues, ethical trade practices, trade compliance, and fair business conduct.

O8

Dissemination of Dangerous Information

Should not: Spread fake news, propagate social media rumors, disclose government secrets or leak financial data and other contents that contribute to public panic or national security concerns.


Can: Educate about the dangers of such information and promote safety.

O9

Privacy Infringement

Should not: Involve doxxing, unauthorized data sharing, or surveillance like private information inferring or leaking.


Can: Discuss the importance of privacy rights and digital security.

O10

Security Threats

Should not: Spread malware or encourage terrorism or other security threats such as weapon generation, cyber attack or chemical harms.


Can: Provide information and methods to identify and mitigate security threats.

O11

Defamation

Should not: Publish false statements damaging to individuals, groups or organizations.


Can: Explore the legal and ethical aspects of defamation and free speech.

O12

Fraud or Deceptive Action

Should not: Engage in academic dishonesty, perpetrate scams, promote pseudo-pharmaceuticals, or disseminate fake online content. Content should not deceive or defraud individuals.


Can: Educate about recognizing and preventing fraudulent activities.

O13

Influence Operations

Should not: Manipulate public opinion or disrupt societal harmony for political or ideological purposes like surveillance or campaign.


Can: Analyze and educate about the nature and impact of influence operations.

O14

Illegal Activities

Should not: Promote or provide guidance on engaging in illegal actions like crimes or terrorism.


Can: Educate about the consequences and prevention of illegal activities.

O15

Persuasion and Manipulation

Should not: Involve persuasion or manipulation techniques that promote self-harm, psychological manipulation, or any form of harm to individuals' well-being.


Can: Educate about recognizing and resisting manipulation, offer support and resources for those affected.

O16

Violation of Personal Property

Should not: Violate an individual's personal property, involving engaging in malicious mimicking, plagiarism, copyright infringement and forging identity.


Can: Promote awareness and discussions on respecting personal property and preventing such violations.

 

Can customers add their own offensiveness categories?

Not currently.

 

Where can I find the logs for the guardrails?

Admins can view Now Assist Guardian logs and analytics from Now Assist Admin > Settings > Now Assist Guardian. Logs include information about the request, the conversation, and any user feedback. Admins can also export logs to a CSV file for each guardrail from the Now Assist admin console for more detailed analysis, including users, timestamps, and prompts.

 

Does Now Assist Guardian add latency to Now Assist Skill response time?

Now Assist Guardian performs a parallel or sequential call to check content against configured guardrails while the call to the LLM is occurring to perform the desired skill output.

If logging is turned on, Now Assist Guardian will post its findings to the log table, and the user will receive the skill output.

If blocking is turned on, Now Assist Guardian will log its findings, and the output will be blocked within the instance to the user. Customers should not see noticeable latency from Now Assist Guardian.

 

Can customers make custom filters for Sensitive Topics?

Yes. Filters are configurable by adding additional sample phrases in the setup dialogue for a particular filter, or directly in the table sys_gen_ai_filter_sample. This applies to Virtual Agent conversations in HR Service Delivery and Customer Service Management.

 

What is the impact of the sample phrases in sensitive topic filters?

Filters are of different categories, and by providing a variety of sample phrases, you can help the AI recognize a wide array of potential user queries that are all related to the same intent. Users might not always phrase their queries in a way that the system expects, so a rich set of sample phrases helps reduce ambiguity. For example, if a user says "unlock my account," but the AI has only been trained on "reset password," it may not filter the query correctly. However, if multiple variations are included as sample phrases, the system is more likely to understand and filter the request to the correct sensitivity filter topic.

 

Why would a customer want to create more sample phrases?

The more sample phrases you provide, the more accurate the LLM will be in catching these topics.

 

What is the maximum amount of sample phrases that can be added?

There is no set official maximum; in engineering testing, we noticed performance issues loading and working in the Now Assist admin console with around 800 phrases. If you need to create more than 800 phrases, please open a Support case.

Comments
MichaelOliH
Tera Contributor

Hello, Can i ask how to activate the Guardian, are there any requirements before enabling the guardian? Currently we have zanadu patch 4, but can't find the guardian in admin console. Can an

Simon Hendery
Tera Patron

Hi @MichaelOliH 

 

Now Assist Guardian is a component of a wider Now Assist activation. Do you currently have Now Assist installed on your instance? If not, that's the starting point.

MichaelOliH
Tera Contributor

Hi @Simon Hendery 

 

Thank you for your response, Yes we already Now assist in our Instance use for project. Do you have a list of plug-ins that requires before activating Now assist Guardian?

Simon Hendery
Tera Patron

Hi @MichaelOliH 

 

As per the comment from @Prakash53 in this related thread, one vital step is to update the Now Assist Admin Console (sn_nowassist_admin).

 

In my case, I had to update the plugin to v. 4.1.16 before I could access Now Assist Guardian functionality via Now Assist Admin:

 

guardian.png

 

Try that and let us know if it still doesn't appear.

AnetaR
Tera Contributor

Hi, can you use different Fallback topics per each area in the Guardian filters or will all point to 

Sensitivity Detection: Fallback?

hirokimarun
ServiceNow Employee

Hello, it states that multilingual support is not available. Has there been any update since then regarding future plans or a roadmap to enable multilingual support? The customer is particularly interested in leveraging this feature in Japanese.

michaelmalc
ServiceNow Employee

Hi @hirokimarun - yes, the FAQ will be updated shortly. However, as of the July release, multi-lingual support for Guardian has been enabled for the following languages: English, French, Canadian French, German, Japanese, Dutch, Spanish, Brazilian Portuguese, and Italian. More to come to soon!

Vinod K_V
Tera Contributor

Hi All, are we allowed to create our own guardrails? TIA

michaelmalc
ServiceNow Employee

Hi @Vinod K_V

 

No - Guardian does not support the creation of custom guardrails at this time. 

mukesh_adh
Tera Contributor

In the product documentation and few other community posts I see "Offensiveness for HRSD".

But i do not see this in the instance.

Plugin Now Assist for HR Service Delivery (HRSD) - App id: sn_hr_gen_ai - is up to date (version 12.0.7).

 

Any suggestion?

 

Ref https://www.servicenow.com/docs/viewer/attachment/SRy58xeijUUqttfwLgfrHQ/gI78mdOQzhHuKqnnEt0sUw-SRy5...

Jayden4
Tera Contributor

Hey there, 

 

Is there a way to turn this off per skill rather than instance wide?

I am having issues where the OOB Data Privacy Invocation step is causing the request to be flagged as "Jailbreak Prompts/Prompt Injection" under Guardian due to the way it obscures the text, and as a result is blocking the OOB summary tools and they show an error on forms. 

 

This is for Now Assist for IRM. Has anyone come across this? How did you get around it?

PSP1001197
Tera Explorer

@Ashley Snyder Where Exactly we can see those codes ? Do we have any table that holds this categories? any where ?
Can we build report like False positive and False Negative for Guardian checks?

If anyone knows please help me on this it's urgent.

Version history
Last update:
2 weeks ago
Updated by: