Interested in a ServiceNow event built for developers? Registration for now[dev]26 is officially open!

Nicola Attico
ServiceNow Employee

A world model predicts how an environment changes over time, including how actions might influence what happens next.

 

The term appears in reinforcement learning, robotics, video generation, and AI agents. Its meaning varies across these settings. Here, I’ll focus on a definition rooted in model-based reinforcement learning:

A world model is a learned model that predicts how an environment changes, including how it may change when an agent acts.

The “world” does not have to be the entire physical world. It can be a bounded environment: a robot’s surroundings, a simulated game, or an enterprise service undergoing incident resolution.

What matters is that the model captures how that environment evolves.

Let st represent the environment’s state at step t, and st+1 its state at the next step. A model can predict:

p(st+1 | st)

This asks: given the current state, which next states are likely?

Here, p represents a probability distribution, and the vertical bar means “given.” The model can assign probabilities to several possible outcomes. The environment may behave unpredictably, or the model may lack enough information to predict the outcome with certainty.

“State” means more than a status field such as “open” or “resolved.”

For an IT incident, a useful state representation might include its category, priority, description, work notes, and attachments such as documents and images.

Category and priority can be represented as structured values. Text and images can be converted into embeddings—numerical representations of their content. Together, they provide context about the situation we want to predict.

Now introduce an agent. Its actions can change the environment, so we need to include the proposed action, at:

p(st+1 | st, at)

The question becomes: given the current state and a chosen action, which next states are likely?

For an agent that uses tools, an action can specify a tool and its inputs:

at = (tool, inputs)

For example:

at = restart_service(service="SAP")

This is an illustrative tool call. The world model predicts possible consequences before the agent makes the call. These predictions can support comparisons between restarting the service, running a diagnostic check, and taking other available actions.

That comparison should also include doing nothing.

Let ∅ represent a no-op action. We can define the environment’s “free evolution” as:

pfree(st+1 | st) := p(st+1 | st, at = ∅)

The symbol := means “defined as.” The label “free” means that the agent does not intervene.

The environment may still change. Another process might complete, a service might recover, or a failure might worsen.

Doing nothing therefore provides a useful baseline: what is likely to happen without intervention, and how would an action change that expectation?

So far, we have assumed that the agent can observe the full state of the environment. In practice, it often sees only part of what is happening.

An incident record may describe symptoms without revealing their cause. A “service unavailable” alert tells the agent something is wrong, but leaves many possible explanations open.

This is the distinction between full and partial observability. The agent can be the same in both cases; what changes is the information available to it.

Under partial observability, the agent receives observations: alerts, user reports, tool results, and service-status checks. Let ot represent the observation at the current step.

A single observation may not provide enough context. Previous observations and actions can help.


Consider two situations:

  • The service is unavailable, and no recovery action has been attempted.
  • The service is unavailable, and a restart has already failed to help.

The current alert is identical. The evidence available for predicting what happens next is different.

We can represent that evidence as a history:

ht = (o1, a1, …, ot−1, at−1, ot)

The history contains observations and actions already taken, ending with the current observation, ot. The proposed next action, at, has not happened yet.

History helps the agent infer what may be happening now. A failed restart, for example, provides additional evidence about the underlying fault.

This does not mean that a complete current state would be insufficient for prediction. History is useful because the agent cannot directly observe that complete state. It helps fill the information gap, although uncertainty may remain.

The world model can then predict:

p(ot+1 | ht, at)

In plain language: given what has happened so far and a proposed action, what might the agent observe next?

The earlier equation predicted the next state of the environment. This equation predicts the next observation available to the agent.

That distinction matters in practical systems. Agents often assess progress through tool responses and monitoring signals, while the underlying state remains partly hidden.

Return to the unavailable service. The history shows that one restart has already failed. The agent can compare what it might observe after another restart, after a diagnostic check, or after no intervention.

Its decision can now draw on predictions about outcomes as well as descriptions of the tools available to it.

The world model predicts possible consequences. The agent compares those consequences against its goals.

These are separate responsibilities. Predicting what may happen does not, by itself, determine which outcome is desirable or which action is worth taking.

A world model can learn from historical observations, actions, and outcomes. Its predictions can also represent uncertainty.

For a numerical outcome such as time to resolution, the model might estimate an expected value and a variance describing the spread of possible results. For categorical outcomes such as “resolved” or “still unavailable,” it can assign a probability to each possibility.

These estimates help an agent compare actions while considering how uncertain their outcomes are.

The practical value of a world model lies in connecting the current situation, possible actions, and likely consequences.

Before an agent changes the world, a world model helps it predict what that change might bring.

Version history
Last update:
15m ago
Updated by: