ServiceNow AI Research

Benchmark

Context is Key: A Benchmark for Forecasting with Essential Textual Information
Forecasting is a critical task in decision making across various domains. While numerical data provides a foundation, it often lacks …
Benchmarking Bayesian Causal Discovery Methods for Downstream Treatment Effect Estimation
The practical utility of causality in decision-making is widely recognized, with causal discovery and inference being inherently …
Mastering the Unsupervised Reinforcement Learning Benchmark from Pixels
Controlling artificial agents from visual sensory data is an arduous task. Reinforcement learning (RL) algorithms can succeed but …
The Stack: 3 TB of permissively licensed source code
Large Language Models (LLMs) play an ever-increasing role in the field of Artificial Intelligence (AI)–not only for natural …
In search of robust measures of generalization
One of the principal scientific challenges in deep learning is explaining generalization, i.e., why the particular way the community …