How should GPU dependencies of self-hosted LLMs be modeled in AI Control Tower and CMDB?
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
28m ago
Hi community,
I’m looking for guidance on connecting a self-hosted LLM’s AI inventory in AI Control Tower to its runtime and underlying GPU infrastructure in CMDB.
I reviewed the AICT with CSDM white paper and inspected the collection scripts and transform maps in our installed AI Service Graph Connectors for Amazon and Microsoft. These helped clarify the Product Model, Digital Asset, and operational CI layers. However, I’m still unclear about the recommended modeling of GPU dependencies.
Example environment
- Open WebUI: chat interface
- TensorFold: OpenAI-compatible local LLM inference API
- LLM: served through TensorFold
- Linux server with an NVIDIA GPU: hosts the applications and inference workload
My test environment uses DGX Spark, but the question applies more broadly to self-hosted LLMs, including servers with multiple GPUs or multiple inference services.
Current working model
| Local Chat system | AI System Digital Asset |
| LLM | AI Model Digital Asset, referenced by the AI System |
| Open WebUI | Application CI |
| TensorFold | AI & Model Application CI (cmdb_ci_appl_ai_application) |
| Deployed LLM instance | AI Model Deployment CI (cmdb_ci_ai_model_deployment) |
| Host | Linux Server CI |
| GPU | GPU CI (cmdb_ci_gpu) |
The Digital Assets reference their corresponding Product Models. The precise mapping from the AI System Digital Asset to its operational CI representation is still under consideration.
The current infrastructure relationships are:
- Applications → Runs on::Runs → Linux Server
- GPU → Used by::Uses → Linux Server
The GPU-to-server relationship follows the documented NVIDIA GPU Discovery pattern. I have not established a direct relationship between an AI Application or Model Deployment and a GPU.
Main question: how should the GPU dependency be represented?
Sharing a host establishes where the components reside, but it does not establish which inference runtime or model deployment uses a particular GPU. This becomes more relevant when a server has multiple GPUs, workloads use different GPU subsets, or some models run on CPU only.
What is the recommended way to trace an AI System or AI Model Digital Asset through its operational CIs to the GPU resources it depends on?
In particular:
Is the Linux host the intended connection point, or should there be a direct relationship to the GPU CI?
If a direct relationship is appropriate, should it originate from AI Application, AI Function, or AI Model Deployment? Which relationship type and direction should be used?Should GPU assignment be modeled as a persistent CMDB dependency, runtime telemetry, or both?
For example, how should we distinguish a configured GPU assignment from a model currently loaded into GPU memory or actively performing inference?Is there an out-of-the-box Discovery pattern, connector mapping, or reference blueprint covering this path?
The AICT/CSDM white paper explains the AI asset and operational layers, but I could not find a GPU-specific mapping in it.
Related modeling questions
To make the GPU relationship meaningful, I’d also appreciate clarification on the operational layer:
- For a self-hosted inference API, when should AI Application, AI Function, and AI Model Deployment coexist as separate records?
- Which operational entity should represent an AI System comprising both a chat interface and an inference runtime?
- Which Asset-CI references or relationship types should connect the AI System and AI Model Digital Assets to those operational entities?
I’m keeping deployment availability separate from observed execution: an API advertising a model, or model weights existing on disk, does not by itself establish active inference or GPU residency.
The goal is to support both AI governance and operational impact analysis—for example, understanding which AI systems and model deployments could be affected by a GPU failure—while maintaining the model through Discovery or a custom IRE source.
Implementation examples using actual tables, reference fields, and relationship types would be particularly helpful. Please also mention any relevant platform release or application-version requirements.
References:
