{
    "version": "https://jsonfeed.org/version/1",
    "title": "Armada Documentation Blog",
    "home_page_url": "https://docs.armada.ai/bridge/blog",
    "description": "Armada Documentation Blog",
    "items": [
        {
            "id": "https://docs.armada.ai/bridge/blog/armada-nvidia-ai-grid",
            "content_html": "<p>Real-time AI is reshaping infrastructure requirements.</p>\n<p>Inference workloads such as conversational AI, real-time video generation, AR/XR streaming, visual search, and large-scale personalization demand ultra-low latency, predictable performance, and geographic proximity to users and data sources. Centralized AI factories remain essential for training, but for many AI-native services, inference at scale requires AI Grids: geographically distributed GPU infrastructure operating as a unified, policy-controlled system.</p>\n<p>Armada is collaborating with NVIDIA to enable NVIDIA AI Grid on Armada Edge Platform (AEP), providing telecommunications operators, service providers, and enterprises with a validated architecture for deploying and operating distributed AI infrastructure at global scale.</p>\n<p>This post explores the architecture and operational model behind that system.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"reference-architecture-nvidia-ai-grid--armada-edge-platform\">Reference Architecture: NVIDIA AI Grid + Armada Edge Platform<a href=\"https://docs.armada.ai/bridge/blog/armada-nvidia-ai-grid#reference-architecture-nvidia-ai-grid--armada-edge-platform\" class=\"hash-link\" aria-label=\"Direct link to Reference Architecture: NVIDIA AI Grid + Armada Edge Platform\" title=\"Direct link to Reference Architecture: NVIDIA AI Grid + Armada Edge Platform\" translate=\"no\">​</a></h2>\n<p>Armada Edge Platform is aligned with the NVIDIA AI Grid reference design and integrates with key NVIDIA technologies, including: NVIDIA RTX PRO Servers, NVIDIA HGX B200 systems, NVIDIA Spectrum-X Ethernet networking, NVIDIA BlueField DPUs, and NVIDIA AI Enterprise software.</p>\n<p>Together, these components form a distributed AI infrastructure stack designed for production-scale inference across centralized and edge environments.</p>\n<p>Where NVIDIA provides the accelerated compute, networking, and software stack, Armada provides the distributed control plane and operational layer required to manage this infrastructure coherently across thousands of locations.</p>\n<p><img decoding=\"async\" loading=\"lazy\" alt=\"Live view of AI Grid\" src=\"https://docs.armada.ai/assets/images/ai-grid-live-view-4a8140b82242764660967cc2757b55e2.png\" width=\"2500\" height=\"1405\" class=\"img_ev3q\">\n<em>Live view of AI Grid</em></p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-core-challenge-distributed-gpus-as-one-system\">The Core Challenge: Distributed GPUs as One System<a href=\"https://docs.armada.ai/bridge/blog/armada-nvidia-ai-grid#the-core-challenge-distributed-gpus-as-one-system\" class=\"hash-link\" aria-label=\"Direct link to The Core Challenge: Distributed GPUs as One System\" title=\"Direct link to The Core Challenge: Distributed GPUs as One System\" translate=\"no\">​</a></h2>\n<p>Deploying GPUs at multiple sites is not equivalent to operating a distributed AI platform.</p>\n<p>AI Grid deployments must support:</p>\n<ul>\n<li class=\"\">Unified lifecycle management</li>\n<li class=\"\">Deterministic workload placement</li>\n<li class=\"\">Resource-aware scheduling</li>\n<li class=\"\">Secure multi-tenancy</li>\n<li class=\"\">Policy-based network control</li>\n<li class=\"\">Centralized observability</li>\n<li class=\"\">Compliance-aware orchestration</li>\n</ul>\n<p>Armada Edge Platform provides a unified control plane spanning centralized AI factories, regional hubs, and edge locations.</p>\n<p><img decoding=\"async\" loading=\"lazy\" alt=\"Monitoring and Observability for AI Grid\" src=\"https://docs.armada.ai/assets/images/ai-grid-monitoring-3e1b1b01179ed605fd1fc0fd23d5a931.png\" width=\"2500\" height=\"1406\" class=\"img_ev3q\">\n<em>Monitoring and Observability for AI Grid</em></p>\n<p>Rather than treating each GPU site as an isolated cluster, AEP stitches distributed GPU infrastructure into a single operational domain. This enables:</p>\n<ul>\n<li class=\"\">Workload-aware and resource-aware orchestration across sites</li>\n<li class=\"\">Intelligent placement decisions based on latency, proximity, GPU utilization, cost, compliance, and policy</li>\n<li class=\"\">Consistent software lifecycle management across hundreds to thousands of locations</li>\n<li class=\"\">Centralized monitoring and observability across the AI Grid</li>\n</ul>\n<p>The result is a globally distributed GPU fabric that behaves like a coherent platform.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"orchestration-model-latency--and-policy-aware-placement\">Orchestration Model: Latency- and Policy-Aware Placement<a href=\"https://docs.armada.ai/bridge/blog/armada-nvidia-ai-grid#orchestration-model-latency--and-policy-aware-placement\" class=\"hash-link\" aria-label=\"Direct link to Orchestration Model: Latency- and Policy-Aware Placement\" title=\"Direct link to Orchestration Model: Latency- and Policy-Aware Placement\" translate=\"no\">​</a></h2>\n<p>AI inference workloads differ significantly from traditional cloud workloads. They are often latency-sensitive, data-locality constrained, burst-driven, GPU-intensive, and multi-tenant in nature.</p>\n<p>AEP's orchestration layer evaluates placement decisions across multiple dimensions:</p>\n<ul>\n<li class=\"\">Latency requirements</li>\n<li class=\"\">Proximity to users and data sources</li>\n<li class=\"\">Real-time GPU availability and utilization</li>\n<li class=\"\">Network characteristics</li>\n<li class=\"\">Cost models</li>\n<li class=\"\">Regulatory and compliance policies</li>\n</ul>\n<p>This enables deterministic and optimized inference placement across distributed sites.</p>\n<p>For example, conversational AI workloads can be pinned close to users for minimal response time, while batch inference or less sensitive workloads can be placed in regional hubs or centralized AI factories, all under a single operational framework.</p>\n<p><img decoding=\"async\" loading=\"lazy\" alt=\"Grid Capacity and Composition\" src=\"https://docs.armada.ai/assets/images/ai-grid-capacity-ef32603dc4e1c3e48f65e7615d98b394.png\" width=\"2500\" height=\"1387\" class=\"img_ev3q\">\n<em>Grid Capacity and Composition</em></p>\n<p>Each AI Grid location operates as a secure, multi-tenant environment.</p>\n<p>Armada platform layer supports Bare metal provisioning, virtual machines, storage and networking services, managed Kubernetes, Model-as-a-Service, managed SLURM, Jupyter notebooks, and end-to-end ML workflows.</p>\n<p>Isolation is enforced across CPU, GPU, network, and storage resources to ensure strong tenant separation, predictable performance, compliance alignment, and maximized GPU efficiency. This allows operators to safely expose distributed GPU capacity as monetizable services, including GPUaaS and inference platforms.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"network-integration-deterministic-ai-delivery\">Network Integration: Deterministic AI Delivery<a href=\"https://docs.armada.ai/bridge/blog/armada-nvidia-ai-grid#network-integration-deterministic-ai-delivery\" class=\"hash-link\" aria-label=\"Direct link to Network Integration: Deterministic AI Delivery\" title=\"Direct link to Network Integration: Deterministic AI Delivery\" translate=\"no\">​</a></h2>\n<p>AI Grid performance depends not just on GPUs but on network architecture.</p>\n<p>Armada Edge Platform integrates directly with the service provider's network layer to establish dedicated, policy-controlled connectivity from data sources to GPU workloads.</p>\n<p>This ensures predictable latency, secure data paths, traffic isolation, QoS enforcement, and end-to-end performance guarantees. By combining distributed GPU placement with network-aware orchestration, AI inference becomes a deterministic system rather than a best-effort deployment.</p>\n<p><img decoding=\"async\" loading=\"lazy\" alt=\"AI Grid Workload Allocation\" src=\"https://docs.armada.ai/assets/images/ai-grid-workload-allocation-c9110a9a0652a92694b401488d4e029f.png\" width=\"2500\" height=\"1413\" class=\"img_ev3q\">\n<em>AI Grid Workload Allocation</em></p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"physical-infrastructure-data-centers-as-ai-grid-foundation\">Physical Infrastructure: Data Centers as AI Grid Foundation<a href=\"https://docs.armada.ai/bridge/blog/armada-nvidia-ai-grid#physical-infrastructure-data-centers-as-ai-grid-foundation\" class=\"hash-link\" aria-label=\"Direct link to Physical Infrastructure: Data Centers as AI Grid Foundation\" title=\"Direct link to Physical Infrastructure: Data Centers as AI Grid Foundation\" translate=\"no\">​</a></h2>\n<p>Distributed AI requires standardized infrastructure at the physical layer. AEP supports AI Grid across brick-and-mortar data centers as well as Armada's Galleon modular data centers when data centers are unavailable or can't be built rapidly enough.</p>\n<p>Galleon, Armada's modular data center platform, provides a ruggedized, rapidly deployable, high-density AI infrastructure foundation for AI Grid deployments.</p>\n<p>Galleon integrates Power systems, Cooling, Networking, Compute and Storage into a standardized, edge-ready form factor. This enables accelerated deployment timelines, consistent hardware profiles across sites, repeatable rollout models along with edge and remote environment operations.</p>\n<p>When combined with AEP's control plane, Galleon allows operators to treat distributed AI infrastructure as a scalable system rather than a collection of bespoke deployments.</p>\n<p><img decoding=\"async\" loading=\"lazy\" alt=\"AI Grid workload allocation\" src=\"https://docs.armada.ai/assets/images/ai-grid-workload-allocation-2-883eb7307175e7458ea4aceb70ecee50.png\" width=\"2500\" height=\"1423\" class=\"img_ev3q\">\n<em>AI Grid workload allocation</em></p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"from-gpu-infrastructure-to-revenue-infrastructure\">From GPU Infrastructure to Revenue Infrastructure<a href=\"https://docs.armada.ai/bridge/blog/armada-nvidia-ai-grid#from-gpu-infrastructure-to-revenue-infrastructure\" class=\"hash-link\" aria-label=\"Direct link to From GPU Infrastructure to Revenue Infrastructure\" title=\"Direct link to From GPU Infrastructure to Revenue Infrastructure\" translate=\"no\">​</a></h2>\n<p>The transition to AI Grids is not solely a technical evolution but it is an economic one.</p>\n<p>Telecommunications operators and service providers can leverage distributed GPU infrastructure to:</p>\n<ul>\n<li class=\"\">Offer low-latency AI inference services</li>\n<li class=\"\">Support enterprise AI workloads at the edge</li>\n<li class=\"\">Enable real-time consumer AI applications</li>\n<li class=\"\">Provide GPUaaS across geographic markets</li>\n</ul>\n<p>Armada provides the operational layer that transforms geographically distributed GPU deployments into unified, revenue-generating AI platforms, including sovereign GPU cloud.</p>\n<p><img decoding=\"async\" loading=\"lazy\" alt=\"AI Grid Revenue Metrics and Cost Optimization\" src=\"https://docs.armada.ai/assets/images/ai-grid-revenue-metrics-3b98664599331eaee9cc494c1ee8d7aa.png\" width=\"2500\" height=\"1403\" class=\"img_ev3q\">\n<em>AI Grid Revenue Metrics and Cost Optimization</em></p>\n<p>By abstracting complexity across hardware, networking, orchestration, and lifecycle management, AEP enables operators to scale AI infrastructure from dozens to thousands of sites without multiplying operational overhead.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-architecture-of-distributed-ai-at-scale\">The Architecture of Distributed AI at Scale<a href=\"https://docs.armada.ai/bridge/blog/armada-nvidia-ai-grid#the-architecture-of-distributed-ai-at-scale\" class=\"hash-link\" aria-label=\"Direct link to The Architecture of Distributed AI at Scale\" title=\"Direct link to The Architecture of Distributed AI at Scale\" translate=\"no\">​</a></h2>\n<p>The next phase of AI infrastructure is defined by distribution, orchestration intelligence, and operational consistency.</p>\n<p>NVIDIA AI Grid provides the accelerated computing foundation. Armada provides the distributed control plane and a physical deployment model through its Galleon and AEP.</p>\n<p>Together, this architecture enables AI infrastructure that is geographically distributed, operationally unified, network-aware, secure and multi-tenant, and policy-driven that is scalable to thousands of locations.</p>\n<p>Finally, distributed inference is no longer an edge experiment and has become core infrastructure. The AI Grid era is here, and it requires systems built to operate everywhere, as one.</p>\n<p><a href=\"https://armada.ai/demo\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">To learn more, schedule a demo.</a></p>",
            "url": "https://docs.armada.ai/bridge/blog/armada-nvidia-ai-grid",
            "title": "Operationalizing Distributed AI: Armada and NVIDIA AI Grid",
            "summary": "Real-time AI is reshaping infrastructure requirements.",
            "date_modified": "2026-03-17T00:00:00.000Z",
            "author": {
                "name": "Anish Swaminathan"
            },
            "tags": [
                "nvidia",
                "ai-grid",
                "distributed-ai",
                "inference",
                "edge",
                "orchestration"
            ]
        },
        {
            "id": "https://docs.armada.ai/bridge/blog/nvidia-dsx-air",
            "content_html": "<p>Armada has been an NVIDIA DSX Air user since late 2024, and we have derived significant benefits from its ability to simulate Spectrum-X Ethernet environments for both internal development and customer proof-of-concept initiatives. NVIDIA DSX Air has enabled us to validate networking configurations and topologies, test multi-tenant configurations, and accelerate Bridge software deployments (Bridge is an on-prem Armada software product that provides multi-tenancy and cloud services on GPU hardware) without relying exclusively on physical hardware.</p>\n<p>We are excited about the launch of NVIDIA DSX Air and the expanded AI Factory digital simulation capabilities it introduces. This next evolution unlocks multiple use cases for Armada and our customers — driving measurable improvements across development velocity, PoC efficiency, operational stability, OPEX reductions and go-to-market. The key benefits are described below.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"1-development-velocity-without-hardware-bottlenecks\">1. Development Velocity Without Hardware Bottlenecks<a href=\"https://docs.armada.ai/bridge/blog/nvidia-dsx-air#1-development-velocity-without-hardware-bottlenecks\" class=\"hash-link\" aria-label=\"Direct link to 1. Development Velocity Without Hardware Bottlenecks\" title=\"Direct link to 1. Development Velocity Without Hardware Bottlenecks\" translate=\"no\">​</a></h2>\n<p>If you're building GPU cloud management software like Bridge, you are constantly running into hardware constraints. Testing against large-scale topologies, DPUs and advanced NVLink fabrics, as well as next-gen GPUs isn't trivial.</p>\n<p>With NVIDIA DSX Air, we can simulate full AI Factory environments — including GPUs, NVLink, Spectrum-X Ethernet switches, ConnectX SuperNICs, and BlueField DPUs — while continuing to validate against physical production-grade clusters.</p>\n<p>That means:</p>\n<ul>\n<li class=\"\">No lab contention</li>\n<li class=\"\">No waiting for hardware allocations</li>\n<li class=\"\">No waiting for the latest generation GPU</li>\n<li class=\"\">Ability to test at scale (including configurations customers may run before we ever see them)</li>\n</ul>\n<p>For Bridge, this translates to faster releases, better validation, and stronger resiliency. Simulation accelerates validation, while hardware testing ensures production-grade performance and interoperability.</p>\n<p>With the advent of AI Grid, NVIDIA DSX Air will become even more valuable. We will be able to simulate multiple edge sites and rapidly develop our policy-based application placement capabilities.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"2-removing-friction-from-bridge-evaluations\">2. Removing Friction from Bridge Evaluations<a href=\"https://docs.armada.ai/bridge/blog/nvidia-dsx-air#2-removing-friction-from-bridge-evaluations\" class=\"hash-link\" aria-label=\"Direct link to 2. Removing Friction from Bridge Evaluations\" title=\"Direct link to 2. Removing Friction from Bridge Evaluations\" translate=\"no\">​</a></h2>\n<p>Today, most AI factory or multi-tenancy PoCs depend on real infrastructure. That slows everything down:</p>\n<ul>\n<li class=\"\">Provision racks</li>\n<li class=\"\">Provide access</li>\n<li class=\"\">Isolate environments</li>\n<li class=\"\">Justify costs</li>\n</ul>\n<p>NVIDIA DSX Air breaks that dependency. Armada can provision simulated AI factory environments where customers can evaluate Bridge and its numerous capabilities:</p>\n<ul>\n<li class=\"\">Multi-tenancy policies</li>\n<li class=\"\">Hard isolation between tenants</li>\n<li class=\"\">IaaS (BMaaS, VMaaS, storage, network)</li>\n<li class=\"\">PaaS (managed Kubernetes)</li>\n<li class=\"\">AIaaS (LLMaaS, Jupyter Notebooks, managed SLURM, Kubeflow)</li>\n<li class=\"\">Billing</li>\n<li class=\"\">User management</li>\n<li class=\"\">NVAIE and 3rd party AI software</li>\n</ul>\n<p>This allows early-stage validation without waiting on hardware availability, compressing PoC timelines from months to weeks — before transitioning to hardware-based validation with greater confidence. That shortens sales cycles for us. More importantly, it lowers psychological friction. When customers can \"try before they rack,\" adoption accelerates. Once this phase completes, the customer can ultimately move to a hardware based PoC.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"3-production-digital-twins-for-operational-safety\">3. Production Digital Twins for Operational Safety<a href=\"https://docs.armada.ai/bridge/blog/nvidia-dsx-air#3-production-digital-twins-for-operational-safety\" class=\"hash-link\" aria-label=\"Direct link to 3. Production Digital Twins for Operational Safety\" title=\"Direct link to 3. Production Digital Twins for Operational Safety\" translate=\"no\">​</a></h2>\n<p>This is where the long-term value resides. Bridge customers operating AI factories at scale deal with:</p>\n<ul>\n<li class=\"\">Complex tenant segmentation</li>\n<li class=\"\">Reserved vs on-demand workloads</li>\n<li class=\"\">Brownfield network configurations</li>\n</ul>\n<p>Making changes in production is risky. With NVIDIA DSX Air, customers can build a high-fidelity digital twin of their environment and:</p>\n<ul>\n<li class=\"\">Confirm that Bridge coexists with manual configurations</li>\n<li class=\"\">Test new configuration changes safely</li>\n<li class=\"\">Validate policy updates</li>\n<li class=\"\">Simulate tenant onboarding</li>\n<li class=\"\">Expand model capacity expansion</li>\n<li class=\"\">Run and identify failure scenarios</li>\n<li class=\"\">Detect configuration drift</li>\n</ul>\n<p>NVIDIA DSX Air features extensibility, which will allow us to model our modular datacenters branded Galleon. With Galleon modeling, we will be able to use our datacenter infrastructure management software to simulate actions such as changing chillers on the digital twin before applying the same action to the physical environment.</p>\n<p>Instead of configure and hope, customers can simulate and verify. For brownfield environments especially — where some multi-tenancy rules were implemented manually — this reduces the fear of Bridge \"messing up\" existing configurations. Customers can easily reduce costly configuration errors that can result in six-figure downtime events.</p>\n<p>Operational risk is one of the biggest hidden OPEX drivers in AI infrastructure. Digital twins reduce that risk.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"4-a-new-distribution-channel-for-bridge\">4. A New Distribution Channel for Bridge<a href=\"https://docs.armada.ai/bridge/blog/nvidia-dsx-air#4-a-new-distribution-channel-for-bridge\" class=\"hash-link\" aria-label=\"Direct link to 4. A New Distribution Channel for Bridge\" title=\"Direct link to 4. A New Distribution Channel for Bridge\" translate=\"no\">​</a></h2>\n<p>The NVIDIA DSX Air blueprint marketplace introduces another dimension. Bridge based pre-built blueprints will be instantiated within the NVIDIA DSX Air environment; customers can experience our value immediately at a fraction of the cost compared to the hardware — integrated with simulated GPU, networking, storage, and security stacks. Now we go beyond simulation to becoming discoverable.</p>\n<p>It positions Bridge inside the AI factory design phase — not after hardware has already been deployed.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-bigger-picture\">The Bigger Picture<a href=\"https://docs.armada.ai/bridge/blog/nvidia-dsx-air#the-bigger-picture\" class=\"hash-link\" aria-label=\"Direct link to The Bigger Picture\" title=\"Direct link to The Bigger Picture\" translate=\"no\">​</a></h2>\n<p>For Bridge, NVIDIA DSX Air unlocks:</p>\n<ul>\n<li class=\"\">Faster development</li>\n<li class=\"\">Faster sales cycles</li>\n<li class=\"\">Safer production evolution</li>\n<li class=\"\">Stronger ecosystem presence</li>\n</ul>\n<p>AI factories are the future of infrastructure, and high-fidelity digital twins will become standard practice. And Bridge will be native to that future — not bolted on after the racks are live.</p>",
            "url": "https://docs.armada.ai/bridge/blog/nvidia-dsx-air",
            "title": "How NVIDIA DSX Air Reduces Dev/Test Costs, Accelerates PoCs, and Lowers Production Risk for Armada and Its Customers",
            "summary": "Armada has been an NVIDIA DSX Air user since late 2024, and we have derived significant benefits from its ability to simulate Spectrum-X Ethernet environments for both internal development and customer proof-of-concept initiatives. NVIDIA DSX Air has enabled us to validate networking configurations and topologies, test multi-tenant configurations, and accelerate Bridge software deployments (Bridge is an on-prem Armada software product that provides multi-tenancy and cloud services on GPU hardware) without relying exclusively on physical hardware.",
            "date_modified": "2026-03-16T00:00:00.000Z",
            "author": {
                "name": "Pavan Samudrala"
            },
            "tags": [
                "nvidia",
                "dsx-air",
                "simulation",
                "digital-twin",
                "ai-factory",
                "spectrum-x"
            ]
        },
        {
            "id": "https://docs.armada.ai/bridge/blog/distributed-ai-edge",
            "content_html": "<p>All AI is not created equal. While centralized inference serves some use-cases well where long thinking times are acceptable, new use cases such as physical AI, real-time agentic AI chatbots, digital avatars doing real time dialog, and computer vision require faster response times. It is not just about network latency, but compute latency becomes important, mandating computation closer to data sources, and lower bandwidth usage across the network in order to scale cost effectively.</p>\n<p>These applications can't tolerate the latency of round trips to centralized data centers nor can they afford the cost of constantly transferring large volumes of data. Instead, they require inference that is geographically distributed, dynamically orchestrated, and tightly optimized for latency and bandwidth.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-growing-need-for-edge-inference\">The Growing Need for Edge Inference<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#the-growing-need-for-edge-inference\" class=\"hash-link\" aria-label=\"Direct link to The Growing Need for Edge Inference\" title=\"Direct link to The Growing Need for Edge Inference\" translate=\"no\">​</a></h2>\n<p>This is fueling a surge in demand for distributed inference infrastructure—capable of running AI models across clusters of GPUs residing at regional data centers and edge sites, while maintaining cloud-like flexibility and scale. The distributed inference market is poised for exceptional growth between 2025 and 2030, with projections indicating an expansion from USD 106.15 billion to USD 254.98 billion at a CAGR of 19.2%.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"why-nvidia-mgx-servers-are-a-game-changer-for-edge-inference\">Why NVIDIA MGX Servers Are a Game-Changer for Edge Inference<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#why-nvidia-mgx-servers-are-a-game-changer-for-edge-inference\" class=\"hash-link\" aria-label=\"Direct link to Why NVIDIA MGX Servers Are a Game-Changer for Edge Inference\" title=\"Direct link to Why NVIDIA MGX Servers Are a Game-Changer for Edge Inference\" translate=\"no\">​</a></h2>\n<p>NVIDIA MGX servers, based on a modular reference design, can be used for a wide variety of use cases, from compute-intensive datacenter to edge workloads. MGX provides a new standard for modular server design by improving ROI and reducing time to market and is especially suited to distributed inference. Some of the reasons for this are:</p>\n<ul>\n<li class=\"\"><strong>Modular design</strong> allows core and edge sites to scale from 1 RU to multiple racks of servers</li>\n<li class=\"\"><strong>High performance per watt</strong> allows maximum GPU compute capacity to be deployed at the distributed inference site</li>\n<li class=\"\"><strong>Integration with NVIDIA Cloud Functions (NVCF) and NVIDIA NIM</strong>, both included in the NVIDIA AI Enterprise suite, providing access to a large number of vertically oriented models and solutions</li>\n</ul>\n<p>When combined with NVIDIA Spectrum-X Ethernet networking platform for AI, customers can extract the full performance of the underlying GPUs.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"challenges-in-building-an-edge-inference-stack\">Challenges in Building an Edge Inference Stack<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#challenges-in-building-an-edge-inference-stack\" class=\"hash-link\" aria-label=\"Direct link to Challenges in Building an Edge Inference Stack\" title=\"Direct link to Challenges in Building an Edge Inference Stack\" translate=\"no\">​</a></h2>\n<p>While MGX servers along with Spectrum-X and NVIDIA AI Enterprise offer an integrated solution stack, distributed inference presents a number of infrastructure challenges for a GPU-as-a-Service (GPUaaS) provider:</p>\n<ul>\n<li class=\"\"><strong>Managing multiple sites</strong>: Distributed GPUaaS providers typically have multiple sites that are often in light-out environments. The infrastructure consisting of compute, storage, networking, and WAN gateways has to be managed remotely with the lowest possible OPEX</li>\n<li class=\"\"><strong>Managing isolation between multiple tenants</strong>: Distributed GPU sites have multiple tenants that demand the highest level of security between tenants</li>\n<li class=\"\"><strong>Matching workloads to the correct GPU site</strong>: Workloads have to be mapped to the appropriate site for latency, bandwidth, compliance, or data gravity reasons</li>\n<li class=\"\"><strong>Maximizing utilization</strong>: Given the high cost of GPUs, utilization has to be as close to 100% as possible at all times</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-need-for-secure-dynamic-tenancy-and-isolation\">The Need for Secure, Dynamic Tenancy and Isolation<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#the-need-for-secure-dynamic-tenancy-and-isolation\" class=\"hash-link\" aria-label=\"Direct link to The Need for Secure, Dynamic Tenancy and Isolation\" title=\"Direct link to The Need for Secure, Dynamic Tenancy and Isolation\" translate=\"no\">​</a></h2>\n<p>The above challenges require a secure and dynamic tenancy software layer for distributed inference. The ideal software solution must offer:</p>\n<ul>\n<li class=\"\">Zero touch management of the underlying hardware infrastructure potentially across 10,000s edge and core sites to slash OPEX</li>\n<li class=\"\">Isolation between tenants for security and compliance</li>\n<li class=\"\">Dynamic resource scaling for maximizing GPU utilization</li>\n<li class=\"\">Registration of underutilized resources with NVCF for maximizing GPU utilization</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"introducing-bridge-gpu-cloud-management-software\">Introducing Bridge GPU Cloud Management Software<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#introducing-bridge-gpu-cloud-management-software\" class=\"hash-link\" aria-label=\"Direct link to Introducing Bridge GPU Cloud Management Software\" title=\"Direct link to Introducing Bridge GPU Cloud Management Software\" translate=\"no\">​</a></h2>\n<p>Bridge GPU CMS provides the following functionality:</p>\n<ul>\n<li class=\"\"><strong>On-demand isolation</strong> spanning CPU, GPU, network, storage, and the WAN gateway</li>\n<li class=\"\"><strong>Bare metal, virtual machine, or container instances</strong></li>\n<li class=\"\"><strong>Automated infrastructure management</strong> for tenants with scale-out and scale-in</li>\n<li class=\"\"><strong>Admin functionality</strong> to discover, observe, and manage the underlying hardware across 10,000s of sites</li>\n<li class=\"\"><strong>Billing and User management</strong> with RBAC</li>\n<li class=\"\"><strong>Integration with open source</strong> (Ray, vLLM) or 3rd party PaaS (Red Hat OpenShift and more)</li>\n<li class=\"\"><strong>Integration with NVIDIA Cloud Functions (NVCF)</strong> to monetize unused capacity</li>\n<li class=\"\"><strong>Centralized Management</strong> for managing and orchestrating multiple Edge locations</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"reference-architecture-nvidia--bridge-for-edge-inference\">Reference Architecture: NVIDIA + Bridge for Edge Inference<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#reference-architecture-nvidia--bridge-for-edge-inference\" class=\"hash-link\" aria-label=\"Direct link to Reference Architecture: NVIDIA + Bridge for Edge Inference\" title=\"Direct link to Reference Architecture: NVIDIA + Bridge for Edge Inference\" translate=\"no\">​</a></h2>\n<p>Bridge GPU CMS when coupled with NVIDIA MGX, Spectrum-X, and NVIDIA AI Enterprise solves the above-listed problems for GPUaaS providers. The components for this architecture are:</p>\n<ul>\n<li class=\"\">NVIDIA MGX servers with NVIDIA Bluefield-3 DPU or CX7 ethernet card</li>\n<li class=\"\">NVIDIA Spectrum switches for communication (East-West and North-South)</li>\n<li class=\"\">NVIDIA Spectrum switches for OOB Management</li>\n<li class=\"\">NVIDIA AI Enterprise (NVIDIA NIM, NVCF)</li>\n<li class=\"\">Optional NVIDIA Quantum InfiniBand switches for East-West communication</li>\n<li class=\"\">External High Performance Storage (HPS) from partner solutions</li>\n<li class=\"\">Bridge GPU Cloud Management Software (CMS)</li>\n</ul>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"installation\">Installation<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#installation\" class=\"hash-link\" aria-label=\"Direct link to Installation\" title=\"Direct link to Installation\" translate=\"no\">​</a></h3>\n<p>The infrastructure is installed at the distributed edge-core locations, along with other software components including Bridge GPU CMS, and all the hardware related tests are performed, before onboarding the resources.</p>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"onboarding\">Onboarding<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#onboarding\" class=\"hash-link\" aria-label=\"Direct link to Onboarding\" title=\"Direct link to Onboarding\" translate=\"no\">​</a></h3>\n<p>Once this is completed, the site administrator discovers the infrastructure using Bridge GPU CMS, and creates the underlay network using the Spectrum switches and the network adapters or DPUs on the MGX servers. The Admin then goes on to create tenants, which are the logical entities that run different types of workloads on the same physical infrastructure.</p>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"isolation\">Isolation<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#isolation\" class=\"hash-link\" aria-label=\"Direct link to Isolation\" title=\"Direct link to Isolation\" translate=\"no\">​</a></h3>\n<p>The important consideration while allocating these resources is that they need to be fully isolated, so that each tenant's workload can run without any performance or security implication from other tenants. Bridge GPU CMS ensures this by providing hard multi-tenancy at all levels - CPU, GPU, memory, network adapters, network switches, internal and external storage, all the way to the external gateway.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"conclusion-the-path-to-scalable-multi-tenant-distributed-ai\">Conclusion: The Path to Scalable, Multi-Tenant, Distributed AI<a href=\"https://docs.armada.ai/bridge/blog/distributed-ai-edge#conclusion-the-path-to-scalable-multi-tenant-distributed-ai\" class=\"hash-link\" aria-label=\"Direct link to Conclusion: The Path to Scalable, Multi-Tenant, Distributed AI\" title=\"Direct link to Conclusion: The Path to Scalable, Multi-Tenant, Distributed AI\" translate=\"no\">​</a></h2>\n<p>The next wave of AI adoption depends on pushing inference closer to where data is generated—at the edge. Achieving this requires more than raw compute; it calls for an architecture that delivers secure multi-tenancy, dynamic scaling, high utilization, and seamless integration with cloud-native AI services.</p>\n<p>NVIDIA MGX servers, combined with Spectrum-X networking and NVIDIA AI Enterprise, provide the performance and flexibility needed for distributed Edge deployments. Layered with Bridge GPU Cloud Management Software, organizations gain the management and orchestration capabilities essential for turning this distributed infrastructure into a scalable, revenue-generating service.</p>\n<p>The opportunity is clear: distributed inference is becoming central to how next-generation AI will be delivered. Now is the time to explore, pilot, and engage to unlock these capabilities and be part of the ecosystem shaping the future of AI at the edge.</p>",
            "url": "https://docs.armada.ai/bridge/blog/distributed-ai-edge",
            "title": "Delivering Distributed AI at the Edge with Bridge",
            "summary": "All AI is not created equal. While centralized inference serves some use-cases well where long thinking times are acceptable, new use cases such as physical AI, real-time agentic AI chatbots, digital avatars doing real time dialog, and computer vision require faster response times. It is not just about network latency, but compute latency becomes important, mandating computation closer to data sources, and lower bandwidth usage across the network in order to scale cost effectively.",
            "date_modified": "2025-09-02T00:00:00.000Z",
            "author": {
                "name": "Amar Kapadia"
            },
            "tags": [
                "edge",
                "distributed-ai",
                "inference",
                "nvidia-mgx"
            ]
        },
        {
            "id": "https://docs.armada.ai/bridge/blog/nvis-topology-onboarding",
            "content_html": "<p>We recently collaborated with the NVIDIA Infrastructure Specialist (NVIS) team to onboard and validate a complex metadata topology deployed by NVIS into our Bridge GPU Cloud Management Software (CMS). This activity demonstrates how Bridge GPU CMS can take over an NVIS deployed GPU topology and then perform day 1, 2 activities such as discovery, dynamic multi-tenancy, observability, fault management, and more.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-challenge\">The Challenge<a href=\"https://docs.armada.ai/bridge/blog/nvis-topology-onboarding#the-challenge\" class=\"hash-link\" aria-label=\"Direct link to The Challenge\" title=\"Direct link to The Challenge\" translate=\"no\">​</a></h2>\n<p>The initial deployment and configuration of GPU hardware for NVIDIA Cloud Partners (NCPs) and subsequent management is often done by different entities. Case in point, Day 0 tasks for NCP GPU environments are often performed by NVIS. After NVIS hands over the cluster to the NCP effectively with one single tenant, the task of multi-tenancy and other Day 1, 2 tasks can be performed by Bridge GPU CMS.</p>\n<p>In other words, there is a hand-off from an NVIS deployed topology to our GPU CMS. NCPs and GPU-as-a-service providers needed a robust and automated method to:</p>\n<ul>\n<li class=\"\">Onboard a topology created by NVIS onto Bridge GPU CMS</li>\n<li class=\"\">Validate that the metadata topology files created by the NVIDIA NVIS team after deploying the hardware are correctly and completely onboarded</li>\n<li class=\"\">Efficiently provision underlay and overlay network configurations for onboarding infrastructure tenants</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"validation--onboarding-with-bridge-gpu-cms\">Validation &amp; Onboarding with Bridge GPU CMS<a href=\"https://docs.armada.ai/bridge/blog/nvis-topology-onboarding#validation--onboarding-with-bridge-gpu-cms\" class=\"hash-link\" aria-label=\"Direct link to Validation &amp; Onboarding with Bridge GPU CMS\" title=\"Direct link to Validation &amp; Onboarding with Bridge GPU CMS\" translate=\"no\">​</a></h2>\n<p>We validated the successful hand-off of a 16 SU topology deployed by NVIS to Bridge GPU CMS. This validation was performed on NVIDIA Air.</p>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"step-by-step-workflow\">Step-by-Step Workflow<a href=\"https://docs.armada.ai/bridge/blog/nvis-topology-onboarding#step-by-step-workflow\" class=\"hash-link\" aria-label=\"Direct link to Step-by-Step Workflow\" title=\"Direct link to Step-by-Step Workflow\" translate=\"no\">​</a></h3>\n<ol>\n<li class=\"\"><strong>Metadata Onboarding</strong>: Imported the NVIS metadata topology file into Bridge GPU CMS</li>\n<li class=\"\"><strong>RA Compliance Validation</strong>: Automatically validated the metadata against RA compliance rules. Non-compliance feedback was immediately provided to the user with actionable insights</li>\n<li class=\"\"><strong>Topology Discovery</strong>: Dynamically discovered all underlying topology nodes (compute, network, and storage) referenced in the metadata</li>\n<li class=\"\"><strong>Underlay Configuration</strong>: Configured network underlay settings for discovered nodes, ensuring base connectivity across the infrastructure</li>\n<li class=\"\"><strong>Tenant Overlay Creation</strong>: Built tenant-specific overlay networks, enabling scalable multi-tenant operations on top of the validated infrastructure</li>\n</ol>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"impact\">Impact<a href=\"https://docs.armada.ai/bridge/blog/nvis-topology-onboarding#impact\" class=\"hash-link\" aria-label=\"Direct link to Impact\" title=\"Direct link to Impact\" translate=\"no\">​</a></h2>\n<p>By automating and validating the metadata topology through Bridge GPU CMS, the NCPs can achieve:</p>\n<ul>\n<li class=\"\"><strong>Clean hand-off from NVIS to Bridge GPU CMS</strong></li>\n<li class=\"\"><strong>Faster deployment readiness</strong></li>\n<li class=\"\"><strong>Improved reliability of infrastructure metadata</strong></li>\n<li class=\"\"><strong>Streamlined compliance checks</strong>, reducing engineering effort</li>\n</ul>\n<p>This use case illustrates how Bridge GPU CMS can successfully onboard a GPU topology deployed by NVIS. This validation is very important for NCPs as they require a clean hand-off between Day 0 to Day 1,2 activities without any disruptions.</p>",
            "url": "https://docs.armada.ai/bridge/blog/nvis-topology-onboarding",
            "title": "Onboarding NVIDIA NVIS Deployed GPU Topology with Bridge GPU CMS",
            "summary": "We recently collaborated with the NVIDIA Infrastructure Specialist (NVIS) team to onboard and validate a complex metadata topology deployed by NVIS into our Bridge GPU Cloud Management Software (CMS). This activity demonstrates how Bridge GPU CMS can take over an NVIS deployed GPU topology and then perform day 1, 2 activities such as discovery, dynamic multi-tenancy, observability, fault management, and more.",
            "date_modified": "2025-07-02T00:00:00.000Z",
            "author": {
                "name": "Namachi Sankaranarayanan"
            },
            "tags": [
                "nvidia",
                "nvis",
                "topology",
                "onboarding"
            ]
        },
        {
            "id": "https://docs.armada.ai/bridge/blog/ddn-exascaler-integration",
            "content_html": "<p>Managing external storage for GPU-accelerated AI workloads can be complex—especially when ensuring that storage volumes are provisioned correctly, isolated per tenant, and automatically mounted to the right compute nodes. With Bridge GPU Cloud Management Software (GPU CMS), this entire process is streamlined through seamless integration with DDN EXAScaler.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"end-to-end-automation-with-no-manual-steps\">End-to-End Automation with No Manual Steps<a href=\"https://docs.armada.ai/bridge/blog/ddn-exascaler-integration#end-to-end-automation-with-no-manual-steps\" class=\"hash-link\" aria-label=\"Direct link to End-to-End Automation with No Manual Steps\" title=\"Direct link to End-to-End Automation with No Manual Steps\" translate=\"no\">​</a></h2>\n<p>With Bridge GPU CMS, end users don't need to manually log into multiple systems, configure storage mounts, or worry about compatibility between compute and storage. The DDN EXAScaler integration is fully automated—allowing users to simply specify:</p>\n<ul>\n<li class=\"\">The desired storage size</li>\n<li class=\"\">The bare metal node where the storage should be mounted</li>\n</ul>\n<p>Everything else—from tenant-aware provisioning, storage policy enforcement, network isolation, to automatic mount point creation—is handled seamlessly by Bridge GPU CMS.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"simple-and-efficient-flow\">Simple and Efficient Flow<a href=\"https://docs.armada.ai/bridge/blog/ddn-exascaler-integration#simple-and-efficient-flow\" class=\"hash-link\" aria-label=\"Direct link to Simple and Efficient Flow\" title=\"Direct link to Simple and Efficient Flow\" translate=\"no\">​</a></h2>\n<p>The process starts with the NCP admin (cloud provider admin) importing the entire GPU infrastructure (compute, storage, E-W network, N-S network) into the software and setting up a new tenant. Once the tenant is created, the tenant user can allocate a GPU bare-metal or VM instance and request external storage from DDN.</p>\n<p>The tenant simply provides:</p>\n<ul>\n<li class=\"\">The desired storage size</li>\n<li class=\"\">The specific compute node where the storage should be mounted</li>\n</ul>\n<p>Once these inputs are provided, Bridge GPU CMS handles all interactions with DDN, including:</p>\n<ul>\n<li class=\"\">Configuring storage volumes</li>\n<li class=\"\">Assigning tenant-specific quotas</li>\n<li class=\"\">Creating the mount point</li>\n<li class=\"\">Ensuring the mount point is immediately available on the compute node</li>\n<li class=\"\">Tenant specific North-South network isolation for accessing the DDN storage nodes</li>\n</ul>\n<p>This zero-touch integration eliminates any need for the tenant to interact with the DDN portal directly.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"real-time-validation-across-systems\">Real-Time Validation Across Systems<a href=\"https://docs.armada.ai/bridge/blog/ddn-exascaler-integration#real-time-validation-across-systems\" class=\"hash-link\" aria-label=\"Direct link to Real-Time Validation Across Systems\" title=\"Direct link to Real-Time Validation Across Systems\" translate=\"no\">​</a></h2>\n<p>To ensure transparency and operational assurance, the NCP admin or tenant admin can view all configured storage volumes directly within Bridge GPU CMS. For additional verification, they can also cross-check the automatically created tenants, networks, policies, and mount points directly in the DDN admin portal.</p>\n<p>All configurations are performed via APIs with no manual intervention.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"full-tenant-experience\">Full Tenant Experience<a href=\"https://docs.armada.ai/bridge/blog/ddn-exascaler-integration#full-tenant-experience\" class=\"hash-link\" aria-label=\"Direct link to Full Tenant Experience\" title=\"Direct link to Full Tenant Experience\" translate=\"no\">​</a></h2>\n<p>Once the storage is provisioned, the tenant user can log directly into their allocated GPU compute node and immediately access the mounted DDN EXAScaler storage volume. Whether for large-scale AI training data or model checkpoints or inference, this automated mount ensures data is available where and when the user needs it.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"key-benefits\">Key Benefits<a href=\"https://docs.armada.ai/bridge/blog/ddn-exascaler-integration#key-benefits\" class=\"hash-link\" aria-label=\"Direct link to Key Benefits\" title=\"Direct link to Key Benefits\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><strong>End-to-End Automation</strong>: No manual steps—just specify size and compute node, and Bridge GPU CMS handles everything else</li>\n<li class=\"\"><strong>Single Pane of Glass</strong>: Both compute and storage provisioning are managed from a single interface</li>\n<li class=\"\"><strong>Full Tenant Isolation</strong>: Each tenant's storage is isolated with network policies</li>\n<li class=\"\"><strong>Quota Management</strong>: Each tenant user's storage quota can be managed by the tenant admin</li>\n<li class=\"\"><strong>Real-Time Observability</strong>: Both admins and tenants can view and validate storage allocations directly within Bridge GPU CMS portal</li>\n<li class=\"\"><strong>API-Driven Consistency</strong>: All configurations—from mount points to network overlays—are performed through automated APIs, ensuring accuracy and compliance with tenant policies</li>\n</ul>",
            "url": "https://docs.armada.ai/bridge/blog/ddn-exascaler-integration",
            "title": "Seamless Integration of Bridge with DDN EXAScaler for High-Performance AI Workloads",
            "summary": "Managing external storage for GPU-accelerated AI workloads can be complex—especially when ensuring that storage volumes are provisioned correctly, isolated per tenant, and automatically mounted to the right compute nodes. With Bridge GPU Cloud Management Software (GPU CMS), this entire process is streamlined through seamless integration with DDN EXAScaler.",
            "date_modified": "2025-05-08T00:00:00.000Z",
            "author": {
                "name": "Raghuram Gopalshetty"
            },
            "tags": [
                "storage",
                "ddn",
                "integration",
                "ai-workloads"
            ]
        },
        {
            "id": "https://docs.armada.ai/bridge/blog/infiniband-network-isolation",
            "content_html": "<p>Managing network isolation in AI cloud environments is critical for ensuring tenant data security, performance consistency, and compliance. This becomes even more important in high-performance AI clusters that rely on InfiniBand fabric for ultra-low latency communication between GPU nodes.</p>\n<p>With Bridge GPU Cloud Management Software (GPU CMS), cloud providers can achieve complete InfiniBand network isolation for every tenant—all through an automated, policy-driven process. This ensures each tenant's data and traffic are fully segregated, with no manual intervention required.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"fully-automated-infiniband-isolation\">Fully Automated InfiniBand Isolation<a href=\"https://docs.armada.ai/bridge/blog/infiniband-network-isolation#fully-automated-infiniband-isolation\" class=\"hash-link\" aria-label=\"Direct link to Fully Automated InfiniBand Isolation\" title=\"Direct link to Fully Automated InfiniBand Isolation\" translate=\"no\">​</a></h2>\n<p>Bridge GPU CMS achieves end-to-end isolation on InfiniBand fabrics by integrating seamlessly with NVIDIA UFM (Unified Fabric Manager). This allows for:</p>\n<ul>\n<li class=\"\"><strong>Automated Discovery</strong> – The system automatically detects and maps the full InfiniBand topology, including leaf, spine, and core switches, as well as all GPU nodes and their InfiniBand ports (GUIDs)</li>\n<li class=\"\"><strong>Tenant Creation &amp; PKey Assignment</strong> – When a new tenant is created, a unique PKey (Partition Key) is automatically provisioned for that tenant, establishing logical separation at the network layer</li>\n<li class=\"\"><strong>Resource Allocation &amp; GUID Mapping</strong> – When GPU compute nodes are allocated to a tenant, Bridge GPU CMS automatically maps all InfiniBand GUIDs of the tenant's servers to the tenant's PKey—ensuring that all traffic from those servers is restricted to the tenant's own isolated network partition</li>\n</ul>\n<p>This policy-based automation eliminates manual errors, guarantees secure isolation across the entire InfiniBand fabric, and ensures each tenant receives a fully segregated high-performance network.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"seamless-visibility-and-control\">Seamless Visibility and Control<a href=\"https://docs.armada.ai/bridge/blog/infiniband-network-isolation#seamless-visibility-and-control\" class=\"hash-link\" aria-label=\"Direct link to Seamless Visibility and Control\" title=\"Direct link to Seamless Visibility and Control\" translate=\"no\">​</a></h2>\n<p>All discovery, tenant creation, and isolation enforcement actions are fully visible within Bridge GPU CMS Admin Portal. Both NCP admins (cloud provider admins) and tenant admins can track:</p>\n<ul>\n<li class=\"\"><strong>Topology Discovery Results</strong> – Real-time visualization of the InfiniBand fabric, showing all switches, nodes, and links</li>\n<li class=\"\"><strong>Tenant-Specific Isolation</strong> – Full visibility into which PKey is assigned to each tenant</li>\n<li class=\"\"><strong>Server-Level Validation</strong> – Ability to drill down into individual servers and confirm that their InfiniBand ports are correctly assigned to the tenant's PKey</li>\n</ul>\n<p>This centralized visibility ensures operational transparency and gives cloud providers the tools they need to enforce multi-tenant isolation at scale.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"key-benefits-of-infiniband-integration\">Key Benefits of InfiniBand Integration<a href=\"https://docs.armada.ai/bridge/blog/infiniband-network-isolation#key-benefits-of-infiniband-integration\" class=\"hash-link\" aria-label=\"Direct link to Key Benefits of InfiniBand Integration\" title=\"Direct link to Key Benefits of InfiniBand Integration\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><strong>Automated Discovery &amp; Configuration</strong> – No manual effort required for topology discovery or PKey creation</li>\n<li class=\"\"><strong>Guaranteed Tenant Isolation</strong> – Each tenant's traffic is strictly confined to their assigned PKey</li>\n<li class=\"\"><strong>Centralized Management</strong> – All network isolation policies are managed from a single pane of glass within Bridge GPU CMS Admin Portal</li>\n<li class=\"\"><strong>NVIDIA UFM Integration</strong> – Leverages NVIDIA's industry-standard InfiniBand management platform for seamless compatibility</li>\n<li class=\"\"><strong>Real-Time Validation</strong> – Admins can instantly verify that each tenant's compute nodes are correctly isolated at the network level</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"complete-network-isolation-across-ethernet--infiniband\">Complete Network Isolation Across Ethernet &amp; InfiniBand<a href=\"https://docs.armada.ai/bridge/blog/infiniband-network-isolation#complete-network-isolation-across-ethernet--infiniband\" class=\"hash-link\" aria-label=\"Direct link to Complete Network Isolation Across Ethernet &amp; InfiniBand\" title=\"Direct link to Complete Network Isolation Across Ethernet &amp; InfiniBand\" translate=\"no\">​</a></h2>\n<p>While this blog focuses on InfiniBand isolation, Bridge GPU CMS also supports Ethernet network isolation, including full integration with NVIDIA Spectrum-X switches. Whether using Ethernet, InfiniBand, or a combination of both, Bridge GPU CMS ensures complete network separation between tenants—across both the control plane and data plane.</p>",
            "url": "https://docs.armada.ai/bridge/blog/infiniband-network-isolation",
            "title": "Automated InfiniBand Network Isolation with Bridge GPU CMS",
            "summary": "Managing network isolation in AI cloud environments is critical for ensuring tenant data security, performance consistency, and compliance. This becomes even more important in high-performance AI clusters that rely on InfiniBand fabric for ultra-low latency communication between GPU nodes.",
            "date_modified": "2025-03-07T00:00:00.000Z",
            "author": {
                "name": "Raghuram Gopalshetty"
            },
            "tags": [
                "networking",
                "infiniband",
                "isolation",
                "multi-tenancy"
            ]
        },
        {
            "id": "https://docs.armada.ai/bridge/blog/vast-storage-integration",
            "content_html": "<p>Managing external storage for GPU-accelerated AI workloads can be complex—especially when ensuring that storage volumes are provisioned correctly, isolated per tenant, and automatically mounted to the right compute nodes. With Bridge GPU Cloud Management Software (GPU CMS), this entire process is streamlined through seamless integration with VAST external storage systems.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"end-to-end-automation-with-no-manual-steps\">End-to-End Automation with No Manual Steps<a href=\"https://docs.armada.ai/bridge/blog/vast-storage-integration#end-to-end-automation-with-no-manual-steps\" class=\"hash-link\" aria-label=\"Direct link to End-to-End Automation with No Manual Steps\" title=\"Direct link to End-to-End Automation with No Manual Steps\" translate=\"no\">​</a></h2>\n<p>With Bridge GPU CMS, end users don't need to manually log into multiple systems, configure storage mounts, or worry about compatibility between compute and storage. The VAST integration is fully automated—allowing users to simply specify:</p>\n<ul>\n<li class=\"\">The desired storage size</li>\n<li class=\"\">The bare metal node where the storage should be mounted</li>\n</ul>\n<p>Everything else—from tenant-aware provisioning to storage policy enforcement and automatic mount point creation—is handled seamlessly by Bridge GPU CMS in the background.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"simple-and-efficient-flow\">Simple and Efficient Flow<a href=\"https://docs.armada.ai/bridge/blog/vast-storage-integration#simple-and-efficient-flow\" class=\"hash-link\" aria-label=\"Direct link to Simple and Efficient Flow\" title=\"Direct link to Simple and Efficient Flow\" translate=\"no\">​</a></h2>\n<p>The process starts with the NCP admin (cloud provider admin) importing the compute node into the system and setting up a new tenant. Once the tenant is onboarded, the tenant user can allocate a GPU bare-metal instance and request external storage from VAST.</p>\n<p>The tenant simply provides:</p>\n<ul>\n<li class=\"\">The desired storage size</li>\n<li class=\"\">The specific compute node where the storage should be mounted</li>\n</ul>\n<p>Once these inputs are provided, Bridge GPU CMS handles all interactions with VAST, including:</p>\n<ul>\n<li class=\"\">Configuring storage volumes</li>\n<li class=\"\">Assigning tenant-specific quotas</li>\n<li class=\"\">Creating the mount point</li>\n<li class=\"\">Ensuring the mount point is immediately available on the compute node</li>\n</ul>\n<p>This zero-touch integration eliminates any need for the tenant to interact with the VAST portal directly.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"real-time-validation-across-systems\">Real-Time Validation Across Systems<a href=\"https://docs.armada.ai/bridge/blog/vast-storage-integration#real-time-validation-across-systems\" class=\"hash-link\" aria-label=\"Direct link to Real-Time Validation Across Systems\" title=\"Direct link to Real-Time Validation Across Systems\" translate=\"no\">​</a></h2>\n<p>To ensure transparency and operational assurance, the NCP admin or tenant admin can view all configured storage volumes directly within Bridge GPU CMS. For additional verification, they can also cross-check the automatically created tenants, networks, policies, and mount points directly in the VAST admin portal.</p>\n<p>This two-way visibility ensures that:</p>\n<ul>\n<li class=\"\">The tenant's allocated storage matches the requested size</li>\n<li class=\"\">The network isolation policies (north-south overlays) are correctly applied</li>\n<li class=\"\">All configurations are performed via APIs with no manual intervention</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"full-tenant-experience\">Full Tenant Experience<a href=\"https://docs.armada.ai/bridge/blog/vast-storage-integration#full-tenant-experience\" class=\"hash-link\" aria-label=\"Direct link to Full Tenant Experience\" title=\"Direct link to Full Tenant Experience\" translate=\"no\">​</a></h2>\n<p>Once the storage is provisioned, the tenant user can log directly into their allocated GPU compute node and immediately access the mounted VAST storage volume. Whether for large-scale AI training data or model checkpoints, this automated mount ensures data is available where and when the user needs it.</p>\n<p>To further validate, the tenant can create and save files to the external storage—confirming that the VAST integration is complete and the storage is fully accessible from their compute instance.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"key-benefits\">Key Benefits<a href=\"https://docs.armada.ai/bridge/blog/vast-storage-integration#key-benefits\" class=\"hash-link\" aria-label=\"Direct link to Key Benefits\" title=\"Direct link to Key Benefits\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><strong>End-to-End Automation</strong>: No manual steps—just specify size and compute node, and Bridge GPU CMS handles everything else</li>\n<li class=\"\"><strong>Single Pane of Glass</strong>: Both compute and storage provisioning are managed from a single interface</li>\n<li class=\"\"><strong>Full Tenant Isolation</strong>: Each tenant's storage is isolated with tenant-specific quotas and network policies</li>\n<li class=\"\"><strong>Real-Time Observability</strong>: Both admins and tenants can view and validate storage allocations directly within Bridge GPU CMS portal</li>\n<li class=\"\"><strong>API-Driven Consistency</strong>: All configurations—from mount points to network overlays—are performed through automated APIs, ensuring accuracy and compliance with tenant policies</li>\n</ul>",
            "url": "https://docs.armada.ai/bridge/blog/vast-storage-integration",
            "title": "Seamless External Storage Integration with VAST Using Bridge GPU CMS",
            "summary": "Managing external storage for GPU-accelerated AI workloads can be complex—especially when ensuring that storage volumes are provisioned correctly, isolated per tenant, and automatically mounted to the right compute nodes. With Bridge GPU Cloud Management Software (GPU CMS), this entire process is streamlined through seamless integration with VAST external storage systems.",
            "date_modified": "2025-03-06T00:00:00.000Z",
            "author": {
                "name": "Raghuram Gopalshetty"
            },
            "tags": [
                "storage",
                "vast",
                "integration",
                "ai-workloads"
            ]
        },
        {
            "id": "https://docs.armada.ai/bridge/blog/spectrum-x-validation",
            "content_html": "<p>The latest Bridge GPU CMS announces network automation, observability, fault management, and multi-tenancy for the v1.3 NVIDIA Spectrum-X Reference Architecture (RA). The Reference Architecture defines an East-West compute network fabric optimized for AI cloud deployments with HGX systems and a North-South converged network for external access, storage, and control plane traffic.</p>\n<p>As part of this announcement, Bridge supports NVIDIA Spectrum-4 SN5000 Series Ethernet switches, NVIDIA Cumulus Linux, NVIDIA BlueField-3 SuperNICs and DPUs, NVIDIA NetQ AI observability and telemetry platform, and NVIDIA Air data center digital twin platform along with NVIDIA HGX H100/H200 nodes.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"multi-tenancy-across-seven-pillars\">Multi-Tenancy Across Seven Pillars<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#multi-tenancy-across-seven-pillars\" class=\"hash-link\" aria-label=\"Direct link to Multi-Tenancy Across Seven Pillars\" title=\"Direct link to Multi-Tenancy Across Seven Pillars\" translate=\"no\">​</a></h2>\n<p>NVIDIA Cloud Partner (NCP) and enterprise AI clouds must provide hard isolation between tenants. Bridge GPU CMS addresses this comprehensively across seven pillars:</p>\n<ol>\n<li class=\"\">High-performance networking</li>\n<li class=\"\">InfiniBand fabrics</li>\n<li class=\"\">NVLink GPU interconnects</li>\n<li class=\"\">Scalable storage</li>\n<li class=\"\">Virtual private clouds (VPCs)</li>\n<li class=\"\">Compute resources</li>\n<li class=\"\">GPUs</li>\n</ol>\n<p>By enforcing isolation and performance in each pillar, we ensure tenants can run demanding AI workloads securely and without compromising the performance on shared hardware.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"switch-fabric-automation\">Switch Fabric Automation<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#switch-fabric-automation\" class=\"hash-link\" aria-label=\"Direct link to Switch Fabric Automation\" title=\"Direct link to Switch Fabric Automation\" translate=\"no\">​</a></h2>\n<p>A modern GPU data center switch fabric typically comprises multiple segmented networks—most notably the East-West (Compute) network and the North-South (Converged) network. The North-South network itself includes Inband management, Storage, and External/Tenant access.</p>\n<p>A single Scalable Unit (SU)—as defined by NVIDIA Spectrum-X Reference Architectures—contains:</p>\n<ul>\n<li class=\"\">32 GPU nodes</li>\n<li class=\"\">12 switches</li>\n<li class=\"\">256 physical cable connections</li>\n</ul>\n<p>This topology serves just 256 GPUs, highlighting the operational complexity and scale. Without automation, managing such fabrics—especially across multiple SUs—becomes impractical and error-prone.</p>\n<p>Bridge CMS addresses this challenge by offering:</p>\n<ul>\n<li class=\"\"><strong>Topology auto-discovery</strong></li>\n<li class=\"\"><strong>Underlay configuration</strong></li>\n<li class=\"\"><strong>Lifecycle automation for switch configurations</strong></li>\n<li class=\"\"><strong>Overlay network creation for tenant-level isolation</strong></li>\n<li class=\"\"><strong>Integration with compute orchestration pipelines</strong></li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"bluefield-3-supernic-spectrum-x-configuration\">BlueField-3 SuperNIC Spectrum-X Configuration<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#bluefield-3-supernic-spectrum-x-configuration\" class=\"hash-link\" aria-label=\"Direct link to BlueField-3 SuperNIC Spectrum-X Configuration\" title=\"Direct link to BlueField-3 SuperNIC Spectrum-X Configuration\" translate=\"no\">​</a></h2>\n<p>Bridge GPU CMS automates all the tasks needed to turn a BlueField-3 equipped server into a Spectrum-X host. When a new host is provisioned, the GPU CMS:</p>\n<ol>\n<li class=\"\">Installs the DOCA Host Packages, enabling key services and libraries</li>\n<li class=\"\">Brings up the DOCA Management Service Daemon (DMSD)</li>\n<li class=\"\">Configures the BlueField-3 SuperNIC with Spectrum-X capabilities like RoCE, congestion control, adaptive routing, and IP routing</li>\n</ol>\n<p>This makes the host fully Spectrum-X aware — ready to participate in high-performance GPU-to-GPU networking, where the multi-tenancy relies on BGP EVPN from Cumulus and Spectrum-X switches.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"network-multi-tenancy\">Network Multi-Tenancy<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#network-multi-tenancy\" class=\"hash-link\" aria-label=\"Direct link to Network Multi-Tenancy\" title=\"Direct link to Network Multi-Tenancy\" translate=\"no\">​</a></h2>\n<p>Bridge GPU CMS creates tenant isolated overlay networks on the ethernet switch fabric using VxLAN and VRF with BGP as the control plane. From the end-user perspective, the GPU CMS provides the ability to define Virtual Private Clouds (VPCs) with multiple subnets, much like a traditional public cloud.</p>\n<p>Behind the scenes, each VPC is backed by an isolated VRF (Virtual Routing and Forwarding) and each subnet corresponds to a unique VxLAN segment.</p>\n<p>This abstraction gives users the flexibility to:</p>\n<ul>\n<li class=\"\">Create isolated environments for different workloads or projects</li>\n<li class=\"\">Define fine-grained IP subnetting and routing policies</li>\n<li class=\"\">Attach load balancers or gateways at subnet edges</li>\n<li class=\"\">Connect VPCs to storage networks or on-prem environments</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"vxlan-and-vrf-based-tenant-segmentation\">VxLAN and VRF-Based Tenant Segmentation<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#vxlan-and-vrf-based-tenant-segmentation\" class=\"hash-link\" aria-label=\"Direct link to VxLAN and VRF-Based Tenant Segmentation\" title=\"Direct link to VxLAN and VRF-Based Tenant Segmentation\" translate=\"no\">​</a></h2>\n<p>To support multiple tenants securely and efficiently on the same physical fabric, Bridge GPU CMS uses VxLAN (Virtual Extensible LAN) in combination with VRF (Virtual Routing and Forwarding) constructs. This ensures complete L2/L3 network isolation across tenants.</p>\n<p>Each tenant's workloads operate within a dedicated overlay network, backed by a separate VRF instance, enabling traffic segmentation, independent routing policies, and security boundaries, without sacrificing performance.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"observability-and-fault-management\">Observability and Fault Management<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#observability-and-fault-management\" class=\"hash-link\" aria-label=\"Direct link to Observability and Fault Management\" title=\"Direct link to Observability and Fault Management\" translate=\"no\">​</a></h2>\n<p>Bridge GPU CMS supports both NetQ and OTLP based telemetry for NVIDIA Spectrum-X switches.</p>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"unified-observability-with-otlp\">Unified Observability with OTLP<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#unified-observability-with-otlp\" class=\"hash-link\" aria-label=\"Direct link to Unified Observability with OTLP\" title=\"Direct link to Unified Observability with OTLP\" translate=\"no\">​</a></h3>\n<p><strong>Switch Telemetry via Cumulus NVUE:</strong></p>\n<ul>\n<li class=\"\">NVIDIA Cumulus Linux on Spectrum-X switches exports telemetry via NVUE commands</li>\n<li class=\"\">Metrics exposed: Buffer occupancy histograms, interface-level stats (bandwidth, errors, drops), platform metrics (temperature, fan speed, power)</li>\n<li class=\"\">Telemetry is exported in OTEL format</li>\n</ul>\n<p><strong>SuperNIC Telemetry via DOCA DTS:</strong></p>\n<ul>\n<li class=\"\">DOCA Telemetry Service (DTS) collects real-time SuperNIC metrics</li>\n<li class=\"\">Supports telemetry types: High-Frequency Telemetry (HFT), Programmable Congestion Control (PCC)</li>\n</ul>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"netq-integration\">NetQ Integration<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#netq-integration\" class=\"hash-link\" aria-label=\"Direct link to NetQ Integration\" title=\"Direct link to NetQ Integration\" translate=\"no\">​</a></h3>\n<p>Infrastructure administrators can subscribe to NetQ events directly from the GPU CMS. When subscribed events are triggered, Bridge GPU CMS invokes its own remediation logic to auto-correct faults without manual intervention, helping customers with dramatic OPEX reduction.</p>\n<p>The CMS subscribes to the following fault scenarios from NetQ:</p>\n<ul>\n<li class=\"\"><strong>Switch Failure Detection</strong>: Power loss or hardware faults on switches are detected promptly</li>\n<li class=\"\"><strong>Link Failure Detection</strong>: Real-time monitoring of physical interfaces allows rapid identification of link failures</li>\n<li class=\"\"><strong>Configuration Drift Detection</strong>: Configuration changes are continuously audited to detect and flag drifts</li>\n<li class=\"\"><strong>BGP Session State Monitoring</strong>: Changes in BGP session status are actively monitored</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"nvidia-air-integration\">NVIDIA Air Integration<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#nvidia-air-integration\" class=\"hash-link\" aria-label=\"Direct link to NVIDIA Air Integration\" title=\"Direct link to NVIDIA Air Integration\" translate=\"no\">​</a></h2>\n<p>Bridge GPU CMS integration with Spectrum-X is extensively validated on NVIDIA Air (data center digital twin), enabling realistic simulation and demonstration of features in a virtual environment.</p>\n<ul>\n<li class=\"\"><strong>Storage Integration</strong>: Supports vanilla NFS servers and is integrated with major storage vendors such as DDN and VAST</li>\n<li class=\"\"><strong>External Gateway Integration</strong>: Border leaf configurations are tested with F5 BIG-IP gateway</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"global-customer-deployments\">Global Customer Deployments<a href=\"https://docs.armada.ai/bridge/blog/spectrum-x-validation#global-customer-deployments\" class=\"hash-link\" aria-label=\"Direct link to Global Customer Deployments\" title=\"Direct link to Global Customer Deployments\" translate=\"no\">​</a></h2>\n<p>Our platform is already being used by NCPs in multiple regions along with Spectrum-X. For example, a leading Southeast Asian telecom operator is deploying an AI compute grid for distributed inference. Similarly, a top U.S. NCP is building an AI-and-RAN edge platform with multi-tenancy support for multiple use cases or customers.</p>",
            "url": "https://docs.armada.ai/bridge/blog/spectrum-x-validation",
            "title": "Bridge GPU CMS Announces Network Automation and Multi-Tenancy for NVIDIA Spectrum-X",
            "summary": "The latest Bridge GPU CMS announces network automation, observability, fault management, and multi-tenancy for the v1.3 NVIDIA Spectrum-X Reference Architecture (RA). The Reference Architecture defines an East-West compute network fabric optimized for AI cloud deployments with HGX systems and a North-South converged network for external access, storage, and control plane traffic.",
            "date_modified": "2025-01-15T00:00:00.000Z",
            "author": {
                "name": "Sriram Rupanagunta"
            },
            "tags": [
                "networking",
                "spectrum-x",
                "nvidia",
                "multi-tenancy",
                "automation"
            ]
        },
        {
            "id": "https://docs.armada.ai/bridge/blog/nvidia-bluefield-rtx-pro-integration",
            "content_html": "<p>As organizations continue to build AI factories capable of handling massive-scale inference and data processing, one challenge looms large: how to deliver secure, multi-tenant infrastructure that keeps GPUs fully utilized without adding operational complexity.</p>\n<p>At Armada, we're solving that challenge head-on. Bridge product now integrates with NVIDIA BlueField-3 data processing units (DPUs) and NVIDIA RTX PRO Servers featuring NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs — creating an end-to-end foundation for high-performance, automated AI infrastructure.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"offloading-complexity-unlocking-performance\">Offloading Complexity, Unlocking Performance<a href=\"https://docs.armada.ai/bridge/blog/nvidia-bluefield-rtx-pro-integration#offloading-complexity-unlocking-performance\" class=\"hash-link\" aria-label=\"Direct link to Offloading Complexity, Unlocking Performance\" title=\"Direct link to Offloading Complexity, Unlocking Performance\" translate=\"no\">​</a></h2>\n<p>GPU infrastructure deployed manually could result in inefficiency, underutilization, and data exposure. With Bridge, enterprises can automate the entire lifecycle of GPU infrastructure while enabling secure multi-tenancy across compute, storage, and networking layers.</p>\n<p>By leveraging NVIDIA BlueField DPUs, network, storage and security workloads for AI infrastructure are offloaded from host CPUs, freeing up valuable compute cycles for business applications. Further, BlueField-accelerated data movement and zero trust security result in faster inference, stronger isolation, and more predictable performance for every AI workload.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"building-the-fully-automated-ai-factory\">Building the Fully Automated AI Factory<a href=\"https://docs.armada.ai/bridge/blog/nvidia-bluefield-rtx-pro-integration#building-the-fully-automated-ai-factory\" class=\"hash-link\" aria-label=\"Direct link to Building the Fully Automated AI Factory\" title=\"Direct link to Building the Fully Automated AI Factory\" translate=\"no\">​</a></h2>\n<p>Bridge integrates with NVIDIA RTX PRO Servers and NVIDIA BlueField DPUs to power fully automated AI Factories.</p>\n<p>Key capabilities include:</p>\n<ul>\n<li class=\"\"><strong>Secure multi-tenancy</strong> across CPU, GPU, NVIDIA Spectrum-X Ethernet, storage, and external connectivity</li>\n<li class=\"\"><strong>IaaS capabilities</strong>: Bare-Metal-as-a-Service, VM-as-a-Service, Storage-as-a-Service, and dedicated Kubernetes clusters</li>\n<li class=\"\"><strong>PaaS capabilities</strong>: LLM-as-a-Service, job submission, a third-party app catalog, and integration with NVIDIA Cloud Functions (NVCF) to monetize idle GPU capacity</li>\n<li class=\"\"><strong>Automated orchestration</strong> of NVIDIA DOCA microservices on BlueField for network, storage, and security acceleration leveraging the DOCA Platform Framework (DPF)</li>\n</ul>\n<p>Together, these features enable enterprises to deploy secure, high-performance AI environments faster — maximizing GPU efficiency while reducing operational overhead.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"upcoming-features\">Upcoming Features<a href=\"https://docs.armada.ai/bridge/blog/nvidia-bluefield-rtx-pro-integration#upcoming-features\" class=\"hash-link\" aria-label=\"Direct link to Upcoming Features\" title=\"Direct link to Upcoming Features\" translate=\"no\">​</a></h2>\n<p>Armada continues to expand Bridge to leverage additional capabilities of the BlueField. Upcoming features include:</p>\n<ul>\n<li class=\"\">Zero-trust networking and distributed firewalls</li>\n<li class=\"\">DOCA-based storage applications</li>\n<li class=\"\">Kubernetes control plane on DPU</li>\n<li class=\"\">Secure boot and AI runtime protection</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"looking-ahead-bluefield-4-and-gigascale-ai\">Looking Ahead: BlueField-4 and Gigascale AI<a href=\"https://docs.armada.ai/bridge/blog/nvidia-bluefield-rtx-pro-integration#looking-ahead-bluefield-4-and-gigascale-ai\" class=\"hash-link\" aria-label=\"Direct link to Looking Ahead: BlueField-4 and Gigascale AI\" title=\"Direct link to Looking Ahead: BlueField-4 and Gigascale AI\" translate=\"no\">​</a></h2>\n<p>The journey doesn't stop here. With NVIDIA BlueField-4 on the horizon — offering 6X the compute power of NVIDIA BlueField-3 and 800 Gb/s throughput — Bridge will extend its capabilities even further.</p>\n<p>These innovations will enable AI factories to handle gigascale workloads while maintaining tight isolation, blazing-fast data access, and best-in-class security.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-future-of-ai-infrastructure\">The Future of AI Infrastructure<a href=\"https://docs.armada.ai/bridge/blog/nvidia-bluefield-rtx-pro-integration#the-future-of-ai-infrastructure\" class=\"hash-link\" aria-label=\"Direct link to The Future of AI Infrastructure\" title=\"Direct link to The Future of AI Infrastructure\" translate=\"no\">​</a></h2>\n<p>As enterprises evolve from AI pilots to full-scale production, they need infrastructure that's not only powerful — but programmable, secure, and efficient. The integration of Bridge with NVIDIA RTX PRO Servers and NVIDIA BlueField DPUs represents a major step toward that future.</p>\n<p>We're excited to help our customers build the world's most efficient and secure AI factories — one GPU at a time.</p>",
            "url": "https://docs.armada.ai/bridge/blog/nvidia-bluefield-rtx-pro-integration",
            "title": "Armada Powering the Next Generation of Secure, Multi-Tenant AI Factories",
            "summary": "As organizations continue to build AI factories capable of handling massive-scale inference and data processing, one challenge looms large: how to deliver secure, multi-tenant infrastructure that keeps GPUs fully utilized without adding operational complexity.",
            "date_modified": "2024-12-15T00:00:00.000Z",
            "author": {
                "name": "Amar Kapadia"
            },
            "tags": [
                "nvidia",
                "bluefield",
                "dpu",
                "rtx-pro",
                "multi-tenancy",
                "ai-factory",
                "automation"
            ]
        },
        {
            "id": "https://docs.armada.ai/bridge/blog/soft-isolation-risks",
            "content_html": "<p>Relying solely on Kubernetes Namespaces or vClusters for multi-tenant isolation in GPU clouds is risky — especially when hosting untrusted or external workloads.</p>\n<p>In September 2024, Wiz discovered a critical NVIDIA Container Toolkit vulnerability (CVE-2024-0132) that allowed GPU containers to escape soft isolation and gain root access to the host. This flaw impacted over one-third of GPU-enabled environments and exposed the limits of Kubernetes-based isolation.</p>\n<p><strong>Soft isolation is not secure isolation.</strong> For environments like Neoclouds, NVIDIA Cloud Partners (NCPs), or regulated industries, only hard or hybrid isolation strategies — such as dedicated Kubernetes clusters, MIG-based GPU partitioning, VPCs, VxLAN, VRFs, KVM virtualization, IB P-KEY, and NVLink partitioning — can protect against container escapes.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"context\">Context<a href=\"https://docs.armada.ai/bridge/blog/soft-isolation-risks#context\" class=\"hash-link\" aria-label=\"Direct link to Context\" title=\"Direct link to Context\" translate=\"no\">​</a></h2>\n<p>When building a multi-tenant GPU cloud, it's common to rely on Kubernetes-native tools like Namespaces or vClusters for tenant isolation. This \"soft isolation\" approach offers ease of deployment and simplicity — but it carries serious security risks in GPU-accelerated environments.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"soft-isolation-is-not-security-isolation\">Soft Isolation is Not Security Isolation<a href=\"https://docs.armada.ai/bridge/blog/soft-isolation-risks#soft-isolation-is-not-security-isolation\" class=\"hash-link\" aria-label=\"Direct link to Soft Isolation is Not Security Isolation\" title=\"Direct link to Soft Isolation is Not Security Isolation\" translate=\"no\">​</a></h2>\n<p>Soft isolation assumes that tenants can safely share the same Kubernetes control plane, node kernel, and container runtime. But when it comes to multi-tenant GPU workloads, this assumption breaks down fast — and can open the door to container escapes, privilege escalation, and cross-tenant compromise.</p>\n<p>This isn't just theoretical.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"real-world-example-the-nvidia-gpu-container-vulnerability\">Real-World Example: The NVIDIA GPU Container Vulnerability<a href=\"https://docs.armada.ai/bridge/blog/soft-isolation-risks#real-world-example-the-nvidia-gpu-container-vulnerability\" class=\"hash-link\" aria-label=\"Direct link to Real-World Example: The NVIDIA GPU Container Vulnerability\" title=\"Direct link to Real-World Example: The NVIDIA GPU Container Vulnerability\" translate=\"no\">​</a></h2>\n<p>In September 2024, security researchers at Wiz discovered a critical vulnerability — CVE-2024-0132 — in the NVIDIA Container Toolkit (v1.16.1 and earlier), a core component used to enable GPU access for Docker, containerd, and CRI-O containers.</p>\n<p>Here's what went wrong:</p>\n<ul>\n<li class=\"\">The bug was a <strong>time-of-check/time-of-use (TOCTOU) race condition</strong></li>\n<li class=\"\">It allowed a malicious GPU container to <strong>mount the host filesystem</strong> using NVIDIA's container runtime hooks</li>\n<li class=\"\">Once the container had access to the host filesystem, it could interact with the container runtime socket (e.g., containerd.sock or docker.sock) and <strong>gain root-level access to the host</strong></li>\n<li class=\"\">This effectively allowed a tenant to <strong>escape their container sandbox</strong> and compromise the entire node</li>\n</ul>\n<p>The vulnerability was widespread. Wiz estimated that <strong>over one-third of cloud GPU environments</strong> — across AWS, Azure, GCP, and on-prem systems — were running the vulnerable software. Kubernetes clusters using the NVIDIA GPU Operator v24.6.1 and earlier were especially affected.</p>\n<p>Critically, <strong>this vulnerability bypassed all protections offered by Kubernetes Namespaces or vClusters</strong>. Even if workloads were logically isolated at the K8s level, any tenant with the ability to run a malicious GPU container could break out and attack the host.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-this-means-for-gpu-multi-tenancy\">What This Means for GPU Multi-Tenancy<a href=\"https://docs.armada.ai/bridge/blog/soft-isolation-risks#what-this-means-for-gpu-multi-tenancy\" class=\"hash-link\" aria-label=\"Direct link to What This Means for GPU Multi-Tenancy\" title=\"Direct link to What This Means for GPU Multi-Tenancy\" translate=\"no\">​</a></h2>\n<p>This incident underscores a hard truth: <strong>soft isolation is not sufficient</strong> in environments where untrusted or external tenants run GPU workloads. Kubernetes Namespaces and vClusters only provide logical segmentation — they don't protect against kernel-level or runtime-level exploits.</p>\n<p>For Neoclouds, NVIDIA Cloud Partners (NCPs), or regulated enterprises, soft isolation leaves the door wide open. A single misconfigured container or a compromised customer workload could escalate into a full infrastructure breach.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-you-should-be-doing-instead\">What You Should Be Doing Instead<a href=\"https://docs.armada.ai/bridge/blog/soft-isolation-risks#what-you-should-be-doing-instead\" class=\"hash-link\" aria-label=\"Direct link to What You Should Be Doing Instead\" title=\"Direct link to What You Should Be Doing Instead\" translate=\"no\">​</a></h2>\n<p>To protect against these risks, multi-tenant GPU clouds must adopt <strong>hard or hybrid isolation</strong>, including:</p>\n<ul>\n<li class=\"\"><strong>Dedicated Kubernetes clusters or node pools</strong> per tenant</li>\n<li class=\"\"><strong>CPU partitioning</strong> via KVM (virtualization) or bare-metal allocation</li>\n<li class=\"\"><strong>GPU resource partitioning</strong> via MIG or bare-metal allocation</li>\n<li class=\"\"><strong>Spectrum-X or other Ethernet network segmentation</strong> using VPCs or VxLANs</li>\n<li class=\"\"><strong>InfiniBand isolation</strong> using P-KEYS</li>\n<li class=\"\"><strong>NVLink partitioning</strong></li>\n<li class=\"\"><strong>Storage-level isolation</strong> using VRFs and tenant-specific volumes</li>\n<li class=\"\"><strong>Runtime controls</strong> that avoid shared container runtimes for untrusted workloads</li>\n<li class=\"\">Use of <strong>CDI (Container Device Interface)</strong> mode instead of NVIDIA's older hooks (CDI is not affected by CVE-2024-0132)</li>\n</ul>\n<p>Additionally, any GPU cloud environment should now be running the latest version of NVIDIA Container Toolkit and GPU Operator.</p>\n<p>While deploying a dedicated, multi-node Kubernetes cluster is the gold standard for large production tenants, this model is not always cost-effective for smaller teams, university labs, or development workloads. To address this, lightweight, dedicated Kubernetes clusters using efficient, fully compliant distributions like K3s or MicroK8s could also be considered. These clusters are deployed within a single virtual machine, providing a complete, self-contained environment with its own dedicated kernel and control plane, while physical GPUs are attached via passthrough.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"bottom-line\">Bottom Line<a href=\"https://docs.armada.ai/bridge/blog/soft-isolation-risks#bottom-line\" class=\"hash-link\" aria-label=\"Direct link to Bottom Line\" title=\"Direct link to Bottom Line\" translate=\"no\">​</a></h2>\n<p><strong>Soft isolation is fine — until it isn't.</strong> In GPU-accelerated environments, especially those involving multi-tenancy, the risks are too high to depend solely on Kubernetes constructs like Namespaces or vClusters.</p>\n<p>Real-world vulnerabilities like CVE-2024-0132 show that container escapes are not just possible — they're happening. If you're serving customers, partners, or even internal departments with access to GPU workloads, you need real, infrastructure-level isolation.</p>\n<p>The good news? With tools like Bridge GPU CMS, you can build secure, tenant-isolated GPU clouds that automatically configure isolation across compute, storage, networking, VPC, NVLink, InfiniBand, and GPUs — without taking on all the complexity yourself.</p>",
            "url": "https://docs.armada.ai/bridge/blog/soft-isolation-risks",
            "title": "The Hidden Risks of Soft Isolation in Multi-Tenant GPU Clouds",
            "summary": "Relying solely on Kubernetes Namespaces or vClusters for multi-tenant isolation in GPU clouds is risky — especially when hosting untrusted or external workloads.",
            "date_modified": "2024-10-15T00:00:00.000Z",
            "author": {
                "name": "Amar Kapadia"
            },
            "tags": [
                "security",
                "isolation",
                "multi-tenancy",
                "kubernetes"
            ]
        }
    ]
}