{"id":16352,"date":"2026-06-23T22:02:06","date_gmt":"2026-06-23T16:02:06","guid":{"rendered":"https:\/\/dtasiagroup.com\/?p=16352"},"modified":"2026-06-23T22:02:08","modified_gmt":"2026-06-23T16:02:08","slug":"monitoring-ai-infrastructure-network-visibility-for-security-and-operations-teams","status":"publish","type":"post","link":"https:\/\/dtasiagroup.com\/vi\/monitoring-ai-infrastructure-network-visibility-for-security-and-operations-teams\/","title":{"rendered":"Monitoring AI Infrastructure: Network Visibility for Security and Operations Teams"},"content":{"rendered":"<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"576\" src=\"https:\/\/dtasiagroup.com\/wp-content\/uploads\/2026\/06\/image-1024x576.png\" alt=\"\" class=\"wp-image-16353\"\/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Organizations deploying GPU infrastructure for AI training and inference are facing a new set of networking challenges. Unlike traditional enterprise applications, AI workloads generate unique traffic patterns, including large-scale data ingestion before training, high-bandwidth model checkpoint transfers to distributed storage, continuous inference requests from thousands of users, and ongoing management traffic that crosses cluster boundaries.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For network and security teams, the issue is not whether this traffic exists\u2014it travels across standard IP networks and generates NetFlow records on routers and switches at the cluster boundary. The real question is whether that traffic is being monitored, baselined, and analyzed for anomalies that could signal performance issues or security incidents.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This article explains where NetFlow Optimizer (NFO) provides measurable value for AI infrastructure monitoring and where its visibility naturally ends.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Understanding What NetFlow Can See in an AI Environment<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">AI cluster networks typically consist of two separate traffic planes. Knowing which one NetFlow can observe is critical for setting realistic expectations.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">The East-West Compute Fabric: Outside NetFlow Visibility<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">During model training, GPU-to-GPU communication occurs across a dedicated compute fabric. By 2026, RoCEv2 over Ethernet has become the dominant technology for enterprise AI clusters, while InfiniBand remains common in the largest hyperscale environments. Both technologies rely on RDMA.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Because RDMA bypasses the traditional IP networking stack, this traffic is not visible to NetFlow or sFlow collectors, regardless of whether the underlying transport uses Ethernet or InfiniBand.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This distinction is important: NFO cannot observe east-west GPU training traffic in most enterprise AI deployments. Visibility into the compute fabric requires specialized RDMA monitoring solutions.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">The North-South Front-End Network: Fully Visible to NetFlow<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Every AI cluster connects to the wider enterprise environment through a standard IP-based front-end network. This network carries:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Data ingestion from storage platforms and data lakes<\/li>\n\n\n\n<li>Model checkpoint transfers<\/li>\n\n\n\n<li>Inference traffic between users, applications, and AI services<\/li>\n\n\n\n<li>Cluster management and orchestration communications<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Because this traffic runs over standard TCP\/IP, it generates NetFlow records on boundary routers and switches.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">While NFO cannot see inside the GPU compute fabric, it provides complete visibility into traffic crossing the cluster boundary, including inbound and outbound data transfers, inference requests, management communications, and unauthorized external connections.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Three Areas Where NFO Delivers Clear Value<\/h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"583\" src=\"https:\/\/dtasiagroup.com\/wp-content\/uploads\/2026\/06\/image-1-1024x583.png\" alt=\"\" class=\"wp-image-16355\"\/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>1. Monitoring Data Ingestion and Storage Traffic<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">AI training workloads depend on moving large datasets from object storage, data lakes, and NFS repositories to GPU nodes before and during training operations. These transfers generate sustained, high-bandwidth flows that are fully visible through NetFlow records collected at the cluster boundary.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">NFO provides per-flow and per-application bandwidth visibility, allowing teams to understand:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Which storage systems are serving specific GPU nodes<\/li>\n\n\n\n<li>The volume of data being transferred<\/li>\n\n\n\n<li>When transfers occur<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This information is valuable for capacity planning and performance troubleshooting. Teams can determine whether storage network congestion is contributing to training slowdowns, identify heavily utilized storage tiers, and detect preprocessing pipelines generating unexpected traffic that competes with training workloads for bandwidth.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">2. Baselining Inference Traffic and Detecting Anomalies<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Inference workloads create a very different traffic profile from training workloads. Instead of large internal data exchanges, inference environments handle high volumes of concurrent requests from users and downstream applications accessing AI services.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Many organizations now operate hybrid AI architectures, using dedicated RDMA fabrics for training while serving inference traffic over standard Ethernet networks. The inference side remains fully visible through NetFlow telemetry.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">NFO enriches flow records with application intelligence, GeoIP information, and, where applicable, user identity data from Active Directory, Entra ID, Okta, and VPN authentication logs before forwarding the information to SIEM and monitoring platforms.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This enriched telemetry enables organizations to establish normal traffic baselines and identify anomalies such as:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Unexpected spikes in inference request volumes<\/li>\n\n\n\n<li>Connections originating from unauthorized source IP addresses<\/li>\n\n\n\n<li>Access attempts outside approved operational windows<\/li>\n\n\n\n<li>Requests from unusual geographic locations<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">For organizations operating AI services in regulated industries, this visibility also supports ongoing monitoring and audit requirements associated with applicable compliance frameworks.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">3. Strengthening AI Infrastructure Security Monitoring<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">AI infrastructure has become a high-value target for attackers. Proprietary model weights, training datasets, and expensive compute resources represent attractive assets for cybercriminals and nation-state actors alike.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">NFO delivers enriched flow telemetry from the cluster boundary, providing upstream security platforms with the information needed to detect several key threat scenarios.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Model and Dataset Exfiltration<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Large outbound transfers from storage systems or AI cluster nodes to unexpected external destinations are often the primary indicator of data theft.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">By analyzing flow duration and cumulative transfer volumes over extended periods, security teams can identify both immediate and low-and-slow exfiltration attempts. This approach aligns directly with the detection methodology discussed in <em>Defeating the Low and Slow<\/em>.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Unauthorized Access to Inference Services<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Connections originating from source IP addresses outside approved access lists can be identified immediately through NFO-generated flow telemetry.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">When enriched with GeoIP information and threat intelligence context, these events become significantly easier to investigate and prioritize.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Indicators of Supply Chain Compromise<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Unexpected outbound communications from AI infrastructure during or after model deployment may indicate compromise.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Examples include connections to unfamiliar package repositories, unknown internet destinations, or external hosts flagged by threat intelligence feeds. Because these communications traverse the north-south network boundary, they are visible within NetFlow data and can be investigated before they develop into larger incidents.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">NFO Visibility Summary for AI Infrastructure<\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><td><strong>Traffic Type<\/strong><\/td><td><strong>NFO Visibility<\/strong><\/td><td><strong>Value Delivered<\/strong><\/td><\/tr><\/thead><tbody><tr><td>Data ingestion from storage to GPU nodes<\/td><td><strong>Full<\/strong><\/td><td>Bandwidth monitoring, storage capacity planning, training bottleneck identification<\/td><\/tr><tr><td>Model checkpointing to distributed storage<\/td><td><strong>Full<\/strong><\/td><td>Checkpoint frequency and volume tracking, storage utilization visibility<\/td><\/tr><tr><td>Inference serving (users and applications to GPU nodes)<\/td><td><strong>Full<\/strong><\/td><td>Enriched data foundation for upstream baselining, anomaly detection, unauthorized access detection, and audit trail for regulated environments<\/td><\/tr><tr><td>Management and orchestration traffic<\/td><td><strong>Full<\/strong><\/td><td>Unexpected management connections, configuration change indicators, first-contact external destinations<\/td><\/tr><tr><td>Outbound connections from AI infrastructure (potential exfiltration)<\/td><td><strong>Full<\/strong><\/td><td>Enriched outbound flow data enabling exfiltration detection, supply chain compromise identification, and threat intelligence screening in upstream security systems<\/td><\/tr><tr><td>East-west GPU training traffic (RoCEv2 or InfiniBand compute fabric)<\/td><td>Not visible<\/td><td>RDMA bypasses the IP stack; dedicated RDMA monitoring tools required for compute fabric visibility<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h5 class=\"wp-block-heading\"><strong>Deploying NFO for AI Infrastructure Visibility<\/strong><\/h5>\n\n\n\n<p class=\"wp-block-paragraph\">NFO is software-only and deploys on standard Linux or Windows Server with no hardware changes to AI infrastructure. The deployment model is straightforward: NFO collects NetFlow or IPFIX from the boundary switches and routers connecting the AI cluster to the storage network and enterprise network, enriches the data, and delivers it to your SIEM or monitoring platform in under one hour.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For organizations delivering AI services in regulated environments, NFO\u2019s on-premises, air-gap-compatible architecture ensures that AI infrastructure telemetry stays inside the security boundary. See the\u00a0NFO Government Solution Brief\u00a0for deployment architecture details relevant to classified and sensitive environments.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h5 class=\"wp-block-heading\"><strong>T\u00f3m l\u1ea1i<\/strong><\/h5>\n\n\n\n<p class=\"wp-block-paragraph\">NFO does not provide visibility into the GPU compute fabric. That requires dedicated RDMA monitoring tools. What it provides is continuous, enriched network telemetry for everything that crosses the AI cluster boundary: the storage traffic that feeds training runs, the inference traffic that serves users, the management traffic that operates the cluster, and the outbound connections that could indicate a security incident.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For most enterprises deploying AI infrastructure, the cluster boundary is where the security and operational visibility gaps are largest and least addressed. The RDMA fabric has specialized monitoring tooling built around it. The front-end network is frequently less instrumented.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The GPU cluster is the new crown jewel of enterprise infrastructure. The network around it deserves the same visibility as any other critical asset.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Gi\u1edbi thi\u1ec7u v\u1ec1 DT Asia<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">DT Asia \u0111\u01b0\u1ee3c th\u00e0nh l\u1eadp v\u00e0o n\u0103m 2007 v\u1edbi s\u1ee9 m\u1ec7nh r\u00f5 r\u00e0ng l\u00e0 x\u00e2y d\u1ef1ng b\u01b0\u1edbc th\u00e2m nh\u1eadp th\u1ecb tr\u01b0\u1eddng cho c\u00e1c gi\u1ea3i ph\u00e1p b\u1ea3o m\u1eadt CNTT ti\u00ean phong kh\u00e1c nhau t\u1eeb M\u1ef9, Ch\u00e2u \u00c2u v\u00e0 Israel.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Ng\u00e0y nay, DT Asia l\u00e0 nh\u00e0 ph\u00e2n ph\u1ed1i gi\u00e1 tr\u1ecb gia t\u0103ng khu v\u1ef1c v\u1ec1 c\u00e1c gi\u1ea3i ph\u00e1p an ninh m\u1ea1ng, cung c\u1ea5p c\u00e1c c\u00f4ng ngh\u1ec7 ti\u00ean ti\u1ebfn cho c\u00e1c c\u01a1 quan ch\u00ednh ph\u1ee7 tr\u1ecdng \u0111i\u1ec3m v\u00e0 c\u00e1c kh\u00e1ch h\u00e0ng h\u00e0ng \u0111\u1ea7u thu\u1ed9c khu v\u1ef1c t\u01b0 nh\u00e2n, bao g\u1ed3m c\u00e1c ng\u00e2n h\u00e0ng to\u00e0n c\u1ea7u v\u00e0 c\u00e1c c\u00f4ng ty thu\u1ed9c danh s\u00e1ch Fortune 500. Ch\u00fang t\u00f4i c\u00f3 c\u00e1c v\u0103n ph\u00f2ng v\u00e0 \u0111\u1ed1i t\u00e1c kh\u1eafp khu v\u1ef1c Ch\u00e2u \u00c1 - Th\u00e1i B\u00ecnh D\u01b0\u01a1ng nh\u1eb1m th\u1ea5u hi\u1ec3u r\u00f5 h\u01a1n c\u00e1c th\u1ecb tr\u01b0\u1eddng v\u00e0 cung c\u1ea5p c\u00e1c gi\u1ea3i ph\u00e1p mang t\u00ednh b\u1ea3n \u0111\u1ecba h\u00f3a.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Ch\u00fang t\u00f4i h\u1ed7 tr\u1ee3 nh\u01b0 th\u1ebf n\u00e0o<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you need to know more about Monitoring AI Infrastructure: Network Visibility for Security and Operations Teams, you\u2019re in the right place, we\u2019re here to help! DTA is Netflow Logic\u2019s distributor, especially in Singapore and Asia, our technicians have deep experience on the product and relevant technologies you can always trust, we provide this product\u2019s turnkey solutions, including consultation, deployment, and maintenance service.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Nh\u1ea5n v\u00e0o \u0111\u00e2y \u0111\u1ec3 t\u00ecm hi\u1ec3u th\u00eam:&nbsp;<a href=\"https:\/\/dtasiagroup.com\/vi\/netflowlogic\/\">https:\/\/dtasiagroup.com\/netflowlogic\/<\/a><\/p>","protected":false},"excerpt":{"rendered":"<p>Organizations deploying GPU infrastructure for AI training and inference are facing a new set of networking challenges. Unlike traditional enterprise applications, AI workloads generate unique traffic patterns, including large-scale data ingestion before training, high-bandwidth model checkpoint transfers to distributed storage, continuous inference requests from thousands of users, and ongoing management traffic that crosses cluster boundaries.<\/p>","protected":false},"author":11,"featured_media":16353,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[56],"tags":[],"class_list":["post-16352","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-articles"],"_links":{"self":[{"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/posts\/16352","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/users\/11"}],"replies":[{"embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/comments?post=16352"}],"version-history":[{"count":2,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/posts\/16352\/revisions"}],"predecessor-version":[{"id":16358,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/posts\/16352\/revisions\/16358"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/media\/16353"}],"wp:attachment":[{"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/media?parent=16352"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/categories?post=16352"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/tags?post=16352"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}