跳转至主要内容

DII for proactive observability, automation, and security

Two IT professionals monitoring servers and reviewing system data in a data center.
Contents

分享该页面

NetApp 架构标识
Ramanjaneyalu M.

DII for proactive observability, automation, and security

When you operate infrastructure at the scale required to support NetApp’s global engineering organization, observability isn’t optional; it’s foundational.

Across our engineering data centers and cloud environments, we manage tens of thousands of virtual machines, petabytes of storage, Kubernetes platforms, and multi-cloud resources that support critical development, testing, and lab workloads. Maintaining stability, security, and efficiency across this landscape requires more than visibility alone. It requires a platform that turns insight into action. That’s why NetApp Engineering IT relies on NetApp Data Infrastructure Insights (DII) as a core operational capability.

As Customer Zero, Engineering IT runs DII in production, not as a proof point or QA exercise, but as a real-world operational tool. Every capability we adopt is measured against a simple standard: does it meaningfully improve how we operate at scale? If it doesn’t, we don’t use it.

How we evaluate DII features

The role of Customer Zero is often misunderstood. Engineering IT is not validating whether features work as designed because that’s a QA’s job. Instead, we evaluate post-GA capabilities in live environments and ask a more practical question: Does the product scale, and does this feature help us run Engineering IT better?

When new DII capabilities are released, they are enabled first in our production tenant. We assess how they fit into real operational workflows, including observability, automation, security, and governance. Feedback to product teams is grounded in operational experience rather than theoretical value. Only features that demonstrate a tangible operational benefit are adopted into steady-state operations.

Proactively detecting rogue VMs

One of our most critical proactive use cases involves identifying rogue virtual machines in our private cloud environment.

Engineering IT routinely deploys thousands of transient VMs. Occasionally, a misconfigured or runaway VM generates an unexpected surge in IOPS, putting backend storage performance, network infrastructure, and overall system stability at risk. DII allows us to monitor VM-level performance metrics in real time and quickly identify abnormal behavior.

When DII detects a sudden and sustained spike in IOPS, it initiates an automated response. Alerts are generated based on performance thresholds, the responsible VM owners are notified, and a defined remediation window is provided. If the anomaly persists beyond that window, the VM is automatically shut down.

This closed-loop, automated workflow protects shared infrastructure before users experience any impact, without requiring manual intervention.

Unified visibility across hybrid and multi-cloud environments

Engineering IT operates workloads across on-premises VMware and OpenShift platforms, as well as public cloud environments including AWS, Azure, and Google Cloud. DII provides a unified view of resource consumption across all of these platforms, enabling us to clearly understand what resources are being consumed, where they’re running, and how efficiently they’re being used.

While DII is not a billing or chargeback tool, it delivers the operational visibility governance teams need to identify underutilized workloads, highlight potential capacity waste, and inform capacity planning decisions. This insight is shared directly with internal governance teams, eliminating the need for Engineering IT to compile reports and enabling faster, data-driven decisions manually.

Enforcing security and configuration standards at scale

Consistency becomes a security requirement at the scale Engineering IT operates.

We use DII to validate that standard configurations are applied across data centers continuously. This includes verifying that storage encryption is enabled on critical volumes, enforcing best-practice volume configurations, and validating growth policies to prevent unexpected capacity issues.

Beyond configuration compliance, DII actively plays a role in workload security. We monitor file- and folder-level activity to detect anomalous behavior, such as potential ransomware activity or destructive actions. When suspicious patterns are identified, DII can automatically trigger snapshots, helping preserve data integrity while reducing the need for constant manual oversight.

Deep Kubernetes observability and automation through APIs

As our OpenShift footprint continues to expand, Kubernetes observability has become increasingly critical.

DII provides complete visibility into Kubernetes configuration, performance, and overall health, allowing platform teams to detect cluster-level issues early through metrics, alerts, and centralized dashboards tailored to their needs. However, dashboards are only part of the story.

Engineering IT also leverages DII’s extensive API set to integrate observability data directly into custom automation workflows. Teams can programmatically retrieve cluster inventories by site, query performance data, and feed DII insights into other internal tools. This API-driven approach allows teams to consume observability data within the workflows they already use, rather than introducing new operational friction.

ServiceNow CMDB as a trusted source of truth

DII plays a critical role in maintaining an accurate and reliable ServiceNow CMDB.

We use DII as a primary source of truth for configuration items, automatically populating the CMDB with virtual machine metadata, infrastructure relationships, and environment ownership details. This integration significantly improves change and incident correlation, reduces manual CMDB maintenance, and ensures configuration data remains up to date.

When incidents or changes occur, teams can trust the data they’re working with because it’s sourced directly from live infrastructure telemetry.

Why engineering IT uses DII instead of third-party tools

Customers often ask why Engineering IT doesn’t rely on generic observability platforms.

The answer is straightforward. DII understands NetApp infrastructure natively, integrates directly with ONTAP telemetry and logs, and aligns with how we operate storage, compute, Kubernetes, and cloud as a unified platform. Most importantly, DII enables automation-first operations, turning observability into action rather than static dashboards.

Customer zero lessons for the field

Running DII at scale reinforces a core NetApp on NetApp principle: the most valuable IT tools aren’t the ones with the most features, but the ones that quietly prevent problems before anyone notices.

From rogue VM detection and Kubernetes observability to security automation and CMDB integration, DII helps Engineering IT stay ahead of issues, reduce risk, and operate with confidence.

NetApp 架构标识

Ramanjaneyalu M.

Ramanjaneyalu M is a Senior Information Systems Engineer at NetApp with expertise in cloud infrastructure, Kubernetes, storage, and enterprise IT operations.

查看 Ramanjaneyalu M. 的所有文章
DII: Proactive observability and automation for hybrid IT | NetApp