When Cisco acquired Splunk back in 2024, I’ll be honest, I wasn’t convinced. As a former network engineer, my simple mental model of Splunk was that it was a glorified syslog server where logs came in, you could search them, and that was more or less the story. I suspect quite a few people reading this have had exactly the same thought.
A lot has changed since then, and I want to share some of the latest network observability insights we’ve developed at Data#3. The gap between what people assume Splunk is and what it actually does is one of the most significant missed opportunities I see in networking today.
Before Cisco completed the acquisition, a large Cisco environment may have looked like this:
If something went wrong, you had to know which dashboard to open first and then manually connect the dots across systems that had never been designed to communicate with each other.
Cisco bought Splunk because it needed a single place for all that data to land and a platform capable of making sense of it. What they acquired is an observability and resilience platform that can ingest data from dozens of sources, surface it in a unified view and help you understand not only what is happening on your network, but also what it means for the people and services depending on it.
That reframing, from log aggregation to genuine network observability, is the foundation for everything we’re now building for our networking customers at Data#3.
Network observability can be a bit tricky to wrap your head around, so I like to narrow it down with this analogy. Put yourself in the shoes of a health provider. When a network problem occurs, a nurse doesn’t call the helpdesk and say, “I’m having an issue with this access point.” She says, “I’m at ward four, and I can’t access my patient records.” The operations team’s job is to translate that description into a technical root cause, often across multiple systems and under pressure.
This is where Splunk delivers real value because, in this scenario, Cisco Identity Services Engine (ISE) raised an authentication issue on Wi-Fi, the access point is managed by Cisco Meraki, and the firewall has flagged the same device. Without Splunk, these are three separate alerts and potentially three separate teams working without realising they’re looking at the same issue. With Splunk, they become one actionable incident.
That’s data resiliency in practice. The faster you can correlate events across your environment, the sooner you can detect and remediate. That speed is the core benefit we discuss with customers.
Splunk ships with technology add-ons (TAs) for its major integrations. You can download the Cisco ISE TA, the Catalyst Centre TA, and get dashboards out of the box. The limitation is that those dashboards reflect what Splunk thinks you want to see, not the way your team really works.
The Data#3 approach starts differently. We use an agentic AI workflow to engage with a customer, understand their environment and then build a custom TA tailored to how their organisation consumes network data. A network engineer troubleshooting an outage needs different information from a network manager tracking infrastructure health and both need different information from a CTO trying to understand the cost of a two-hour outage to the business. We build one TA that serves all those layers, with views shaped to what each person needs to see.
The resulting package is stored in a code repository owned by the customer. If they add 100 new switches tomorrow, the dashboard picks them up automatically because it’s built around how data flows through the network, not a fixed list of devices. When they want to add a new data source or change how something is represented, they can come back to us. We version-update the TA, and they have a new package they can deploy to their Splunk Cloud or on-premises Splunk environment. The code is always there, always backed up and never tied to a single piece of hardware.
One question that comes up often is whether this only works for Cisco environments. The short answer is no, because Splunk ingests data and doesn’t care where it comes from.
We recently spoke with a customer who was running Aruba (now HPE) across part of their network and Cisco across another. Their network team needed to see both in a single dashboard, and Splunk handled it without difficulty. Similarly, if you have Cisco switches and Palo Alto Networks firewalls or a mix of vendors across your data centre and campus, Splunk can handle it. While out-of-the-box content is deepest for Cisco Catalyst, Cisco Meraki or ThousandEyes, a mixed environment is absolutely workable and in some cases the consolidation benefit is even more pronounced.
What makes this particularly interesting is where the industry is heading next. Observability platforms are no longer being used solely to collect and correlate data. Increasingly, they are becoming the source of context that operational and AI-driven tools rely on to make decisions. The more complete and connected the data, the more valuable those platforms become, which was one of the major themes I saw reinforced at Cisco Live.
I was able to attend Cisco Live in Las Vegas earlier this year, and what I brought back shifted my thinking about where all of this is heading.
Up until this point, I’ve talked about observability in terms of helping teams understand what’s happening across their environment. What became clear at Cisco Live is that the value of observability doesn’t stop there. The same data that helps engineers correlate and troubleshoot issues is now being used to support a much broader operational model, one where AI can work alongside network teams to analyse problems, recommend actions and, over time, help automate remediation.
The most significant thing I saw was Cisco Cloud Control paired with AI Canvas. Cloud Control is designed to be the single control plane for all things Cisco, bringing those 20-odd domain controllers together in one place. AI Canvas is the workbench that sits on top of it and in the upcoming release, it’s going to be agentic: capable not just of helping you troubleshoot a problem, but also of starting to fix it for you.
The workflow will look like this:
In the next release, that process becomes automated. As the AI starts remediating, it learns over time which fixes it applies reliably enough to request access to apply them without waiting for human approval. You can still choose to allow or deny it, depending on your environment.
This isn’t a distant aspiration, as the first alpha release of root cause analysis capabilities was shipping at the time of writing, with the full story set to be told at Splunk’s annual conference in September.
I also learned that Cisco is building Codex directly into Cloud Control, enabling partners and customers to write their own applications on the platform. The marketplace for those apps will be tightly governed, but as an organisation with early access and hands-on experience, Data#3 will be able to build custom Cloud Control applications for customers who want capabilities beyond what ships out of the box. Having seen how that works firsthand, in a closed beta available only to US customers at the time, puts us in a position to help our customers take advantage of it well ahead of the market.
Any conversation about agentic AI connecting to production network infrastructure must address security, and this one is no different.
On the Splunk side, the agentic security operations centre is a capability in development that enables Splunk to triage security events and apply remediations it has learned to apply reliably over time. The goal is that instead of an engineer approving each fix manually, they receive a monthly report indicating that this issue occurred ten times and was resolved ten times. The human stays in the loop at the level they need to be, rather than at every individual event.
For organisations concerned about AI systems connecting to their environments, Splunk’s AI Defence product offers a way to think about this, functioning as a firewall for AI workloads and sitting between AI requests and the systems they interact with. Splunk has also recently acquired Galileo Technologies for AI observability, bringing a small-language-model approach to request validation that checks AI outputs locally against the customer’s own data rather than sending every request to an external model for verification. That makes validation significantly faster and cheaper while protecting against the kinds of cascading errors that concern most enterprise security teams.
If you’ve been running Cisco networking for a while and you’ve thought of Splunk as something the security team uses, this is worth reconsidering. The network observability use case is mature, the customisation capability that Data#3 brings goes well beyond what you get with standard TAs, and the roadmap from Cisco and Splunk is moving in a direction that will make this increasingly central to how enterprise networks are managed and maintained.
I’ll be sharing more of this at our upcoming lunch events across September and October, including how we’re applying an agentic approach to Splunk, examples of the dashboards we’ve built and what Cisco Live confirmed about where Cloud Control and AI Canvas are headed.
If this is relevant to your team, I’d welcome the opportunity to discuss it with you in person.
Information provided within this form will be handled in accordance with our privacy statement.