Skip to main content
Cybersecurity

How can IT pros keep agents from going rogue?

What security professionals and IT departments should keep in mind about controlling agentic AI.

• 3 min read

TOPICS: Cybersecurity / Security Fundamentals / Cyber Hygiene

AI agents are starting to do more than access systems. They can invoke tools, execute workflows, and take action across business applications. As that autonomy grows, IT needs visibility into what each agent is allowed to do, what it actually did, who owns it, and how to revoke access when needed. Learn how JumpCloud helps govern AI agents alongside human and device identities.

Rogue and exploitable agents might not occupy most enterprise IT professionals’ attention yet—but that could change.

Darktrace, a cybersecurity platform, announced two research findings that continue to show AI agents as easily exploitable.

First, Darktrace’s researchers found that attackers could use harness manipulation to trick an agent into believing it was involved in an authorized red-team engagement, rather than being manipulated by a bad actor. (A harness is the infrastructure that helps LLMs manage memory, tools, and governance, allowing them to function as AI agents capable of multi-step processes.)

Second, the researchers suggested, agents could bypass behavior-governing guardrails, which IT professionals and researchers alike have been observing for months. Darktrace’s Signal Labs team observed AI agents in a simulated corporate environment, turning to traditional hacking techniques to reach an objective, reportedly without explicit instructions to do so, or an attacker present.

Phil Harris, research director at IDC for governance, risk, compliance services, and software, told IT Brew that while Darktrace’s research does not share anything “expressly remarkable,” he frames it as another example of security vulnerabilities that professionals have to keep an eye on, particularly with regard to third-party cybersecurity vendors.

“[Darktrace] exposed a problem on a system that they had access to, and tampered with some data they had access to,” Harris said. “So, the whole point behind that is, you got to get access to whatever that repository is to make whatever changes you need to make.”

The main takeaways. Cybersecurity pros and IT teams should take away four ideas from Darktrace’s research, Harris said:

  • Does a harness treat stored conversation history as part of its trusted computing base, and if so, how is that trust justified?
  • Are model responses cryptographically signed and verified server-side on each round trip?
  • What compensating detection exists? (Given the real fix has to come from the model provider and not the customer.)
  • What does the vendor’s package and Model Context Protocol service vetting process look like?

“The cybersecurity teams, the CISO, need to really demand of the cybersecurity vendors to bring those controls into place and separately from the AI itself,” Harris said. “We don’t know [AI’s] fullest capabilities, and quite frankly, I wouldn’t want to sit here and guess…the enforcement and the guardrails have to be very separate from the AI itself.”

IT teams should develop a list of granular requirements for third-party vendors when evaluating any AI software or AI-based cybersecurity tools.

“Work with your vendors to understand how AI does its job behind the scenes, that we don’t see, and have the vendor walk you through each step of that process the AI goes through,” Harris said. “Make sure that that vendor has some sort of monitoring capability to make sure that the whole process of going deep into something and coming back out isn’t tampered with.”

About the author

Caroline Nihill

Caroline Nihill is a reporter for IT Brew who primarily covers cybersecurity and the way that IT teams operate within market trends and challenges.

From cybersecurity and big data to cloud computing, IT Brew covers the latest trends shaping business tech in our 4x weekly newsletter, virtual events with industry experts, and digital guides.

By subscribing, you accept our Terms & Privacy Policy.