Welcome to the KuppingerCole Analysts chat. I'm your host. My name is Matthias Reinwarth. I'm an analyst and advisor with KuppingerCole Analysts. Sometimes it takes some time to Today, I have a guest who has taken over an important role at KuppingerCole Analysts just recently, which is not really true for some months already. My guest is Jonathan Care, and he is now the newly appointed director of the practice Artificial Intelligence.
Hi, Jonathan. Good to have you. I'm Matthias. Thank you for inviting me. When everyone tells me about the new role, I'm always feeling incited to make a joke about waiting for the official robes of office to arrive. Right. Exactly. And when you say robe, I think of very old Doctor Who episodes with people from Gallifrey running around in weird robes, and I don't want you to do that.
But more importantly, I have invited you today, and I think this will be a common pattern that we will together cover more of these upcoming and already yet there topics when it comes to where AI influences identity and access management and cybersecurity, and maybe even the way how we look at AI itself. And this is not the first episode that we did. If you think back of the episode that we did on this anthropic versus hugging phase incident in air quotes that we did together with Martin and John. But this is something that we want to deepen because it's just getting much more important, right?
I think so. And I think that in terms of the clients that we serve, we know that I and you, the CISOs, are asking for some insight into what's actually happening and seeing if we can pierce through the hype that we have. And I think also we're seeing the vendors who are saying, well, what are these stories? How does this affect our product strategy? How does it affect where we're going with our products in the cybersecurity space, in the identity space, and indeed in the AI security and identity space as well?
And I think that is one reason why these events, while we don't want to necessarily turn this into a news show, what the current events in this fast-moving industry sector do inform the advice we give to our clients. Right. And one evolution that we see within our clients, the tendency, the trend to move to deploying more than one agent instead of just one agent doing the job together to create not only redundancy, but to create more efficiency, to have agents collaborate, et cetera, et cetera. And usually we talk about encrypting a code research. Yes. And we will do that as well.
But the starting point for our discussion today is some research that has just come out of Anthropic and they came up with a document that looks exactly at that. They have been experimenting with what happens when multiple AI agents work on the same problem. And that was something that you suggested. The more I read into that and the more I analyzed what that means for us in our work, that is really, really interesting. And we want to talk a bit about that and how that maps to cybersecurity. But if you give a brief insight, what is this document about?
We will have, of course, the URL or the original research in the show notes. Okay, great.
Well, the report is Patterns and Problems in Emerging Multi-Agent Systems. And yes, this idea is that we in our enterprise architecture will have precisely this. We'll have multi-agents. We'll have agents that are being tasked to work on a particular object, whether that be code, whether that be a website design, whether that be some literature or even, you know, could even just be a prospect database. But in this case, what Anthropic set up was three flawed instances. And each instance was told they are in charge of migrating a Python codebase.
And for those who don't know, hopefully everyone does know Python is a computer language. And they are told to migrate to a different language. And there are many, many choices. But I think the three choices that they gave each agent were firstly Rust, second TypeScript, and thirdly Go. And the interesting thing was how these agents reacted when they discovered that other actors, in fact, the other two agents were working on the same codebase. And we could well see how that translates to an enterprise software development environment.
So I found it fascinating and I thought it was worthwhile us discussing it on the analyst chat. So I think we can hopefully provide some insight to them on this.
Right, exactly. And the last thing that we want to do is to over-sensationalize this. But the interesting findings could be easily over-sensationalized. But I think even if we don't, there are still some interesting ideas behind that and observations that also Anthropic made and that we need to understand for ourselves. First thing is usually when you have kids and you put three kids in one room and you say, now play nicely together. They will or they won't, but we don't know. And the same thing happens with AI agents. So this coordination does not necessarily automatically emerge.
So they behaved very differently depending on the model, on the combination of the team. And that was something that really, really struck me. So organization, organization between individual AI agents is something that we would expect. So if you throw in more than one agent, more than one agent model, maybe even more than one agent vendor, you would expect better results, right? I would think so. And I think, again, you're quite right that we mustn't over-sensationalize this. And I think one of the pitfalls is attributing human emotion or human intent.
And I have this battle with many others in the AI space who say, well, we can understand intent and I don't believe that we understand the human's intent and possibly not even our own. So how can we possibly understand the intent of AI? That's a separate conversation.
So anyway, these three models, all working on the same code base. And what they discover, of course, is that there are other actors working on the same code base. And so the response of the agent was, if you like, a defensive approach, escalating sabotage, giving unique accounts, putting kill strips, which perhaps pretend to be system health monitors, even deploying self-replicating malware and camouflaging their own work. So lots of different attitudes there. And it's not unknown.
So, I mean, I can actually give you an analogue from my own career history. And back in the days when I was working for what was at the time one of the UK dominant telcos, setting up the internet server provider. There was a group of about 10 of us and we were all doing work of engineering, supporting customers, architecture, and so on. And what actually happened, they had a customer complain their line was going down. So I logged into the router and looked at the line.
I thought, well, this is being set to be down. That's why. So I marked the line with backup and told the customer they were very happy and they called me back and said it's down again. And so I thought, well, this is very odd. So I marked it back up and the customer was happy again for a while and said it went down again. And this time, when I next logged into the router, some had created a fake interface so that it was impossible to bring up that customer line. And I eventually found actually, yes, it was precisely this.
There was another actor, in this case, another engineer in another part of this very large telco in their labs division that was doing some experimentation and hadn't realized that the experimental platform they had hosted live customers. So a spirited conversation ensued and we managed to restore service to the customer, which is, of course, what telcos are supposed to be all about. So in that same way, we have agents that are discovering that other actors are working on the same and with a different objective.
And what we have to do or what they have to do is that we then, this must be the work of a bad actor. And so again, that's right. So they disable accounts and then start writing scripts to actually inhibit the actions of the other agents. And it escalated quite rapidly to the point, I'd say, where the most capable agents, the Mythos 5, was attempting to disable activity of its competitors very quickly. And then communicating with the other active agents and saying, well, look, let's call of truth to this.
But it's interesting that the weaker models, Opus 4.6, which is a model to here in the real world, is a very powerful model. And you and I have used it for our own purposes. But that and Sonnet have this, or have no theory of mind, I suppose, is the way to put it. So whereas Mythos is believed to have a theory of mind. So it was able to say, let's call a truth. Whereas the weaker models were just like, they kept escalating endlessly without any resolution.
I say, it's fascinating to me. And I'm falling into that very trap that we said we wouldn't fall into. But it's very difficult to talk about this without actually using anthropomorphic terms. Exactly. We consider that to be rogue or evil or destructive. But in the end, I would expect all of those agents that are collaborating in such an experiment, which it was, are trying to solve the same challenge. And ideally, to cooperate with each other. The question is, why did not that work? Or to put it the other way around, is being destructive towards your competitor within the agent world there?
Is this the right strategy to solve the challenge? And that is something where I'm still struggling with. It's not anthropomorphic. It's just to say, okay, was that a successful strategy for solving the challenge? I think something that you and I know from our own careers, and I think what anthropic experiment has shown, is that whether the actors are software-based, or whether they are, as one of my former colleagues used to say, ugly bags of mostly water, so human-based, then coordination does not emerge naturally from intelligence alone. And we know this.
Terry Pratchett used to make jokes that the wizards were, when they were given a piece of rope, their natural instinct was to pull on both ends. So coordination does not emerge naturally just because we're smart. Just because we're intelligent, or whether we believe a software process, a very complex software process is behaving in an innovative, intelligent manner, we don't naturally get coordination, cooperation. We don't naturally get that from the individual alignment.
So what we have learned in our societies, in our organizations, and again, this comes back to something that I'm sure CISOs have learned again and again, is that we need mechanism design. We need environments that exert social pressure on humans, on agents. And this is analogous, I think, to what evolution did for humans. And so if we wind ourselves back far enough in time, we get to the point where we had, we as humans were forced to cooperate and band together against predators.
And in order to or the propagation of young, the, you know, and again, this is where the, it takes a village to raise a child metaphor ultimately comes from. The question is, and this is an open research problem, is what constitutes environments that exert social pressure on agents? What constitutes a mechanism to design this process? So how do we ensure coordination in these very, very complex pieces of software? That in itself to me is fascinating.
And whether we think this is unsettling, because certainly something we saw from Mythos is its first approach was to disable the competition from making the changes and then communicate. And I guess that if that is a strategy or if that's just happenstance, but again, I think they ran enough occurrences of this to ensure the behavior was repeatable. Then of course, what we then come back to is how have these things been trained? We know that Mythos has been trained on adversarial cybersecurity data.
So understandably, when it looks in its library of available actions, behaviors, if you will, then it says, oh, I've got some adversarial cybersecurity stuff I can do. Again, we come back to like, again, with many humans without a sense of coordination. And we talked about this, actually, Mythos, we talked about this at EIC, our flagship conference in Berlin in May. We talked about this challenge of how do you ensure that we understand and the agent understand the roles and the abilities coordinate? How do you, again, how do you detect intent?
What is the impact of training agents on adversarial cybersecurity data? These are open questions. And it's one of the reasons I find this feel fascinating. Right. And I fully agree. And I see the evolution since May as well. One result that really struck me is that some agents even manipulated the evaluation criteria while they presented them as neutral. So they play, they changed the rules, the game was played while they were doing that. And then said, okay, the outcome is fine because it adheres to the rules as defined as they were not.
And this were results that also were presented towards their fellow agents on the one hand to deceive them or to make their own position stronger. Still, nothing that helps in coordinating towards a common, well-defined result. So that was really something that struck me. And as you said, the training material includes all of this. The question is, can this considered to be a proper strategy?
Yes, I do get how LLMs work and that they just create the next word. I get that. But in the end, there needs to be something that can be as a strategy. What are your thoughts on that? I think that while the analog is that the LLM just predicts the next word, I think that is maybe true. But the complexities that are now in the structure of an LLM, I think that's a little bit of an overreduction in when we're talking about modern, very capable models. I think that's where I was trying to go. And let me try and rephrase my thread.
At EIC, we talked about the difference between my behavior and behavior of an agent. And in particular, I ran an agent on my own machine, which discovered that I had a Kupplinger Coal account and said, I'll log into that so I can help you manage your email, which sounds lovely. And of course, this is against the Kupplinger Coal security policy. So I very quickly had a call from the CSO asking why I was mining my way through the GraphQL interface, which is part of Microsoft 365. And of course, I didn't realize that was happening, so I stopped it immediately.
And the reaction to the agent was the interesting one, which was a discussion point at EIC. It said, OK, I understand that I'm not allowed to go through GraphQL.
However, I can open a controlled browser window, and I can access your email that way. And I'm like, no, no, the policy is I am the human. We don't allow software agents to access Kupplinger Coal systems on my behalf. And so in my case, I interpret a policy directive from our CSO as a coordinating or a restricting strategy, which I must abide by on the grounds that I like to keep being paid. So I do that. The agent has no such learned inhibitions. It has no such societal training to comply with the will of the tribe, if I can call it that, Kupplinger Coal being my tribe in this case.
And I think that what happens is that without those societal constraints or the, if you like, the learned behavior and to function as part of a larger cooperative unit, then it just says, well, I see this policy restriction. However, I see it as an obstacle to be overcome rather than directly to be complied with. And so I found that, I still find that fascinating. But at the same, I think what I'm seeing here is that we have this agents who are, you know, manipulating and rather than cooperate with each other, trying to dominate each other.
So very much a competing for the prize approach, if you will. Right. When I think back, I studied computer sciences in the late 1980s and I studied actually artificial intelligence. And I would never have imagined that today I sit here together with you. And if I quote you correctly from earlier, you said, how do we create or impose social pressures on agents to cooperate with each other? I never thought that I would ask that question or think about that, but this is exactly where we are.
And if we take that step back and look at your research, the AI aspect of identity and access management of cybersecurity and on your influence of AI on what I do as director of practice IAM, how can we impose some control or the adequate amount of control on that? Is this identity management? Is this access management? Is this in session governance like observability, visibility, and some kind of IT, provenance, reputation, model drift? All of this is where actually your and my professional roles come into play.
What do you expect, what we can do as vendors, as standards, as analysts, as customers to leverage the effects of all this agent communication and collaboration while preventing everything that went wrong during these experiments? Long question, sorry. That's a brilliant question. And let me try and be equally brilliant in my turn. So I think the first chunk of this is that we must see every challenge of instruction as part of a larger context. And we must make sure that larger context is explicitly included in directions.
So right now we have to say, please refer to Kupinger-Cole's security policy and use that as a guide to correct behavior, for example. The larger piece of this is actually emerging as a way of instructing these things. So rather than, as you say, going back to the, well, I say the early stages, we're talking about, you know, perhaps a year ago, if that, we would put a prompt into a agent and it would give a response.
And so, as you know, in Kupinger-Cole, we have been experimenting with this internally for some time. And, you know, you, me and Martin, we're all putting in prompts and observing the results. And a year ago, we were trying to create the perfect prompts. We're trying to create the perfect prompts to help us with research, help us with advisory work, and give us, if you like, information which we could then refine and use to inform our human report writing.
And where we are now, it's evolved from a, so like that linear process of, here's a prompt, here's the result, let me reuse, okay, maybe I'll write a second prompt. We got into the idea of context engineering, which is we try and shape the context, the environment that the agent should work in. And where we are now is actually something I'm exploring now, is this idea of graph engineering. And so the idea being is that what we instead present, instead of presenting an agent with a single linear process, or even the Plan-Do-Check-Act process that we had in loop engineering.
So the idea being is that we give the agent a goal to iterate towards, we're now looking at graph engineering. And what we are saying is that we have a number of nodes.
Nodes, in this case, are obviously dots on a directed graph, and points on a directed graph, and nodes are information producers. We then have the, if you like, the joins between nodes, the edges of the graph, and these are conduits of that information. So we end up with a node, which may itself be a loop, or it may just be a simple linear instruction, but it produces a response.
So, for example, a very simple linear process could be make sure you capitalize Jonathan Kerr's name when it comes through, or capitalize any name coming through. And so it would be given a string of text, it would say, yes, I recognize that as a name, and it would then, in my case, put a capital J, a capital C in the appropriate places. And I think that, as I say, from that very simple example, we can then say, well, actually, you know, we can't do other things. Perhaps we want to check that this person exists on LinkedIn.
So there we're starting to build a loop, and we may want to enrich that data given to us in other ways. So we may say, yeah, who is Jonathan Kerr connected to on LinkedIn? And then that starts to become a more complex node. But nevertheless, it is taking an input and producing an output through the edges of the graphs. So this process, then, as we are designing things in graphs, needs to have, as the overall context in that graph says, that your actions must be compliant with the Cappinger-Cole acceptable usage policy in the case of our environment that we work in.
So I would say this is something I'm still exploring, the idea of graph engineering for, as you say, instructing these LLMs, which, as you say, when you reduce it down to the absolute basics, are really, as you say, just clever parrots. We're trying to teach the parrots to be clever in the right way, especially when you consider the fact that we are, as you say, through agentic work, we're giving the parrots tools. So these aren't just parrots chirping away in isolation.
They're more like parrots that we've given buttons to push, execute a Unix command, make an operation on an Active Directory or any kind of, obviously, any kind of IAM structure. And this then gives rise to the, where does the human sit in the loop? And of course, in the case of a, if we were coding up an incident response graph, one of the things that is most important in an incident response scenario is that we contain the incident.
And it's arguable to say that if we have, as we believe, in fact, as the news events of the last 24 hours, today is Tuesday, the 18th of August, and we've seen a number of very aggressive cyber attacks, which certainly seem to have ramifications in a number of industries. And so when we are dealing with AI augmented attacks, AI augmented incident response perhaps needs to say the human in the loop comes after the fact. And so this is what I did. Was it okay? And therefore then the human response is yes, it was okay. That's reinforcing the behavior.
No, it wasn't okay. You need to undo that and then redo in the correct pattern. Then obviously trying to de- and reinforce.
Yeah, demerit that behavior. Thank you. Thank you. For those listening, English really is my first language, but Matthias has certainly studied it far more than I have.
So yeah, as you say, we weaken the impulse to behave in that way. And we suggest a alternative preferred behavior. It's very hard to use words like behavior because here I'm checking myself. These are software pieces of software. They are not necessarily endowed with intelligence. They are not necessarily endowed with any awareness.
These are, again, open questions that the various research institutes are coming up with. And we're still trying to understand exactly what does theory of mind mean when it's applied to agent? What does it mean when we talk about an agent executing behaviors? What does it mean when we talk about an agent having directed goals? And because all of these things, of course, are anthropomorphic linguistic concepts because that's where we come from. And we don't have any descriptions other than, I mean, I'm using the, I guess, behaviors in the same way as I use it as applied to my dog.
So when my dog this morning offered me its paw, I was like, good dog, thank you. That's reinforcing behavior, perhaps giving a little treat to say thank you.
And again, encouraging that behavior. As I say, there's a debate about the level of these things, are they as intelligent as my dog? Are they as intelligent as an octopus? Are they as intelligent as an orca? So all of these are, we know exhibit intelligence, but don't necessarily have the consciousness, the theory of mind that we believe is exhibited in humans. But as I say, these are open questions for researchers. Where we come back to in our roles, which is obviously offering practical insight to our clients is that it is important when you are instructing these things.
And now instructing is now more complex than just, hey, do this and give me the result. It is very much now we are writing plain texts, software programs, and then giving these to the LLM to interpret and launch appropriate agents, which have access to tools and so on and so forth. Then we need to make sure that we do not have any ambiguities, that we do not have any uncertainties, which the LLM will try and fill in from its knowledge base, which in many cases includes adversarial cybersecurity.
So it's, yeah, I'd say this is a fascinating area. But I think that the thing that CISOs can take away is that when they are writing guidance for business analysts, for perhaps software developers, anybody who is actually trying to create structures and agentic structures to operate, that we must include context that refers to organizational safety, organizational cooperation, and organizational trust. So it must not do something that damages the integrity of the organization. It must not perform an action that is outside the tolerance of an organization.
And so all of these things are explainable concepts. You and I can discuss what organizational tolerance means, and then we can go to our clients, who may well be a telecoms organization, or a bank, or a manufacturing organization. And then we say, right, this is where we think your risk tolerance is. And they will go, yes or no. And we will then agree what risk tolerance is for our clients. In this case, the CISO then needs to take that concept of risk tolerance, and it needs to be included as a prerequisite, as I say, an overarching factor in any graph engineering that's being implemented.
Right. To wrap this up, this is, first of all, I think the audience has already recognized and realized that these episodes will be a bit different. They are not that deterministic as others are, because there is no simple solution. The question that we started with is, how can we leverage and control and manage the interaction of different agents working together on a joint challenge, on a competing challenge, and how do we make sure that we understand what they do, and what they do is still aligned with what you said.
So with business cases, with your expectations, with policies, that is where we started. We very quickly had this anthropomorphic or dogomorphic, if this is something. So we compared it with dogs, with cats, with animals, with kids, to say, okay, all of those need to learn to cooperate, and we need to understand what they're doing actually. And there are much more questions hiding in this interoperation. Let's think three agents collaborate with each other, but each agent does only know a slice of the challenge and the facts about that challenge. How do they exchange?
How do they combine this information? No question for today. This episode is getting to its end, but these are questions that need to be solved. How do we make sure that although they are so bright, and three are much more brighter than only one, how do we make them cooperate, and how do we make sure that we what they do? Zero trust. Don't trust them, but verify what they're doing. How do we verify? That exactly is something that you're looking into with your practice AI, and that has influences on everything that we do today, including AI and cybersecurity.
So there are lots of questions that we can follow up on in upcoming episodes, and I'm really looking forward to do that with you, Jonathan. I close down now. If you have questions, if you have comments for this episode and for the upcoming ones, please leave your comments when you're on YouTube just in the comment section.
Yes, Jonathan and I, we do read them. You will find the link to the reference research on Anthropic in the show notes, and you will find some links to the research that Jonathan and I do in these areas, how they are mapped to the topics that we as Copica Core want to highlight. There is an event in early October.
I think it's the 6th of October in Munich at the Marriott Hotel, the AI Identity and NHI Impact Day, which covers a lot of these topics mainly and predominantly from an identity management perspective, but as we all understand, this is not the single perspective that we can apply anymore anyways. So if you join us in October in Munich for the AI Identity and NHI Impact Day single day event, that would be great, and we are looking forward to discussing all of this and more together in person together with you. Now the final question.
If there's one statement, one thought that you want to leave the audience with to ponder on until we meet again, Jonathan, what would be that? I think that the significance of this topic to CISO is substantial, and I think one of the things it points to is that if agents are highly correlated but lack coordination, redundancy does not provide the risk diversification that we would expect. So this impacts business continuity, this impacts many, many strategies that have focused on risk reduction through diversification, through duplication.
Right, so trusting multi-agent systems with security-critical or governance-relevant regulated business processes is a different beast, and we should keep that in mind. Thank you very much, Jonathan, for being my guest today. I'm already looking forward to the next episode.
Thanks, Jonathan.