Welcome to the KuppingerCole Analysts Chat. I'm your host. My name is Matthias Reinwarth. I'm analyst and advisor with KuppingerCole Analysts. Our guest today is Alexei Balaganski, and we want to talk about a topic we have not yet covered. But first of all, hi Alexei, good to have you. Hello Matthias, and great to be back again. Great to have you. We want to talk about a rather large topic and we want to scratch the surface for the first time. We want to talk about AI data provenance and AI data management in general. It's difficult to start.
There's maybe a brief definition of what you consider to be AI data provenance and everything that's related from your perspective. Right.
Well, Matthias, first of all, you are right. It's kind of a new, barely explored topic for not just KuppingerCole, but for everybody, I guess, including our listeners, simply because on the one hand, it's so obvious, but on the other hand, nobody is actually talking much about it until, I guess, it will be too late because we will be hit by some kind of regulation.
And yeah, one of the biggest challenges we had in living in the AI age is all those huge piles of AI slop, if you will. Low quality generated data that's obviously unusable and not true. The results of hallucinations and poor data decisions and low quality inputs, you know, garbage in, garbage out. That's what we get a lot. And I mean, if you are watching a poorly made, generated video on YouTube showing some hitters playing music instruments, I guess it's still okay. But if the same level of data provenance and quality management applies to your business decisions, you are in huge trouble.
This is exactly what we want to discuss and how to prepare for this and how to avoid all the consequences in the future. Right. So we all have learned the hard way, that providing good data as the foundation for good results of AI is essential. But on the other hand, this usage of the data and the overall topic of our topic today, the AI data provenance, is becoming an urgent issue. What exactly do you mean when you say that we have to expect that crisis, that we might get hit by the regulation, as you said? Where does it start and where can we see that as indication? Right.
Well, it all started basically when all those AI giants like OpenAI and other companies have started training their models. They needed huge amounts of data for training. And basically, what they decided to do is to scrape the entire internet with complete disregard to copyright and other regulations, without asking permissions from anyone. They just basically have taken the entirety of the data generated by humans and somehow accessible through the internet and used all the data to train their models. I guess it worked to an extent.
So we have all those wonders of modern chatbots and AI assistants. Unfortunately, it's now very difficult to differentiate whether this response was triggered by some training data from a reputable expert or a copyright owner or it was something from a Reddit post. You remember all those stories like a chatbot recommending you to put glue on your pizza or to misidentify poisonous mushroom or anything like that.
So again, garbage in, garbage out. If your data model is trained entirely on Wikipedia and Reddit, well, you cannot expect any sufficiently useful results.
And again, while it still kind of works for consumer-grade decisions, it would definitely catastrophically fail when you start applying it for critical business decisions or for critical infrastructure management. Even if there is no existing regulations for this yet, sooner or later, it will appear simply because, well, we cannot leave like that. We have to do something.
And again, at least when it comes to copyright, obviously, this has already started. Movie studios, artists, singers, you name it, all the content creators who actually invested a lot of the real human effort and money into the content are now complaining because all those things were used without their permission and without any compensation. So this is probably the first and the most important driver behind those future regulations, but we will definitely see more. Right.
And if I think just briefly out of the IAM context and the auditors coming in and saying, okay, why does Matthias have this access right? You need to have proper auditing, a complete trail of why this access has been assigned, and you need to prove that. So this is logging, this is documenting what has happened. If an AI does such a decision that Matthias should or should not have an access right, it gets blurred already to say, okay, will this decision be correct? On what is it based?
And if we get to complete AI-generated content, then it gets much more fuzzy and much more interesting when auditors come into play. So it's much more than traditional logging and auditing, right? Right.
Well, I guess you are completely right. You are hitting the nail on its head here. It is actually, on the one hand, this is nothing new. It is based on some kind of audit trail for everything. So obviously this is not something you have to invent from scratch just for AI purposes. You already have tons of the data as application logs, database logs, just basically traces of everything which is happening within your business, both business related and IT related, if you will, or identity related. The only problem is that it's not enough.
First of all, all the data is scattered across different silos. Second, as soon as AI models come into play, you have this kind of media break. You know that AI models are by definition non-deterministic because they never actually deal with the original training data. They deal with weights, tokens, hashes, some basically random numbers which were derived from the original data. And often this process is completely blind and proprietary. So it's very difficult to recreate that path back into the past reliably.
So yes, you have to think in more than one way. You have to have all that historical original data readily available and prepared for AI consumption. But you also have to have some kind of a race and explainability trail for the actual model reasoning. And of course, when you actually have an architecture like REG, where you reach out to third party data, you have to have the data traced as well, both the actual data and all the operations you have performed. So it's not new. There is no need to reinvent the wheel. You just have to juggle a lot of different wheels to make it work.
And this is perhaps the biggest challenge for the actual business. Right. And if you use your own data, your own business data as training material, I think, as you said, that can be well documented, the lineage of data, how you use it, how it was used to train the LLM. But if you consider everything that is within your favorite AI provider's model, if you consider this as your supply chain in the end, then it really becomes tricky.
And I think many organizations are not yet really well prepared to understanding that there needs to be auditability, traceability in the end, a proper way of dealing with, again, back to AI data governance. So this is really a challenge for many organizations where they just, where many just resign to say, OK, I can't do that. Can they?
Well, absolutely. I mean, the data lineage as a domain thing, like a problem is not new.
I mean, we have been doing this for decades. And in some industries, it has been a mandatory requirement anyway, for decades as well. When it comes to AI, you're right. It is a massive chunk of your software supply chain nowadays. And unfortunately, we do not have an S-bone. We don't have a software bill of materials for AI models yet. And definitely not for AI data.
Well, yeah, at the very best, you can trust the AI vendor claims that they were not using Reddit to train that model. But how do you verify that? Do you actually get some kind of a step-by-step explanation of the model thinking process showing that, yes, it was definitely not touching Reddit or Wikipedia or someone else's unproven data and so on and so forth?
No, not yet. But this will definitely be a critical requirement in the future, because again, this is what differentiates AI from AI slop. Right. And I think these are also different types of expectations when it comes to what you see. But when I use an AI, I want to create a proper set of results that meet my expectations. And that is quality, that is style, that is overall applicability for my given task.
But auditors or somebody who wants to make sure that everything has been processed the right way based on data that is well-documented, they have a different focus and there is a clash between the two. Can that be synchronized? Can that be made work in parallel?
Well, yes and no, because this is what we usually believe, that the auditor is out there to punish us for not doing the paperwork right, but not really. All that paperwork, I mean, it was created because of a long trail of blood lost by businesses of not following that paperwork in the past. So actually, the compliance auditor is on your side, because they want to make sure that you are doing your job properly. Because if you won't, well, you will have many problems in the future, which are not related to compliance at all.
It will be security issues or quality issues, obviously, legal repercussions if you fail to, for example, respect copyright owners and so on. So yes, it's up to you to ensure that your work is done properly, that all your feed, all the data that you feed into your AI models has proper provenance and lineage and quality, and you have the means to prove it. The only thing you have to do specifically for compliance is basically follow best practices and regulations and prescribed formats and protocols and so on.
But again, this is secondary. You have to be vigilant. You have to follow all this AI data provenance things by default, because if you won't, you will fail spectacularly sooner or later. The quality of your digital product will suffer.
Again, your legal standing and your brand reputation will suffer. And again, the compliance fine would be just icing on the cake. So you should never think about compliance as a means to an end. It's just the way to ensure that you are doing the rest of your job properly. Right. And the fact that we're talking about it right now, of course, the disclaimer, we are not lawmakers, we are not lawyers, we're not auditors. But in the end, why are we talking about it right now? What has changed?
Can we see that in real life, that the times of the Wild West and AI at least partially are over, so that we need to do things properly? What has changed?
Well, again, we are definitely not lawyers, but we say that lawyers are actually the first driving force behind this. Again, as I mentioned earlier, copyright is probably the most sensible driver behind it, because, again, you can easily move and demonstrate that the current AI implementations are disregarding copyright laws on a huge and massive scale. It kind of used to work for a few years. And this is how companies behind all those huge and successful AI models actually managed to amass all that knowledge.
But sooner or later, we are coming to the realization that, no, it will no longer work. You have to do something, you have to compensate the efforts of the original content creators, you have to follow the existing legal frameworks.
So, yes, legal people and regulations are behind the first push. But, of course, they are not the only ones.
We are, I mean, we're keeping a call, we are operating within the cyber security and identity markets, and we see a lot of similar push as well. If you are making security or identity decisions on the data, which you cannot prove is even remotely relevant and good enough to make those decisions, those decisions will be garbage. And which would mean that you will not block the data breach, you will let the malicious actor in on the stolen credential, and so on and so forth. The rule is universal. Garbage in, garbage out. And nowhere it applies more than in cyber security.
Because, again, a wrong decision, a single wrong decision is enough to open your entire digital infrastructure to a ransomware attack or a massive data breach. And it will be too late to fix. This is not something you can kind of backtrack. This is not something you can investigate with forensics. Too late. You have to prepare in advance. You have to make sure that all your decisions are made on proper data. Same story, same issue, same AI data provenance. Right. And when organizations are doing things properly and want to do things properly, how can you add governance to existing systems?
When we are all using AI, we are providing a chatbot as keeping a call. Many others are doing this as well. We are using that, our own data as training material. But how to add that proper level of governance to that? Run a system for half a year, then I say, okay, I need to add some proper governance to allow for this data provenance. Is this possible in retrospect?
Well, I mean, at least it's the billion-dollar question, if you will, because it's definitely not something that we can properly answer in five minutes or even half an hour. This is something that the entire industry and the entire humankind needs to decide on a global scale, and it will probably take years to figure out properly. But one thing which is clear is that you cannot bolt it on in retrospect, on the one hand. You have to actually ensure that everything follows the same principles.
All your data, all your third-party data, all your internal and third-party AI models, and of course, all the data flows between those endpoints, they all have to be covered by this governance framework. It is extremely difficult.
Yes, this is true. But on the other hand, again, you don't have to reinvent everything from scratch. Parts of this are surely already covered, especially if you are in a regulated industry already, like if you are a financial organization, a bank, you already have a lot of governance controls for your own data. Somehow we have to force your primary third-party suppliers to do the same.
Again, for AI models, it is difficult because we are very early on the stage of developing those formats and regulations, but this is definitely coming soon. There are some developments out there already. You remember we've talked about AI explainability for years. This is actually one of the primary use cases for explainability because explainability means traceability and auditability. If you can explain how exactly this decision was made, you are already halfway on showing the auditor that you are doing it properly. And if you cannot, well, you failed already.
So, yeah, this is difficult. This is very fragmented. This is the same. This is like an extension of your software supply chain security, if you will, which is a huge topic on its own and which is a topic which we are covering exactly now as well. Exciting times, tons of upcoming research, but again, we are not lost, you are not lost, we are in it together.
So, let's work together. Right, and I think you've mentioned both aspects, explainability on the one hand and non-deterministic systems on the other hand, which of course produces this inner conflict of achieving a full level of governance, of achieving the right level of explainability. But nevertheless, we can start to make things better or properly. What would be practical approaches to say, okay, this is the minimum to start with, documenting your data catalog, documenting your data, having the data still available for reference purposes, would that be a good starting point?
Well, obviously, if you cannot tell even to yourself where your data is, well, you are already failed. You have already failed even before the auditors come.
So, yes, this is absolutely a foundational crucial requirement. You have to know where your data is.
So, I guess you start as usual with almost every kind of security strategy. You start with your posture management. You have to catalog everything and not just once, but all the time. That's exactly the point. That's exactly the difference of dealing with non-deterministic decision chains. You have to basically keep traces of everything, of every decision, because you cannot expect that when you are redoing the same steps again, that they will produce the same result. It will never happen just by nature of the AI.
So, either the model, the actual decision maker has to have this tracing, auditing, protocoling capabilities built in. Maybe this will happen sometime. Or you have to build a universal extension of your existing data governance solution to cover that quote-unquote attack surface as well.
So, somebody has to do it. There is no single way yet, but again, the industry is evolving. What we observe, for example, among the data security suppliers, that they are absolutely thinking about this already. And they are already working on some solutions. I can only recommend reading our existing research on data security platforms, for example, and on data quality and lineage we've done earlier. All those things plug in together, at least they should. And of course, some specialized generative AI defense tools, if you will, for the lack of better terms, they've come into play as well.
They are very early in their kind of maturity and productization stages, but they are emerging. And we're also doing some research on those tools as well. They all somehow have to come together in another kind of a fabric, if you will. Kuping et al. has defined identity fabrics and security fabrics, and everybody is talking about data fabrics now.
Well, this is your AI fabric, AI data fabric, AI provenance fabric, whatever you want to, but you have to know what's going on. You have to trace every AI decision from the actual start, from the raw data you had supplied for training to the actual operational data you have supplied as context, to every step of the inference process, and finally, to all the activities which were made, let's say, by an agent based on that decision. Everything has to be tracked, traced, audited, and somehow combined to be able to provide a single protocol of everything. Do they have tools like that already?
Probably not, but they will have to, because again, this is what's definitely coming in the future, sooner or later. AI data provenance regulations will be the next GDPR, if you will. Exactly. I wanted to get to that quote exactly, because you said we were all struggling with GDPR compliance, because there is no compliance, there's only incompliance. If you fail, then you can prove it, but you cannot prove that you're compliant, which is always an issue. So it's always like a security issue or a security incident. As long as you think you're secure, nothing has happened, all fine.
If there's an incident, okay, then you have just failed. Maybe also just as a side note or a thought, if we think back to the high times of big data, I think many of the principles that had to apply back then need to apply to AI as well. So if you combine data of different origins, of different lineage, you need to understand what the result is, what the product, the report is, and how you apply proper access control and management to that as well.
With AI, this is the same principle on steroids, much faster, much more agile, but the same principle behind that. So if provenance is the new GDPR readiness, what should executives, those who are responsible, those who will most probably get punished when something fails, what do they do in the next 12 months?
Well, first of all, of course, they have to start thinking about it. Because if they don't do it today, they will be woken up in the middle of the night sometime next year and told that something really bad happened. So they have to start thinking about it today. But at the same time, I would urge them not to make too hasty decisions.
Again, just like with quote-unquote AI security a couple of years ago, there will be a lot of snake oil peddlers, if you will, who will tell you that just buy our product and you will have all this figured out for you. No, this is not how it's supposed to work. Primarily because nobody can decide for you your priorities, your risk environments and risk habitats. You have to make decisions, you have to prioritize what data you want to capture and trace and protect first, because it's your business decision.
And also you have to remember that at least parts of those requirements and future regulations are probably already addressed by existing tools. Again, if you have a database, you probably already have some data security controls in place. And you don't have to replace them with new AI aware ones. You just have to probably make sure that they support an upcoming logging format or something like that, or a connector to feed the data into a third party processing application. Watch this space, understand the strategic decisions, definitely look for interesting developments and announcements.
Come visit our conferences, for example, because this is where we discuss all those things. Watch our webinars.
But again, kind of before you make a purchasing decision, make sure that you understand what you are actually getting. Look behind the labels, never trust, always verify. That's kind of zero trust for AI data provenance, if you will. Should apply everywhere.
But again, your strategic aim should be not just compliance by default, but AI data provenance by default. If you know where all your data is coming from and what's happening to all your data, then you already have just like 90 percent of everything figured out, regardless of a specific use case. That should be your ideal to strive for. Right. And we're talking about an AI that works as it is supposed to be. We have current research ongoing that will be published soon, which looks into LLM security, protecting the AI from external malicious influences of any kind.
That, of course, needs to be a priority as well, to make sure that the AI works as you expect it to be and that it's not influenced by bad actors from outside. This is an additional aspect that we just did not cover in that context, but of course has influences on how the data is actually processed and even extracted or even additional data ingested at runtime. So prompt injection is a thing that we should talk about in another episode, right?
But again, let me just emphasize one takeaway from this. Do not focus on point solutions. Do not think that prompt injection is the root of all AI evils.
No, it's just one fairly well-researched attack vector. And again, if, for example, all your data is always encrypted, if all your LLM interfaces are always monitored, and you have some kind of a guardrail installed in place, then who cares that prompt injection will be handled for you? Because even if they are actually able to steal some of your data, it will still be encrypted. Or if they actually manage to overcome some internal LLM guardrail, they will still not be able to exfiltrate the data.
You have to think in terms of doing everything properly across all your environments, not just AI, all your data, all your data flows, ideally even in use. If you can manage that, at least for your highest and the most critical systems, then it's like a multi-factor authentication. They say it will automatically negate 95% of all the identity-based attacks. I guess you can find the same level of controls for your data as well.
And again, it's not just about tools, it's about processes and workflows. If you have them figured out and configured properly, you are automatically solving a huge chunk of your potential security and compliance problems. So always think strategically, always think from the other end, if you will. Not how to please your auditor or how to fight another ransomware attack. How to make sure you are resilient to both attacks and audits. This is the best way. Prevention is the best medicine. Right. And as you've mentioned already, we are all in this together. We are learning while we're working.
And so this is really maybe one of the biggest takeaways for me to say, okay, yep, we need to start doing things properly. When AI is part of our delivery process of our products, of our services, we need to deal with it as we would do with any other component that we provide. So provide the proper scrutiny, provide the proper governance to that. But nevertheless, this is a problem or a challenge at scale. Any final advice for enterprises that just feel overwhelmed, just over-challenged by this task?
Well, again, this is a billion dollar question. It might actually boil down all those recommendations and strategies into a single takeaway. We probably know we cannot. This is not something which we can do immediately.
But again, one thing is to not think you can sit through it and it will boil off and become another blockchain or something that nobody will carry in five years. No, this will never happen. Even if AI hype will end up with a crash and some large industry leaders will probably cease to exist in five years, it's not going anywhere. We are still a digital society. We are still digital businesses. We still have to deal with huge challenges. And to a large extent, there will always be some kind of a digital data supply chain problem.
And all these things we have discussed now, they apply to any kind of those issues, future or current. AI is just one.
Yeah, maybe a very popular, but just one use case of that. Again, strategically, do not think in labels or buzzwords. Think in risk, if you will, because this is what matters in the end. Exactly. And so all the truisms hold true. Start small, start at all, apply a risk-based approach, and really solve the problems that are closest to your business first, and then continue learning from others. And learning from others, you've mentioned our events, our research. We will cover this, obviously, as this is part of security, and identity management plays an important role.
So this is all a part of the solution, and we are here to support you. And Copenhagen will look at that market in much more detail soon.
Of course, I need to recommend the European Identity and Cloud Conference as one of the key cybersecurity events with a strong focus on IAM, but beyond. And our research is essential, I think. Thank you very much, Alexei, for being my guest today. This will not go away, and we will cover this topic in more detail and with different angles, and looking also really at more of these point solutions sometimes when it comes to really implementing things properly. For the time being, thank you, Alexei.
Well, it was a pleasure, as usual, being in these discussions with you, Matthias. Thanks a lot, and have a nice day. Have a nice day. Bye-bye.