Does it work? Yeah.
Test, test. Okay. All right. Okay.
Hello, everyone. Welcome to our lesson learned session here today. Happy that you're all here. We will talk about the Mercedes project where we migrated their on-prem legacy PAM environment towards the cloud. I'm Erik Siebler. I'm the Digital Identity Capability Lead of DXC and with me is Nico here. Hi. I'm Nico. I'm with Mercedes-Benz since 2019 and took over the part for the PAM service within Mercedes-Benz now two years and also was leading the migration from PAM on-prem to the SaaS environment with all the technical challenges and yeah. Yeah. All right.
And in the agenda today, we wanted to talk about two topics. One is the migration reason and strategy, so a little bit where we were coming from and where we were heading and then of course the lessons learned and we hopefully have still enough time for the Q&A. And on the right-hand side, you also see this little nice Puma over there, which was the internal project name and logo. That's why we have put it on here as well from elevating from on-prem to the cloud. Yeah. Okay.
So, let us start with the reason and the strategy and to understand this a little bit, we will dive back a little bit to show you where Mercedes started their PAM journey. They had their first go live back in 2017, so nearly 10 years back. Over the time, this whole environment has grown quite rapidly.
So, we are talking about more than 9,000 active administrators, over 200,000 accounts, so it's a fairly large environment which scales around the globe. However, the underlying tech stack which has been used has never had any major upgrades, so it was an old version which hasn't kept up with all the challenges which have arised over the last 10 years.
So, we came to the situation that certain use cases and issues were not fully covered anymore. So, when you're looking to the standard use case, as we all use PAM, an administrator connects through a PAM towards resources, most often infrastructure, and that works pretty well. That is where all this legacy PAM tools are working perfectly. You have a native RDP, Linux experience, all is fine.
However, when you're looking towards things like web admin portals, cloud SaaS, you see certain limitations there, which comes to the situation that administrators ask to retrieve a password. So, they're retrieving a password and connect directly to an admin portal, for example. Can be an Azure portal, a SaaS application, whatever you're using, which is then not perfect anymore because you basically lose all the threat intelligence and all the versions. You basically use the whole PAM solution like a really expensive password manager, and that is not what you wanted to achieve.
So, that was one of the underlying pillars that we said we wanted to achieve a more flexible solution. So, we said, all right, we come from an on-prem centric architecture towards the future, which must be flexible and need to adapt to where the use cases are coming and what we wanted to do.
So, on the left-hand side, you see a high-level overview of the old environment, which was fully on-prem, several instances around the globe, and the administrators worked on their instances to achieve access to mainly infrastructure. That's where they started and where they were going. A little bit simplified, there was not much SaaS or cloud connected and fairly limited and not as native as possible.
So, the decision was the new solution must be more flexible and must be able to do more. Then it was clear it needs to go to the cloud to leverage all the features there.
So, the new architecture looks like this, that we have a PAM cloud environment, which now integrates to SaaS, cloud, and still, of course, all the on-prem infrastructure which is there. With now having something in the cloud, which was fully in their control in the past, we of course needed to think about certain disaster situations, since PAM is also rolled out in the production sites of Mercedes. And you can all imagine if cars cannot be produced anymore, that's a big issue.
So, there are certain disaster recovery instances, which allow to connect towards the systems even if the connection to the internet or the cloud or to whatever reason might be not there. So, that is the fallback, which is there to make sure that the operability of the service is there. That's the core. People need to connect to it. It's a critical service. And thinking about this, also the migration is quite a heavy topic. The whole idea was we can have zero downtime for the migration.
So, we needed to move everything to the cloud, but nobody should have any downtime. And this we achieved through, now it's working, through certain strategy ways, which I wanted to touch point on. Because when we talk about over 9,000 administrators, over 200,000 accounts, there was no possibility to do a big bang. That was fairly not possible.
So, we decided to go in certain ways. So, we had the left-hand side, the on-prem world. And on the right-hand side, we had the cloud world, which we wanted to get to.
So, we decided to do a first wave. The first wave were key stakeholders. It was the IAM department and certain key personnel from various business units within their Mercedes-Benz, with various use cases, so that we have good technical contacts throughout the estate of Mercedes who can work with us on the first wave.
With them, we tested all of it. We did certain functionality things.
So, we had a really good first, really small round. What we did then, we stopped for them for a certain time, the password rotation of their accounts in the old environment. And the next step, we started to migrate all their administrative accounts, all the servers or infrastructure which they get access to, and their access rights, so who can access what for this small amount of users.
And then, the good thing of this was, they started to use the new world. In case that they detected any issues with it, the old world was still there and fully operational.
So, in case something was not working as expected, the people were able to work even more still on the old world. Both were there at the same time, fully operational. That allowed us to be without zero downtime during the whole migration, mainly.
And then, after a certain time, when this was there, we started the password rotation on the new platform, and they were fully moved to the cloud. But then, we only had the first couple of people, and now we needed to go in the next way.
So, we decided, or we thought at first that we will do a regional approach, and through different regions, Europe, the Americas, and APJ, but then we decided to do it a little bit different. And we basically took the majority of administrators in a big chunk, and that was the heavy load.
However, as usual, with so many things, you might not get all of them in the first round. So, we still kept the third wave, which was all the remaining things which kept over the wave two.
And then, we did exactly the same procedure again and again. They had old world, new world, at the same time, we were able to check everything is working and move them then also to the modern way.
So, when you're now looking to the final things which we did, we did also a DNS change. So, if someone still wanted to use the old world, it was still there, and they were automatically forwarded to the cloud, so that this is also disappeared, and then we finally did the decommissioning back in March. And in total, this was roughly a project time of 11 months. We did the whole migration of all the users successfully without any major incidents, but of course, there were several war rooms on the way. You can imagine a project this size that will not work without this, and many lessons learned.
And for the lessons learned, the second part now, I will hand over to Nico. Perfect.
Yeah, the lessons learned from the migration from on-prem to SaaS was very difficult because it was a completely new environment for us. It was somewhere in the cloud, no more our environment, no more our infrastructure. But for us, we said, okay, the one-to-one migration makes just sense for us because we know our users, we know our devices, our servers, databases, and so on. Not at all, all the use cases, normal use cases like SSH, RDP, database connection, we know them. But sometimes, maybe also in your environment, the people use the panel solution for maybe some other workarounds.
For that, we said, okay, this could be very tough for us, but we said, okay, we have to make an announcement in the company. We have to share the news that we will move to a new solution to the SaaS environment. And for that, we bring the people in the session to say, okay, what do you do with the panel solution? Do you have some workarounds, or can we just go through the migration with you guys? What was very difficult in the last month for us was the network, because we have data centers all over the world. We have some clouds within the Mercedes network. We are still on-prem.
We have edge data centers. We have the car manufacturer also. And for all the connections, we have to ensure that all the people can reach their devices, that they can work as they know it from the on-prem solution. For the on-prem solution, everything was working, and they expect that it's the same in the SaaS.
For that, it was pretty much easy. We say we do some network tests to compare what is connectable from the on-prem with the SaaS, and then see where is it not working with that.
We said, okay, we have to go with our infrastructure team to solve the issues over the firewalls to troubleshoot it. What was also very interesting, how to connect the SaaS environment, you have only two ways. From the solution was a network connector, which goes just through the public internet, but also BGP was one of the solutions.
What we say, this is how it will work, because public internet, it feels sometimes a little bit difficult when you say you want to secure all your company data, all your connections, and then you go just through public internet, and we say, okay, it's more feasible for us when we just use the BGP technology to connect the SaaS environment. And also within our RFI, we just say, okay, this is one of the leading points where the supplier or the PAM solution must fit to us that we can get the connection through it.
And also what was interesting in the discussion was like, hey, is it like just a central solution in Frankfurt, for example, in the SaaS environment, or like multi-geo, because we have also some data centers and customers all over the world. And what I learned also the last month, all of these users, like in APAC, they are feeling a little bit confused when the latency is not working, and that's why we said, okay, we need multi-instances, so all over the world that the performance is also pretty much good for the customers.
And SaaS is also, as you know, when SaaS is no more there yet, this could be happening also. What we learned also in the past from AWS outages and so on, and for us as a company, for us it's also important that the manufacturer is still able to produce the cars, and for us was DR one of the leading points where we have to say, okay, we have to discuss this also. You're laughing now. We have one DR on-prem solution that if something happens that the guys and the teams from the car manufacturer can still connect to their target devices, that the car manufacturer can still produce it.
And also something like what happens if the IDP, for example, goes down, for that we say, okay, we have a break-last scenario that we have like local users, you can, when it's enabled, you can log in, you can still use the connection as you did it when it's working that the people can just use the PAM solution. And sometimes there happens also sometimes some confusing stuff because some people had some issues with driving mapping and also some copy-paste issues.
For that we just said, okay, we make like a hyper-care with the supplier that we can solve that in the migration phase, and this was very one of the, I think the most benefits that we say, okay, we have all the technical-enabled people together to solve the issue that it's working for the customers. Yeah, the human migration, this was, I think, also very tough for us because technically it's pretty much easy, I think so, to give the people the facts about what we want to do or what we want to achieve also.
But as you know, when you're moving from your tool, what you're using for years, and now you have a completely different tool, which is looking and feeling completely different, the people say, stop, I don't like this tool, I don't work with this. For that, we just said, okay, we have to do some training sessions with the people. We have to share some documentations, also training materials like videos and so on, that the people can see, okay, how the new solution is working. Because a lot of people said it's too complicated and so on, but no, it's not complicated, it's just different.
And for that, you have to prepare and give the people some stuff that they can work with it. Incidents. We thought that there are more incidents appears during the migration, and we were surprised when it's lower than 1% from all the PAM users, also the concurrent sessions, when we all compare that active sessions with the incidents, it was very, very quiet for that, and also some incidents was valid, because they had some technical issues, and some are not valid, it was something like the people didn't read the manual for how they use the PAM solution to get the connection to their machines.
So yeah, it was a little bit quiet for that. And for that was also very important, decommission was for us very simple, because we had a final date, because the licenses was nearly to expire, and for that we have to say, okay, we have a deadline to migrate all the people, that we have also the pressure on all the customers and the leaders and so on, that they will also support us, that we can do this big challenge to migrate the PAM on-prem to the SAS environment. And for that, we just say, okay, we have some spare time for that, but we also say, this is the latest date.
This is, I think, what the people helps us to say, okay, we have here a target, we have to focus on it, and also to split what is necessary to solve the issues until that, what we have to do. And for all other stuff, we can say, okay, it's necessary, but not for the migration itself, we can do it also afterward for that.
Yeah, any Q&A for the migration from on-prem to the SAS environment? Any questions from the audience so far?
Okay, let me get your microphone. Hi, can you say a few more words on the topic of disaster recovery and the local or the on-prem solution you have? How is the synchronization? Yeah.
Just, if you say, I want to say a few words. Yes.
So, in the RFI, we said, okay, we just invited some suppliers and asked, hey, how do you solve this topic? Because as a car manufacturer, we have to, that the service is still working. And they said, we have some app solutions, for example, when people can install the app on their mobile phone, and then they can, it's a synchronized one, and they can just view and copy the password. But as you know, also in the car manufacturer itself, some people don't have a mobile phone.
So for us, that wasn't working for us. So we just said, okay, we need like a on-prem solution, what we can just install on our own environment.
And yeah, it must be fulfilling all the core requirements for PAM. And so we just said, okay, we install the physical machines, install the application on it, doing the same procedure, like that the data are synchronized, and also that it's the connection is working. So we are doing also regularly connection tests also, and also some tests, is the disaster recovery really working or not? And then we can also see if there's an emergent cases that the DR is really working for us. Okay. We have a question for our online attendees. One question is, did you change PAM vendors?
If not, did you evaluate other vendors? You want to answer that? Yeah. So we changed the PAM vendor. Unfortunately, due to corporate policy, we are not allowed to name them openly on a recorded session. Yeah. So possibly you can pick it up with us afterwards. But Nico and us, we went through, and some people here in the room might notice, we went through assessment from the existing with their modern solution together with many others on the market. And out of this, after RFP, the one solution then got picked.
And at the end, it was not necessarily, but at the end, it went also with a vendor change. So I said exact names we can do in one-on-ones, but not on record. I think what is important here is just that you compare all your suppliers. Can they fit all your requirements? Can provide all the stable connection also? I think this is more the leading. But at the end, the vendor, it's not necessary. Are you staying on the same or moving to another one? It's more about, is the technology working? Okay. Then thank you both for your insights. It was quite information, full of information. Thanks a lot.
If you have any questions, feel free to reach out to them. I guess you will be browsing around then. Thank you.