Platform Engineering for Private Cloud
Given at NDC Oslo on .
Recording
Abstract
A platform-engineering talk with a private-cloud bias. Covers what platforms actually contain (everything beyond namespaces and templates), the cost and dev-to-ops ratios of running one well, and the seven ways teams fail at building one.
Slides
Further Resources
- Platform Marketing In 2025 20250312 (PDF)
- Developer Toil: The Hidden Tech Debt (white paper)
- Platform Engineering resources
- CNCF Platforms White Paper - the vendor-neutral reference architecture used in the talk.
- CNCF Platform Engineering Maturity Model - the planning document recommended at the end.
Transcript
Timecodes link into the recording on YouTube at that moment.
Platform engineering, for private cloud
(0:10) Hello, thanks for coming to my talk. This is an exciting fisheye view of a large hall, which will be thrilling. As my talk says, I'm going to go over platform engineering for private cloud - one of the more thrilling titles I've ever used in a presentation. What I want to cover is not only platform engineering itself, but some of the things that are different when you apply it on your dedicated environment, your own infrastructure. Now, what exactly private cloud is, is fluid, but I generally think of it as: primarily, you're running it on premises, on your own infrastructure. And to some extent it could also be if you're treating cloud infrastructure as - and this sounds pejorative - let's call it dumb infrastructure. Something that you build your own stack on top of and run your own thing, instead of using something that is very close to developers.
(1:11) It's very popular nowadays, and there's an evolution from all of these things. What's great about being in the line of work that I'm in is there really is this continuous evolution of trying to finally crack making infrastructure easier for developers to use. Thinking about how we make it easier for them to deal with infrastructure directly, or, as we like to say, write more code and think about their application more. Hopefully this time it'll stick. This is definitely the most characters we've used since IT service management in the 2000s to describe this endeavor. So maybe, instead of shortening it like we've done in the past, once we finally have a long descriptive word, it'll hold. But we'll see - we like to go in cycles for things.
(2:04) So why private cloud? I've been interested over my career in where workloads are, what the movement towards various types of infrastructure is. It's good to check in every now and then - depending on how feisty you're feeling, every year or every two years - and ask the question not so much what kind of excitement and revenue is around the various places, mostly public cloud, that you could move to, but where are the workloads? Where are the applications?
(2:55) This is pretty important for platform engineering, because - unless you've figured out some strange exotic use case - if you don't have applications, all the platform engineering you're going to do is not really important or worth anything. The point of platform engineering is to support applications that you're running. Sure, you want to have things that are secure. It should be performant. You probably want some database services. I always think it's important to remind ourselves of the number one requirement: it should work. But those are things that any sort of engineering or approach could give you, or would purport to.
(3:26) So I'm going to spend the first part talking about that location of workloads, because I don't think it's really discussed very much, and the implications of what you need to do for platform engineering to be successful and useful are always tied to this place that they're running. It's kind of a - if you remember - the medium is the message. Or the medium is the infrastructure. I'm going to work on that phrase a little bit more, because that didn't work out so well.
(3:53) A brief introduction to myself. I've been lucky to work at Pivotal, and then that was VMware, and now it's the Tanzu division by Broadcom, as I recall, for 10 years. In that role I've spent most of my time talking with large organizations about how they get better at software - their custom-written software. And I've gotten to write several little books. Before that I was a software developer. I worked on cloud M&A at Dell - you remember when OpenStack was a big deal and everyone was going to build a public cloud out of OpenStack? There's fun stories there. And I've been an analyst at a place called RedMonk for a while and then 451 Research. And I was a programmer way before that, long ago. I host a couple of podcasts. As my friend Andrew Shafer used to say: engage with my brand.
Where are the apps?
(4:48) So let's get into that thing: where are the applications? And again, to emphasize this, knowing where these applications are is going to drive some significant ways in how you think about doing platform engineering. Finding this out is surprisingly hard. There are estimates locked up behind analyst paywalls, one of which I'll show you here. But if you spend a bit of time looking at the experts on where they say applications are, you can start to get some idea and get to that number I had up there - that it's about 50%. Maybe it's 40/60. I don't know, we'll see.
(5:29) First of all, back in January, we could go to some people who are the best experts at this: the people at Amazon. This is a quote from the AWS CEO, and it is scoped, as you can read, down to large organizations - which are the ones I'm most interested in, and the ones where platform engineering probably has the widest relevance, and where this idea of platform engineering and private cloud certainly has the most relevance. I try not to do math in public, but if you do this math here, I think that amounts to 90 to 60% of workloads in large organizations being on private cloud - or whatever you want to call it, dedicated infrastructure, on premises. They're running it, or outsourcing it to managed service providers, but it's not cloudy infrastructure. That starts to give you an idea, with a crazy big range. This isn't an average, it's just various customers.
(6:33) Here's another one. If you read IDC, like I do - being a former analyst, I have maybe a different relationship with analysts, I think their work is pretty good. If you read them year over year you get a sense of the momentum and who you can trust, who's making up surveys and who's not. I think IDC is quite reliable here. This is some recent findings from their Cloud Pulse report, a survey they do about every year. Unfortunately they get very upset if you use their data from 18 months or older. I'm not sure what that says about its validity 18 months ago. But you can trust me - remember, I wrote all those books, I have a smiling face, I do podcasts, so I'm trustworthy. What you would see is that these numbers more or less have been creeping around the same level for about 10 years or so. I don't know, maybe eight years. I lose track of time as I get older. And this is why I keep saying dedicated environment: they're very prickly about not saying private cloud. It's dedicated environment. Which sure sounds great. But if you add these up, you see that it's just under 45% dedicated environment - although they slip into "dedicated cloud" down there, so maybe we should send them a note. Again, you get this sense that it's around half and half, around 40/60, something like that.
(8:06) There are all sorts of other figures, depending on what kind of charts you like. There are some from Barclays and Goldman Sachs that are particularly good if you really want to get deep into looking at where these applications are. But when you look over all of them, you kind of arrive - or at least I arrive - at this conclusion. And especially, as the AWS CEO was saying, when you look at larger organizations, this holds a whole lot more. Thankfully I have a philosophy degree, so I'm very comfortable with ambiguity and approximations. These numbers make me happy. I don't really need much more precision than that to figure out the area we're working in.
(8:49) So the point being: there is a significant amount of applications running on premises, running on private cloud, running in a dedicated environment. What this starts to mean is that there is, let's say, organizational ownership and responsibility of the stack, in both good and bad ways. So when you come in and you want to build a platform - a platform being to support applications, to support developers - we'll get into what that entails as far as different things that team or those teams need to do, versus if they were just supporting a full stack that they didn't necessarily own and have responsibility for.
What is a platform?
(9:30) With that, let's define what a platform is. It's always good, when you're talking about something, to define it yourself, so you can say "well, that's not my definition of something" when you're objecting to it.
(9:45) First of all, if you haven't been following platforms and platform as a product for a while, there's a great talk that was just given and then published by Paula Kennedy, who I used to work with. You can see that the timeline of platforms is quite long - if you want to throw in platform as a service, if you remember that term, with places like Heroku and others way back when, it goes back further. But this is a really delightful overview of the idea of what a platform is, and you can see the long timeline it's gone through, and a lot of funny subtle jokes about how we keep repeating ourselves over and over again. I realize you can't click on that now, but perhaps later. Maybe there are some talks you're not interested in, you get worn out, you want to enjoy a croissant or a hearty sandwich. It's a good 25 minutes to spend if you're really into platforms.
(10:44) Any talk on platform engineering is required to begin by quoting Evan Bottcher's 2018 definition of a platform, which I have here. So I'll move on.
(10:56) Recently, in the past three or so years, there was a resurgence in - and kind of a creation of - the category of platform engineering. Arguably the good folks at Humanitec kicked this off when they were taking their internal developer portal and they wanted to do some market category creation. Which is great: you want to create a market category that perfectly fits the thing that you have to sell. And if you remember, they declared that DevOps was dead, and that kicked off two years of fantastic conversation. But I think we've arrived at this notion - as I was joking with that Dune picture at the beginning - that platform engineering is what we call all of this stuff. How infrastructure people, operations people, are trying to support developers as much as possible.
(11:38) Now, like I said, I lose track of time, so let's say about two and a half years ago the first thing that emerged was that the P in internal developer platform actually stood for portal. This was when Spotify open sourced Backstage. It was this great internal portal, if you will, for developers - hence an IDP. There was a lot of interest in it, and there still is, and a lot of thinking about it as - this is going to be not quite accurate - the UI for Kubernetes, the UI for a stack that you have, mixed in with the project management that goes with making what various teams are doing inside a company legible, knowable, a lot more standardized.
(12:41) Now, very quickly, as happens often, the analysts swooped in, did some great diagrams, and what happened is the notion of a platform expanded to be more of the infrastructure, going further down than just that portal on top. And then a memo was sent out to all of us in the platform world that the P now changed to platform, from portal. Portal is just a part of what a platform is.
(13:09) Here are two diagrams of what a platform is. The one on your right, of course, is the best platform you will ever encounter. You should definitely be using it and insisting that you have it, because I work there and I like to pay my bills and profit from that. Thank you for sending me here, Tanzu division by Broadcom.
(13:29) Now, if you're not into vendor pitches and you want a vendor-neutral way of thinking about the world, which is always great, the CNCF, the Cloud Native Computing Foundation, has a great reference architecture that they came out with a number of years ago. Bridging to the portal part: you can see that the portal is just one part of that big old - whether you think of it as a cake or a very complex tasty burger, depending on how hungry you are and what you like to eat.
(14:08) As you look through all of those things - maybe it's not too insightful or revelatory - what you see is that the components in there are all oriented around what people whose hair color is getting slightly lighter and whiter, like me, would call middleware. Maybe that's a terrible word. You could call it services if you like, components, frameworks, whatever word you enjoy using that makes you feel younger and more zesty instead of something like middleware. But you can see you've got databases, event queues. There are some things that are a little subtle to see in there: how you package and deploy your applications - we used to call it a release train, I think, then it was a secure software supply chain. All sorts of phrases and words, if you like these things. But there's also a little chunk of infrastructure stuff included in there as well.
(14:56) I think the theory of this - I guess it's the practice of this - is: the more embedded you have all of these things, the more you can integrate all of these things, the easier it becomes for developers. Because the more automated and standardized you can make them, the more self-service you have, and hopefully the less your application developers have to spend time thinking about how they package up this application, how they express the dependencies, or a multi-regional deployment of things like that. And especially not wanting them to invent those things.
(15:30) The third interesting thing about this to have in your head is that, even coming from the home of everyone's favorite Kubernetes, the CNCF, Kubernetes doesn't really show up here. And that's - at least in this diagram, nothing against Kubernetes - more to say that the infrastructure is not really part of the platform. That's not the important part of it. Infrastructure is some other thing, down there in the purple layer. I don't like to criticize people's work, but I'm not really into those Easter-color pastel things, I like darker colors. But that purple Easter egg down there, that's where the infrastructure lives.
(16:11) So that's another thing that you want with a platform: more or less infrastructure agnosticism. Now, is that perfectly possible all the time? We sure would like it to be, every five or ten years. To bring back another phrase from history's past: we want multicloud and portability, we want to avoid lock-in. It'd be great to hear from anyone who's been consistently successful with that over the past few years. That would be a breakthrough paper. Maybe we can go to IT Revolution Press, put together a working group, and have a five- or ten-page PDF about how we did it - multicloud application portability. It only took us two days to move to a completely different stack of infrastructure.
What does a platform team do?
(16:56) So what does a platform team do? Here is a good roles-and-responsibilities division of things. There are a lot more things, but this is a nice summary. You've got your application teams, developers as people would say. You've got your platform teams, and then more that your platform teams do along this spectrum. Your developers are there writing their code, doing the things they're responsible for, hooked up to the tools that you have, the pipelines, all of those things. And the sense you should get from what the platform team is doing is that they're curating and developing the services that are in a platform. They're making sure that it's integrated together. They're taking care of all those thrilling meetings with networking and security people and auditors. They're doing a lot of the work to say this platform not only makes the application developers' work easier - I don't know about their lives.
(18:00) This is a fun thing about AI. You should be asking: I'd love to use AI, and if I'm 25% more productive, does that mean I can go home early and maybe work four days a week? Because really, if you're more productive and you don't get anything from it, you're getting kind of a raw deal. So make sure you pay attention to what your productivity increases are, and be like, now I'd like to not do more work, or get compensated for it somehow. Anyhow, where was I?
(18:27) Yes. So what you see these platform teams doing is kind of upward, making it easier for developers. But a lot of the hidden work they do is making sure that all of the downward-facing things - and there's nothing pejorative about the directions, you could reorient it to 45 degree angles, or if I knew my nautical terms, port and starboard or something - you can see that they're doing a lot of the work to fit it into the rest of the organization.
(18:52) And there is, I think, a little unknown - I haven't heard a lot of discussion of this - about who actually does what I would call the management, the monitoring, the troubleshooting, the root cause analysis, for things up and down the stack. I think the platform engineering team does a fair amount of that. But at some point you probably need to open up some phone call, or create some channel depending on if you're a Slack, Mattermost, Google Chat or Teams person, and talk to your friends the network engineers and the storage people to troubleshoot what's actually happening. Maybe even the developers, if they care to show up at that time. This might be one of those free times they have because they're 25% more productive, enabling you to be 120% productive when they don't show up.
Is Kubernetes a platform?
(19:40) A logical question. Is it - how do you say that guy's name, law or something? Whenever there's a headline with a question, the answer is no. But let me explain why, because I think there has been a lot of positioning, if you will, which is to say thinking of Kubernetes as a developer platform. I know that because I've paid attention over the past 10 years, and I've also talked to numerous operations and infrastructure teams at large organizations, and we have the same conversation. The brief version: they say, these application developers - and then we talk about how great they are, and special snowflakes for a while, little gun shooting, having fun there - they asked us to stand up Kubernetes, and we did, and they're not using it. We don't see them going to it, so what are we doing wrong? I think that is because there's this perception that that's what you would do: that it's this complete developer platform that's fully ready to use.
(20:41) But you see this: this is a survey that we used to do annually. As Kubernetes has spread more and more - that's why there are five years of questions here - as there have been more people using it, you see the benefits of Kubernetes dropping in the survey base. That's what you would expect as a new technology rolls out: new people are using it, they're figuring it out, figuring out how to use it in their organization. But the one highlighted here, to give you a quick summary: the top green one is the most recent year, which is 2024, and the bottom blue one is the first year. You see one of the developer-facing metrics - the ability to deliver code faster, or shorten software development cycles - has been dropping over those past five years. It'd be fun to do a survey this year and see what the result is. A little weird is that you see that similar kind of movement for a lot of the areas there. Again, I think that is largely an effect of: you have a new technology, you've got to figure it out. But I think it's particularly distressing, disheartening, unfortunate for people who think about Kubernetes as the full application developer solution, the thing you would put in place for a platform.
(22:06) And indeed, despite some weirdness in the Kubernetes world over the years, I think the original Kubernetes people have been telling us this for a long time. Here's a summary of the quotes from several years ago. Craig McLuckie used to work in my organization - or I guess technically I used to work in his organization - at VMware, when Heptio was purchased. And you can see this almost like: hey, sorry about all that Kubernetes stuff, it really wasn't our intention that you, the developers - not Google people, not public cloud provider people - would actually be exposed to that.
(22:46) There's a great two-part documentary on Kubernetes, and if you go back and watch it - which you should do after that Paula one, I'm going to assign some more homework later, so you might need a new sheet of paper - it's really a phenomenal look into strategy. Because you realize that the whole strategy of Kubernetes was basically to slow down Amazon. I'll let you look into that, but one thing you should notice, if you pay attention to these things, is they interview everyone in the industry except Amazon. And Amazon is portrayed as a continuous looming, kind of Darth Vader-y drone shot of one of their buildings, which is the only way they show up. Which should tell you a lot about the relationship of Kubernetes, originally, and the strategy that they were going after.
How do you run a platform, in private cloud?
(23:42) So we know that Kubernetes - or at least I believe, and therefore know. Going back to my liberal arts degree, what's the difference between believing and knowing? Effectively they're more or less the same thing, until you stop believing something, I guess, but that's a talk for another time. So how do you run a platform? We know a platform is all that exciting goop on top of the infrastructure, including services that developers might want to use, ways of supporting it, that sort of thing. So what does it look like in private cloud? Well, the first important thing, and why that talk by Paula is a great reminder, is this product part of it. Maybe one day we'll change the P to product - internal developer product - which would be fun. This point you've probably heard a zillion times if you've ever been to a platform engineering thing, but it needs to be emphasized over and over again, because I don't think it's actually practiced that much. Saying that you should treat your platform as a product is kind of like the idea of standing in a stand-up meeting, where everyone obviously sits down the whole time. It's subverting the spirit that's right there in the wording of it.
(24:43) So why is it important to have product management, to have at least one role that is actually a product manager? They can go get their Pragmatic product management certification. Maybe they can just read a stack of books to figure out what it is. But a product manager is what you have when you have a product. Of course it means something different when you're selling toothpaste and cereal, more or less. But in the IT world - having worked mostly in the vendor world for lots of my career - you get very familiar with a product manager, and application developers generally know this idea.
(25:20) But the important thing, other than selecting the features that you have, guiding what this product looks like, is something even further up the concept chain of a product manager: in order to have a product, it's great to have customers. If you don't have customers, as a product, this is the equivalent of having a platform without applications.
(25:42) And so then it's good to ask: who are our customers? As you can imagine, as Thomas Müller at Mercedes-Benz said a couple of years ago - "we are building this platform not for us, we are building it for Mercedes-Benz developers." If you are running a platform, you're the product manager for the platform. You think about the developers as your customers. These are the first people in line. Hopefully they're the people you're doing things for. Sure, the platform needs to run. That would be cool. It needs to be secure. You've got to have all that stuff. But that's like saying it would be cool if Earth was hospitable. Those are basics that you need to have.
(26:12) But the product manager is constantly going through and understanding what the developers are doing, what they do great, what they like, what they dislike, and thinking about how we can change this feature, how we can make it better, how we can add new services to it - that would make our customers happier, that would keep them using our product.
(27:00) If you're more of a developer, you obviously have a notion of how your day-to-day work could be happier. I remember when I was an application developer, we could up our productivity by stopping complaining all the time, because it was a huge amount of the work, so to speak, that we did. But what I advise to people who have more of an infrastructure background - and even if you're from a developer standpoint going down the stack - is that it's good to start asking the developers what they need. The way you do that is either you talk with them directly, which is fantastic if you can do it. A lot of these product managers, especially at large organizations, will travel around. You can imagine at large organizations they have a campus, a hub, whatever you want to call it, various geographic locations, and going around and talking with them - even sitting with them as they do their work - is very valuable. But you can also do things like, to use a word everyone loves, survey them. Ask them questions.
(27:44) So what questions should you be asking them? Here's a set that you can use. There's a white paper I worked on with a couple of people who do this kind of platform stuff, that goes over the theory behind this and some of these questions. It's something you can start with. You can use whatever survey thing you have, put it on a one-to-five Likert scale, and then the first time you do it you can crudely - not crudely as in impolitely, but just the basics of it - stick it in an Excel spreadsheet, and if you know the magic of sort-by-ascending, you can start to get an idea of what the most important problems are.
(28:28) That initial pass is great product management thinking, but what's even better is when you survey them again in three months or six months. You want to see if things have improved. Have some of those high-priority things moved down? Are people happier about things? And if not, that's feedback that you need to improve the things you already shipped, that things could be better. Again, thinking about developers as customers, you're interested in them being happier, I guess. There are some other metrics you can be interested in, like shipping time and stuff like that. But you're constantly - constantly is the wrong word, because people don't like to take a lot of surveys. Although I often say, going back to the 80%-true joke: if you were to offer developers "would you like to take a survey," or if you said "would you like to tell me everything that I'm doing that you think is wrong," they probably would be happy to talk about the second thing. Maybe developers have gotten to be more optimistic and happy since I did it way back when, but probably not.
(29:30) Anyhow, figure out some way of getting feedback from your customers - not just initially, but ongoing - and thinking about how you use that to guide what's in the platform. Product management, just like any product manager would, just like anyone who has a product that they want to, if you will, sell or drive usage of amongst their customer base.
Internal marketing and community management
(29:52) Now this next one I think is something that's especially important in private cloud, in large organizations, and that is marketing. A term I'm very comfortable with, because in my developer relations and whatever role, I've worked in marketing over the years. But if that phrase makes you unhappy, just ease into it - you can also call it community management if you like that better.
(30:24) What you're doing with this internal marketing: think about it, you stand up a platform in a large organization, or you already have one. You're already probably competing with, I don't know, three, four - depending on how big you are, 10 - different platforms. Plus any developer teams you have who are just like, I don't know why I need a platform, that sounds like a weekend's worth of work. Also that DHH guy said that you can just do it, so no problem, right? Building your own platform. Now, why are they going to work on the weekends? Maybe that's what they do with their 25% improvement. I'm not sure, but it seems like you should not do that. Maybe you should say, I could spend a week doing it instead of a weekend.
(31:09) Anyhow, you're competing with a lot of other platforms. Again, thinking about your platform as a product, you're in a market, if you will - an ecosystem, if you don't like that word. What that means is you need to talk about why your platform is good. There's some standard marketing stuff: you want to get your positioning, your messaging, your value propositions, or as us marketing people like to say, value props. Who can be bothered to say two words?
(31:40) But part of that is not only narrowing down what your pitch is - or, if you prefer, having really good documentation going over what it is - but also thinking about core community management or marketing things like branding. I see this over and over again in all sorts of large organizations: one of the first things they do is come up with a t-shirt, as my friend Deshan would joke. And when you come up with a t-shirt, one, you've proven that you can get budget, which is thrilling. But two, you come up with a snappy phrase, a logo, a color scheme, an identity.
(32:20) As developers, we're like: I'm immune to marketing, doesn't work on me. And I love using my Apple Mac. I will defend to the death the type of code editor that I use. I can tell you why the shell I use is extremely important and you're wrong. So forth and so on. Obviously we're affected by identity and marketing and brand. Sure, it's more about the functionality, not the actual identity that you have. But the same is true for any technology, in any platform. So you're trying to define the type of person that someone is, or becomes, or can be - the identity - by thinking about what your marketing is.
(32:41) And again, you see this across the board with companies. A great example is the US Air Force - I think they still do this - they launched this platform and this whole initiative to get better at software, and they called it Kessel Run, because of course they're going to make a reference to Han Solo and Star Wars. That idea of, we're kind of like renegades - not a word you would like your military to use - but we're doing things differently, we're a little more savvy, and we're going to identify with this thing that probably a lot of the people involved really like as well. And then we can also make this t-shirt that has a double meaning. There are other things that AF could stand for instead of Air Force, if you remember those days. But this kind of brand, this kind of internal marketing, helps bring that in to a great degree.
(33:33) Now I joked about calling marketing community management, but genuine community management is also a huge part of what the platform engineering team does. Setting up internal chat channels. Having quarterly internal conferences where you bring application teams that have used your platform to talk with other developers - which, by the way, is a great marketing technique called customer references, or word of mouth. Having people tell you they were successful with your product, your platform, on more of a peer-to-peer basis, rather than having someone who has no idea what they're talking about saying that it's great: the platform engineers being those people versus the application developers. I have this really distressing, ribald idea about application developers. I should probably go revisit that at some point.
(34:32) But you can see that having this - what I would call a developer relations effort - starts to become important for a platform team in a large organization, in a private cloud environment. Indeed, if you look at large organizations like JP Morgan Chase, big gigantic global bank, last I checked they have a team of about seven or eight people internally who are developer advocates for the handful of platforms that they have. And they do all of these activities. They go around. It's not only training, it's not only awareness building.
(35:04) But it also becomes a channel. This is the part about developer relations that all DevRel people aspire to - very few of them get the chance to do it. It's also a feedback channel going back to that product management. So if your internal community people are out there talking with your customers, your developers, they can bring that feedback back to what the product manager, the platform, is doing. They can start to give some advice on how to prioritize features.
Continuous financial metrics
(35:28) So now, I think maybe this is perhaps the most important thing that is different about platform engineering in private cloud, in large organizations, and that is - let's call it continuous financial metrics. The way finance, which is to say oxygen, works in an organization, particularly large organizations, is you're on a 12- or 18-month budgeting cycle. It's supposed to be 12 months, annual, but essentially every year you have to go through an excruciatingly delightful process of saying, this is the budget that we need, proven by the return on investment we had last year and our projected return on investment in years forward. You're basically saying, the reason we need 5 to 10 million is because we're awesome. And they're like, yes, I'm sure you're awesome, but I'm going to need some spreadsheets. What this amounts to is metrics. And these metrics are associated with some aspect of helping the business out. Unless you're working at a tech company - which is great if you are - that metric is usually not oriented around raw growth, because we're looking for a large startup valuation and then an exit, and then actually making money is someone else's problem.
(36:49) So you need to start thinking about how, from day one, you are not only tracking what these metrics are, but very tactically and wisely saying what the metrics are, framing them. Because in my experience, knowing the value of it, if you will, is extremely difficult unless you're in something that's very close to the business's business transactions - like retail, or other things where you can directly connect a line of code changing to people booking more airplane flights or something.
(37:22) But things get a little more obscure. I like to go back to banking. What's the return on investment of me being able to see the balance in my bank account? It's theoretically infinite, because if I couldn't, I wouldn't bank there. But you can't really put infinity in a spreadsheet. That doesn't really work. So you need to start coming up with metrics that map to why the platform is valuable, that allow you to talk upward to your organization - again, I don't mean to be pejorative - to talk in terms of finance and business and that planning.
(37:56) There are all sorts of metrics you can come up with. Here's a default set that's been evolved over, I don't know, the past five or eight years. It's even formatted in a dense way that I think people in large organizations can do. Have you ever noticed they're very good at reading documents in landscape? That's kind of the mindset you have with the rest of the organization. If you want to print something, go up to that little thing - you really want to print it in portrait because you've read about the Amazon memos, all the people complaining about how documenting something is better. Just select landscape, you're good to go, and things will be fine.
(38:37) You can see a lot of these things are capturing what a platform is trying to do. Everyone wants to speed up. Now, speed can be gratuitous, but the point is that the faster your organization - I'll just say business, but you can imagine how this maps to a nonprofit-oriented organization, whether you're helping citizens or otherwise - the faster you can change your software and get that software in front of users, the more that business can evolve. So speed is an interesting indication of something. I guess I could reference the DORA reports - the DevOps one, not the European banking one - to show that that was important as well.
(39:26) A lot of the rest of these are based on that kind of table stakes I was going over. But it's also important to emphasize and track to the rest of the organization why the table stakes are good, and why the investment in your platform is paying off, and to represent that in some way. A lot of these middle ones I think of as indoor plumbing. The value that we put on indoor plumbing is basically zero until it stops working, and then it has infinite value. We really enjoy it when it turns back on after not being on for a while. Thinking about how you represent that so that you don't have to remind people with a catastrophic event is a good number to start tracking.
(40:05) Now, of course, savings is good, and that gets into some weird stuff depending on how you think about things and what the situation is. You've got to be careful if you're trying to get money. It's good to track savings of how much you could have spent versus how much you are spending, because you don't really want to be in a situation where you're like, "and guess what, we need less budget than we had last year, we're saving you money." So think about how it's more of a good investment. Savings is up there because the other ones are S's, so you can call those the five S's. You don't want to be the four S's and the I, which would actually be kind of fun. But think about how you represent the money made, the money you're attached to, the kind of savings you have versus theoretical alternatives, and start tracking those things.
(40:55) Now, for developers, kind of inward facing, there are all sorts of other metrics you can track. Here's a default set as well. This is good for people who aren't so concerned about your budgeting, but they're good things to start measuring and tracking.
(41:09) And I think the point with both of these metric things is, again, thinking of your platform as a product: a product needs to have measurements, needs to have metrics. And I know that once you start measuring something, that becomes the goal of doing stuff. Hopefully now that you know that, you won't fall into that mistake. I always think that's a weird paradox, that people should know that's the issue. So you need to constantly be thinking about, are we just doing the metrics or not? Good luck. But it's good to have some indicators, some dashboards if you will, of how things are going and how you can talk to other people through spreadsheets - metaphorically and literally speaking.
Scaling the platform: pairing and seeding
(41:47) Just a couple more things, thinking about how you spread platform usage through your organization. Again, this is particularly relevant to large organizations, and large organizations use a lot of on premises or private cloud, so it's particularly relevant to that sphere of things. What I've observed, and many organizations have gone through over the years, is this kind of slow-at-first spread of a platform, and then a very rapid spread of it. I'm sure there are some math terms I could use there, but see my credentials earlier when it comes to mathematics - or maths, as people like to say.
(42:27) What you see at the beginning of the platform is selecting an initial development team, maybe even two development teams, and working extremely closely with them. They're your first initial customers. As you're planning and developing the platform, you want to talk with them to figure out what their needs are and build it around them. You're getting good feedback about stuff, and you're also starting small so that your learning function - which is to say, when you fail - is as risk-managed as possible. You don't want to do everything at once. If you're thinking, in the first year we're going to move 120 to 500 application teams to the new platform, because that's the only way we can make the business case - you should probably look around the rest of your organization for alternate jobs that you could have after that doesn't work out. But you see more of this slight rise in development teams on the platform.
(43:26) Now, another side effect that's interesting, that goes back to community management, is that as you build up these initial teams and as they're successful, they become advocates or customer references for you, that you can then use in your efforts to spread the goodness, the great cheer, of your platform. So you're building up the case in these peers that other developers will probably listen to more than weird infrastructure people who have a strange title - aren't they basically just DevOps people, people may be thinking - who are more of a trusted source for spreading the idea that a platform is good. And if you're very lucky, you can even take some of those initial people who are interested in helping others, who have that more senior, principal-engineering mindset of "my job is to help the rest of the organization," and start to seed them in other teams. So you're taking direct knowledge to new teams of how to use the platform, and so on. But planning out that gradual-and-then-sudden ascent becomes very important.
(44:35) And then, getting into more of, let's say, planning - planning out the activities that you'll do, what it's going to look like, even getting into budget, what you might even call strategy if you're not too pedantic about the definition of strategy in a corporate context - there's another great CNCF working group, the platform maturity model crew. That's not their name. But they've been working on a platform maturity model which I think is great. They're one of the few people who have managed to cram a maturity model down into, I think, less than 50 pages - I think it's maybe about 25 or 30. And I hear there's a new revision coming out. This is well worth your time reading if you're doing platform engineering, or especially planning it in a larger organization, because it's the best version of what a maturity model is. Another word you may not like, but a maturity model is sort of like: here's a bunch of mistakes people in the past have made and how you can avoid them, and what you can expect to happen over the course of things, how you can organize the activities and the stages of what you're doing. It's a lot longer title than maturity model, but maybe it makes you more comfortable.
(45:46) But if you read through this, some of the key things you'll see are that you can expect it to take more or less three to five years to go through this entire cycle - to have, let's say, undeniable success with your platform. Which means, to use the technical term, a lot of applications running on it. If you've got a big platform and there's only a few applications running on it, it's kind of questionable what the overall value of that is. There's an important point that all maturity models make, which is that maybe you don't need to go through all five steps. You might reach step two and that's sufficient.
(46:28) There are all sorts of interesting results from the maturity model, and in a good way it's self-reinforcing about a lot of the notions I went over and didn't go over about what a platform engineering team does. In particular, it's paying attention to the developers. One of the steps in the maturity model is: do you have a product manager? Are you getting feedback from the developers? You'll get an idea of what those are and be able to plan this out a lot better.
What about AI?
(46:59) So finally, no talk nowadays is complete without talking about AI. So how am I going to fit that in there? That's a great question. I'll tell you. Here we are talking about large organizations and private cloud. What is relevant to platform engineering with AI is a couple of things. One, the notion of large organizations, where as we saw the CEO of AWS say, a lot of it is private cloud, somewhere between 90 and 60% depending on which organization you talk with.
(47:38) You also survey those organizations and you see that their intention - at least, we'll see how it pans out - is to run a lot of their AI stack on premises. Which makes sense, because of data and wanting to control things. I think it's a little early to get too settled on where the workloads will be, but that's the early feedback you get: the intentions at the moment are to have a lot of it in a dedicated environment, and about 50% of it.
(48:04) Now, what that means is that the way platform engineers - the way I and some of my coworkers and other people in the community are seeing them start to think about how to support AI in their platforms. And by this I don't mean the assistants or copilots coding in your IDE. That's a whole other sort of thing. Here I mean adding actual AI functionality, actual features, to your applications, to how your business is running.
(48:37) So what I think is emerging is that the platform engineers get all enamored of AI. They start freaking out about whether it's going to replace their job, as we all do. And then they actually start using it day to day and they're like, oh, this is fine, it'll be great. And they start to think about that AI functionality as services that they're providing, as middleware.
(49:11) Once you appreciate the magic of what all these AI things do, and the mind-blowing awesomeness, and then you come back down to reality - maybe you've had a huge lunch so you can think more clearly - what you start to think about is: oh, I'm just providing another piece of middleware, a service to developers. How does this fit into the workflow that they have, which I know because I've talked with them, because I'm a great platform engineer, product manager type? What can I do when they need to use this and fold it into the way they're doing application development? It's probably similar to when they use a database or they need to do some sort of messaging thing.
(49:34) And that's where having a platform is going to be a huge advantage, because you'll have so much of that scaffolding, so much of that stuff built in already. You're already providing access to services. You can start to think about how you add this new type of service, how you add access to all sorts of models, all sorts of supporting frameworks - whether you want to run those models on your own, whether you want to broker access from your own models to publicly hosted ones, whatever it may be. If you've been doing platform engineering for even a short amount of time, you're used to dealing with that kind of thing with the other services that you have. It's not that it's not going to be a problem - every service has its own unique cool things that you get to experience - but at least you're going to have done a lot of the conceptual work around it.
(50:18) So, thanks for coming to my talk. It's always a pleasure to talk with people and go over this. If you're interested in clicking on the links in the slides, there they are. And it's always fun to hear people's experience with platform engineering, or other things - it's not like I only get super happy to talk about platform engineering, there are other things that are fun as well. But it's always good to hear from people. Enjoy the rest of the conference. Thank you.