The Product Experience

What is AI product management, really? Jonathan Evens (AI Product Lead, Google DeepMind)

September 16, 2026/43 min read

Jonathan Evens is an AI Product Lead at Google DeepMind, where he has spent more than a decade applying machine learning and AI across industries — from the smart grid at AutoGrid, to detecting roads and buildings from satellite imagery at Planet, to recommender systems, Google Search's AI Overviews and AI Mode, and now live avatars. He is also an advisor to the Evens Foundation, where he is building a "digital citizenry": a democracy sandbox that uses synthetic citizens to pre-test how the public might react to a policy before it is written.

Jonathan returns to The Product Experience, where hosts Lily Smith and Randy Silver pick up the conversation they started at MTPcon London, to dig further into what actually separates an AI product manager from a product manager who simply uses AI tools, why product principles have to come before evaluations, and how synthetic users can help — and mislead — at very different scales of product.

We discuss:

1. Why "AI product manager" has become a near-meaningless label, and the two distinct roles hiding underneath it: the modelling product manager working on core model capabilities, and the AI feature product manager building AI-powered products

2. Why using an LLM as a thinking partner or a coding assistant does not make someone an AI product manager — it makes them a product manager using AI tools, full stop

3. How Google Search's North Star metrics have stayed constant even as the proxy metrics beneath them — side-by-side win rates, user ratings, RLHF signals — have had to be rebuilt from scratch

4. Why product principles, not evaluations, are the real starting point for any AI feature, and how Google Search resolved the problem of trustworthy sources disagreeing on basic facts

5. How Google Search builds trust into its AI Overviews through sourcing, citation placement and UX cues such as highlighting, so users can judge at a glance what to verify

6. Where synthetic users genuinely help — cold-start problems, privacy-sensitive research, automated regression testing — and where they fall short

7. Building the Evens Foundation's "digital citizenry", and the core technical problem behind it: AI-generated personas that are less diverse and more extreme than real people

8. How team size and structure differ between a fully resourced lab like Google DeepMind and a resource-constrained non-profit team, and why Jonathan resists a single answer for the "right" team size

9. How the product manager's job is shifting as engineers absorb more of the evaluation work themselves through prompting and iteration

10. Jonathan's advice for product managers building AI features, and his case for following the Makers Manifesto

Key takeaways

  • "AI product manager" covers two distinct jobs. The modelling product manager defines and measures a model's core capabilities — factuality, reasoning, long context — and that role is concentrated almost entirely inside frontier labs. The AI feature product manager builds a product or feature on top of an existing model, and needs domain expertise and user empathy far more than technical depth. Conflating the two is why the title has become so diluted.
  • Using an LLM to think faster or write code faster does not make someone an AI product manager. It makes them a product manager using AI as part of their toolkit — the same as any other knowledge worker. The distinction matters because it clarifies what skills are actually being tested.
  • Product principles have to come before evaluations, not after. Before Jonathan starts building an eval set for a new product, he first asks what the product is meant to feel like and what values it should encode. Google Search's response to sources disagreeing on a monument's construction date, or to large language models hallucinating at scale, came from principles about trustworthiness established before any metric was built.
  • Trust in an AI feature is built through sourcing and interface design as much as through the model itself. Google Search's AI Overviews are constrained to draw only from ranked, trustworthy documents rather than the model's own memory, and users are given UX signals — citation placement, highlighting — that let them judge at a glance how much to verify.
  • Synthetic users add genuine value in cold-start scenarios, privacy-sensitive research and automated regression testing. Where they fall short is diversity: AI-generated personas tend to be less varied and more extreme than real people, which is the central technical problem behind the Evens Foundation's digital citizenry project.
  • There is no fixed answer to the right team size. Jonathan sees a gradient, from a senior developer working entirely alone, up to the Evens Foundation's single product manager with AI-assisted development skills, up to a fully staffed Google team — with the deciding factor being how unsolved the underlying problem is, not company size.
  • As engineers absorb more evaluation work themselves through prompting and iteration, roles are blending. What still sits with product management is the judgement calls that follow from product principles — deciding, for example, which technical trade-offs actually matter to the use case, rather than which are easiest to measure.

Features links

- Evens Foundation — https://evensfoundation.eu

- Makers Manifesto — https://makersmanifesto.org

- Google DeepMind — https://deepmind.google

- AutoGrid — smart grid AI company where Jonathan began applying machine learning to industry

- Planet — satellite imagery company where Jonathan worked on automated road and building detection

We're refreshing The Product Experience and want your input. Take our two-minute survey and help shape where the show goes next!

Episode transcript

Jonathan Evens 0:00

If you trained a large language model or an AI on all the data up until 10,000 years ago, right? And then you kill all the humans from 10,000 years ago, and you let it run, and then you then have the real world where we didn't have AI, and then you look at both outcomes. Is our world much more exciting and interesting than the world where AI was alone? And I think the answer is always gonna be yes.


Lily Smith 0:31


Hi Jonathan, it's so great to be chatting to you again. Last time I saw you, it was on the main stage at Mind the Product. How are you doing?


Jonathan Evens 0:39


I'm doing great, enjoying the summer, getting a tan. And uh yeah, it's been great.


Lily Smith 0:46


Well, we're gonna carry on the conversation that we kicked off when we were at the Mind the Product earlier this year. But before we do, it would be great if you could give our listeners a real quick intro to who you are and what you do in product.


Jonathan Evens 0:59


Yeah, absolutely. So I've been in product for a little over a decade. My background is is mostly in um general engineering science. And I really got into product thinking about you know what would be a job where I could use my technical skills, but not do those all day, and at the same time use some of my social skills. And then I kind of read this job description about what a product manager was, and as I read it, I just thought this is a job for me, and that's kind of how I got into it. And then yeah, over the last uh decade, I've I've really been focused on trying to apply machine learning AI to different industries. So I first did that in the uh smart grid at a company called Autogrid. Another industry I did that in was uh the geospatial industry at a company called Planets, where the idea was how do you detect buildings and roads automatically across the world with a constellation of satellites? And then more recently, I've been doing this at Google Deep Mind. Uh, first I started on Recommender Systems, and then uh I worked on Google Search, the transformation with uh AI overviews and AI mode. And now I do a few different things, uh live avatars, uh things of that nature.


Lily Smith 2:18


Awesome. And I feel like you're the, you know, when we talk about AI product managers, like you have more of a claim to this title than than most with the history that you have. But also you've launched some incredible AI features, as you mentioned, AI search at Google, but then also with the Evans Foundation as well. Do you want to kind of briefly touch on that?


Jonathan Evens 2:41


Yeah, so I would say there is, you know, as I was saying basically on on stage in in London in the Barbican, there are mature products like Google Search, where the idea was really to launch an AI feature on a on an existing product that actually was itself being transformed. So it's more an AI feature within an AI product, actually, that had already uh started the transformation. And then at the Evans Foundation, it was much more around how do you build an AI product from scratch where you don't really have resources, basically, and you kind of have to figure it out with limited resources. And there the idea was really how can we try to build tools that strengthen democracy? And so the tool that we focused on, or that I focused on on this in this talk was the idea of a digital citizenry, which would be, you know, if you think about it as the acidic sandbox where you can kind of draw up a policy and you would kind of get a pre-run of how the public would likely react, whether the the interest and the values of different individuals in society would be reflected in that policy and why they would or wouldn't agree with it. So that was that was kind of what the talk was was about. But I can also answer your question about AI product management and you know what I think it is, if that's uh of interest.


Lily Smith 4:04


Yeah, that's exactly where I was gonna go next. So let's do it.


Jonathan Evens 4:07


No, no, no worries. Um, yeah, so you know, I I would say that it's interesting that the term AI product management for me is quite weird because I see it pop up on my LinkedIn feed just all the time. And I would say up until maybe 2018, 2019, I even hesitated to put that on my LinkedIn or in my resume. And I was I remember thinking, I think there was a transition basically where you would you would actually call it machine learning product manager or modeling product manager, and then suddenly the term AI started becoming a little bit more used, but just in the industry, not not like today, obviously. And that's when I kind of always put the slash. So I put machine learning slash AI product manager, you know, from 2018, 2019 onwards. And um, yeah, I just thought it was it was interesting. And then actually the years before that, you know, there was a difference between saying you were a machine learning product manager versus a deep learning product manager, because deep learning was was really the more consequential technology. So yeah, so you know, just historically, I just find this fascinating that now everybody's talking about you know AI product management.

Lily Smith 5:24

What do you think qualifies someone to be able to say that they're an AI product manager today? Because it feels like everyone is getting familiar with and learning AI tools and the rate at which technology is changing, you know, everyone's learning. So, you know, is there a level of technical knowledge or whatever that people should be having in order to claim that they are an AI product manager?

Jonathan Evens 5:51

Yeah, I think that's that's a that's a good question. I I would say there's two types, to be honest. Um, the the first type is really what I was doing before, I would say before it was called a you know, the before the last three years, basically, uh which I would just call a modeling product manager, basically. And so when you're a modeling product manager, you're really responsible for what is the model supposed to do? How do you measure whether it's doing what it's supposed to do, and then what kind of error rates are acceptable for that model? And I think that in the the Gen AI world, for me, it really is about finding, defining, measuring these core capabilities of the core model. And what do I mean by core capabilities? I mean things like factuality, you know, whether it's handling long context very well, um, whether it's able to reason really well. And so that modeling PM role that I'm talking about, I think is really centralized in the frontier labs at this point. And so I would really, and then and then by the way, you you even have, as opposed to how it was before, because the model is is really become this complex thing, you don't have just one modeling PM. You have a modeling PM for all of these capabilities, you have a modeling PM for uh pre-training, you have a modeling PM for post-training, etc. And so so I would say that's kind of the evolution of the previous role that that I had, but in this new gen AI world where you have these highly capable models. So I think that's the I would say that is the the real claim to what I would call uh you know a modeling AI product manager. And then I think what's changed is really the fact that now any product manager can use AI as part of its toolkit. Right. Uh now the reason why everybody talks about AI product management so much is because the the tool is so powerful that it completely changes the experience. And so essentially, if you want to build your product, AI is becoming the center of it, right?

Randy Silver 8:11

Yeah, so it sounds like you're making a distinction between people who are building AI systems and then who are building uh products and features using AI. And especially on the on the latter side, you've already touched on there's machine learning, there's LLMs, there's agentic, there's any no, there's rag, there's any number of other things. And I think what most people are doing is they're just opening up, you know, Copilot or Claude or something like that, and and using it as a thinking partner or go build this for me. So when you say somebody you're someone who worked on uh on the makers manifesto and contributes to that as well. So I'm just curious when you're talking about an AI PM or an AI-enabled PM, what does that actually mean? Is it enough to just say I use an LLM to help me with my thinking or to do some of my tasks for me? Or is there something more trust? And it yeah, let's start with that.

Jonathan Evens 9:07

Yeah, no, no, absolutely. So I think there was this first category, first type I talked about, which is the modeling PM. And I think for that you need technical skills in a sense. So if you need to understand machine learning to some extent, and I think you know, the the same skills that you needed in the previous era, you still need those in the current era to be that type of PM. But again, the number of roles available are limited because it's all in the frontier labs, right? Then I think what's exploded is the basically the AI product or the AI feature PM, basically. Um, and in that case, you know, it's basically AI is is a technology and service of your product, of your application. And there, I think the superpower is really to be close to the end use case, to be close to the vertical, the domain expertise. And I think it's more important than ever, actually, and I feel like people are not talking about this enough to understand all the intricacies of the workflows of your industry, of your vertical, of your use cases, and then know how to use AI to automate or do or augment some of those workflows. I think that there's a there's a tremendous superpower there that I think people are not talking about enough. And I think that there are lots of things that don't change in that world, right? Um so for that type of AI product or AI feature PM, and that is that you you still have to design the whole experience around the model, but but with a user interface, with uh user experience, etc. And while the original model can you know can have 80% of the capability, it's maybe the last 20% that you really have to work on to get it right for your use cases. And I and then you were you were mentioning the um you know AI-powered PM or AI-enabled PM. For me, I I don't even think we should really talk about that, actually. Maybe it's uh maybe it's uh maybe it's uh I don't know. I I have a reputation of being uh pretty direct about this, but uh you know for me that's basically using AI for knowledge work, whether you do it in product management or as a financial analyst doesn't really matter.

Randy Silver 11:31

So and I think that was the question I want to get to is is that person an AI PM or that are they just a PM who is using AI as part of their tool set?

Jonathan Evens 11:39

No, they're just they're just a PM using um I would say generally they're there they're they're a PM that is using these tools to improve their productivity. That being said, there are new skills that I think are needed or can be really helpful in the the second world that I described, the AI feature, AI product, product manager, where you know, for example, vibe coding or you know, learning how to prototype pretty fast. Or another one is is, for example, if you want to so one of the core responsibilities of a product manager and one of the added values of a product manager is to bring the right data to the conversation to turn intuition into something that is backed by a little bit more evidence, right? You need the product sense, you need the intuition, but your ability to bring data to the conversation to guide that intuition is quite key. And now you kind of have a data scientist, you know, in your pocket that can just go and grab that data and do some things for you. Whereas in the past, you would have had to rely on, you know, a lot of other people, even if it won't be perfect, it'll still give you a lot of insights. And I think there's still skills there that you can apply. So those skills are part of an I would say just have a PM, actually. Just a PM.

Lily Smith 13:06

I love that we've spent like 15 minutes almost on the very first question that we had. But it's great to, it's great to have this conversation. Um, one of the things you kind of touched on there is measuring success and kind of you've obviously worked on two very different AI features into very different businesses and products. Google, which is a huge organization with like billions of customers, and then the Evans Foundation. I'm not sure how many customers you have, but I imagine it's much smaller.


Jonathan Evens 13:36


Yeah. Um early access program at the moment.


Lily Smith 13:40


Yeah. So, you know, different sort of ends of the spectrum for sure. When you're thinking about measuring success, though, I imagine you kind of approach it in a fairly similar way. And, you know, measuring success for those AI features and that the AI capability, because there's so much sort of personalization or variability in the output depending on the data that you're putting in. So there's been lots of talk before about like evaluations and you know, making sure that you can evaluate your the output of your AI capability like within the feature in the product. But as someone who has got lots of experience of this, like how have you thought about it? Like, how have you approached it in those two different situations?


Jonathan Evens 14:26


So the idea of measuring success, you would say, or yeah. Yeah. So I think I think of it in terms of metrics, right? I would say. And I would say that the thing that that hasn't really changed, and maybe I'll I'll speak in the case of Google search, is you know, what are your North Star metrics? That hasn't really changed necessarily that much, I would say. Um, so you know it's still, you know, you do you have active users that come back daily or you know, weekly or monthly. Right. It's things like does it does your product actually fulfill user needs, depending on how you measure that from an orthometric perspective? Right. So if if you're in the case of Google searches, are people coming back because their needs are fulfilled by Google search? If there were you know YouTube, for example, it might be, you know, did you find the the video that you watched satisfactory? That is, did it satisfy you to watch this video, right? Or do you regret it, for example? And then there's also, are you are you satisfied? So that's more on the Google search side, I would say. Are you satisfied with the information that you got? Is it of high quality? Do you trust it? Right, for example. So those things don't really change, I would say, the the North Star metrics. But I think what's really changed are the the proxy metrics underneath, because your your whole product is transforming, is changing. No matter whether you have an existing product or if you're starting to build a product from scratch, it's a new tech stack, it's a new way of approaching, of using the capability, right? And I would say in that case, what you're looking for are things like you know, side-by-side win rates, right? You want to basically continuously evaluate the outputs of your models and see is it better than the previous one? Is it better than this competing product? Is it better than a baseline that I've established myself, basically? Then you you also want to kind of measure the the user ratings and the user responses. Are they, you know, are they copying the answers? Are they uh following the recommendations? Are they clicking on the links? Are they giving it a thumbs up? If, for example, you're you're doing AI generated music, are they listening to the song all the way until the end, for example, right? So all these signals are extremely valuable, especially for the uh what we call the reinforcement learning uh loop or reinforcement learning from human feedback loop. And then one aspect that I find interesting that you know we hear all the time like evals is all you need. And anytime I start a new, like I work on a which has happened quite a bit in the last like year and a half, I've I've worked on quite a few different new products, let's say. And every time the thing that comes before the evaluations are what are the product principles that your evaluations are going to build on, basically. And so before even the evaluation, the real secret is what are your product principles, right? Like what is it that you're trying to create as an experience? And then the evaluations are the first translation of this into something that you can use to teach the model, basically.


Randy Silver 17:57


Um let's dig into the principles for a second, but you know, you're talking about Google and YouTube, these are established products, established areas. And I think people might be surprised to hear that the underlying principles might still not be clear on some of these things. So what was the experience? Can you be a little more specific about a time when uh things were a little confused?


Jonathan Evens 18:21


Oh, so you mean uh confusing in terms of product principles?


Randy Silver 18:25


Well, yeah, if the if the first thing to do is to get the principles aligned and nailed down, yeah. I'm assuming that means there was a time when they weren't.


Jonathan Evens 18:33


Yeah, absolutely. Well, so one example that's pretty obvious uh in the case of Google Search is why is it that people went to Google search in the first place for let's say, you know, 2000s to 2022, right? And I think that the reason people came to Google was because it had almost always, of course, not always, but it almost always had trustworthy information ranked first. And you were able to get to that information pretty quickly, right? And so there's this trustworthy aspect to the product that is quite key. And at the same time, large language models hallucinated a lot, right? So going back to the product principles, one of the questions that you know came up early on was how do you reconcile these two things? You want to basically build on this amazing capability that large language models have, which is summarizing, synthesizing information that's scattered across the web. Amazing. I don't I won't have to do this work, but at the same time, we have a massive issue if, you know, uh one out of 10 times or you know, one out of five times there is there's a problem there. And I think, you know, I think the principles there are actually machine learning principles, like machine learning product principles that we had already before, which is you have to make an an educated guess based in the case of Google Search, there's you know a lot of history behind it. So you can really look at uh a whole history to to figure out what what those numbers should be. But the the important thing is the error rate. You want to kind of know what is the error rate that is acceptable for your product to be released at what scale. And so that's really, I think, the approach that uh Google took. You know, first releasing to uh power users in in um what was called the search generative experience, and then you know, just getting getting better, you know, over time at uh this factuality problem and studying this problem and seeing examples of where the model fails. Uh, there are some examples that nobody thinks about, which is you might ask, for example, for the the date of construction of a monument. And then over the web, you'll have five trustworthy sources that will basically tell you three different dates, right? And some will overlap, they'll give you ranges and some overlap, overlapping. So, what should Google tell you in that case, right? What should be the answer? And so it's like, how do you design with product principles in mind around these nuances so that in the end the answer you know matches that principle of trustworthiness?


unknown 21:23


Right.


Lily Smith 21:23


And Jonathan, in your talk at MTP, you referenced kind of working with AI research, like doing the sort of AI research as part of working on the features both at um Google on search and also at Evans Foundation. Is that the kind of the types of things that you're looking at when you're doing AI research? Is it essentially sort of an MVP and then um working with real users to get some some proper feedback? Or is there something else to the AI research?


Jonathan Evens 21:55


I think it goes back to this separation or the different types of AI projects. Product managers that I talked about, I would say. So when you're in frontier research, what you're really trying to do is a breakthrough is often defined by whether the model is beating a benchmark. You know, that would that would show basically that you've reached state of the art on a certain capability. But the reality of large language models is that they don't, while they generalize, basic, uh, you know, if they didn't generalize, then we wouldn't have all this fuss about about large language models since you know three years or four years. Uh so they do generalize, but they don't generalize as well as it is claimed in the public sphere in the media, right? And you know, that means that there are gonna be real-world situations where you would expect them to exhibit this particular capability, but where they actually fail, right? And they don't exhibit that capability. And so I think that when you are a product manager that has access to this to this research, often you your job is to think about that model confronting reality and and and the messy reality. And then I think that like the way that you're seeing that it's that it's different is because we have the need for forward deploy engineers, basically, right? And so that's a really good that that really shows you that the AI research is is key, but at the same time, the use case and the messiness of it requires a lot of applied work. And then the other thing is like if you look at, for example, math problems, there was just actually uh some news about about math recently with with 10 problems that were that were solved, right? Um, but generally, I think with with what we see with with math problems is that yes, the AI model is able to solve math problems, but we have a lot of benchmarks with others. But at the same time, it's still struggling to help mathematicians solve problems that are the ones that they're actually focused on. And so, yeah, that's kind of how we think about it.


Randy Silver 24:04


Jonathan, you you talked a moment ago about messy reality. Yes. And I know one of the things you've experimented with in the past is you know, you're working on things that are massive in scale, and doing research at scale is difficult. It's always been difficult. So you've been playing around with synthetic users and and uh using research at that kind of scale. W talk a little bit about that. How do you avoid the the problems of smoothing things out, of you generating really reflecting the messiness of people when you're trying to use synthetic approaches?


Jonathan Evens 24:37


Yeah, yeah. So I I talked about this a little bit in the talk, um, because actually in both cases, we we have uh synthetic users, so both for uh uh Google search AI node personalization, but then also for the um synthetic citizenry at the Evans Foundation. I would say where they genuinely are helpful and add value, the the synthetic users is when you have uh cold start scenarios. So basically you have you don't have much data and you don't have a product or you don't have a feature yet, right? So you're kind of starting from scratch. And the other really important case, which I talked about in the in the talk, actually was the uh when you have privacy concerns, basically. Because actually, with synthetic users, you can kind of do whatever you want uh because they're not real, right? Um so you're not running into these privacy issues. The the other thing that's very helpful is a little bit like auto raters, which basically are large language models that help you rate or judge large language models when you train them. Um synthetic users can really help you with um automated regression testing, basically. So they can really help you understand you know whether if you change something, you've lost on a particular capability or um you're doing worse on a particular use case. And that's why they're really good for these like edge cases that you want to make sure you you know you're doing still well, basically.


Randy Silver 26:15


So are you going to do you primarily use them to help inform you and help refine things? Yeah. I guess the question is would you ever feel comfortable shipping something that had only been seen by tested with synthetic users, or do you always go to to humans as a final pass?


Jonathan Evens 26:32


I think with with trusted testers, but not trusted testers consumers. So, you know, uh B2B, um, yeah, I would I would feel fine if you know there's trust with those users and you're you're you're trying to get feedback, then yes, I wouldn't have a problem with that because I would consider I would consider the PM actually as the the first user, basically, the first human user that will really poke with it and and really QA it as much as possible before they put it in the hands of these uh real users. And then that helps you iterate faster. So and and that's actually what we're doing in the case of the the digital citizenry, actually.


unknown 27:12


Yeah.


Lily Smith 27:13


And again, kind of like talking of trust. Um that's like such a key element when you're introducing AI features, either from a B2B point of view or a B2C point of view. What's been your experience with kind of making sure that you are keeping the trust of the audience? Like, do you have any sort of tips or tricks for people to take away and consider when they're building AI into their own products?


Jonathan Evens 27:39


Yeah. So I guess I'll I'll still lean on the historical context of all of this, actually, because I just don't think that things have changed as much as as we think. So when I was at Planet and we were building a road and building change detector, uh, right? So you basically had satellites that were scanning the earth, you get all this imagery, every day you get the whole world imaged, and you so think of it as a frame, and then from day to day, right, you have new frames of the same place, right? And you're trying to detect changes, right? But then you have all sorts of real-world messiness coming into the picture, and those would be things like clouds, or you know, maybe the atmosphere is a bit different, or maybe the the shade is different coming from the buildings, you know, that are that are coming across the road, for example. So there's many, many factors that come into play. And you know, we were we were running into these problems, which was that the accuracy of the models, which was to detect these new roads and these new buildings every day or every week, right? We were trying to basically pinpoint them for mapping customers and so on. And the the issue was our customers were analysts in those companies or the these public institutions, and they wanted to know whether they could rely on the detection or not, which is very similar to what you know we're talking about now with hallucinations, right? It's just that it moved from you know this select few analysts to the general public now because everybody's using AI. And the answer was you have a product card that tells you, you know, how to deal with the product, basically, right? Like when can you really trust it? When does it kind of fail? What have we not fixed yet? And you know, they would have this card and they would be able to kind of have a better sense, and then they can use other tools to decide whether they want to you know dig deeper basically into the detection. And I think today we actually have a very similar problem, but we just have to adapt it to these, you know, to these consumer-facing products. And so that means that trust, for me, in the case of you know, Google search, it's it's really about factuality and grounding, ensuring that the the sources are, you know, that the information is coming from trusted sources. When you go to Google, you know that the AI, you know, is really taking the information from the sources. It might still hallucinate, but it will not come up with it from its parametric memory. It will it will take it from the documents. Then there is a there's a user experience aspect, which is how do you show citations? How do you kind of, as a user, are you able to quickly at, you know, with a quick glance, are you able to see can I trust this information or not? I think the UX aspect of this is very important uh as well. For example, if something is highlighted, you will have higher confidence in it, right? Being true. Whereas if it's not, you might think, oh, okay, maybe I need to verify that if I'm not sure. And then where the citations are placed and which website is it coming from is also quite important. And then the last one is when when you ask questions that are you know politically sensitive, where the answer is not necessarily super clear, it's about punting or redirecting two links or you know, saying that there are different views on this topic, which is in itself also a problem, by the way, because in some cases, you know, you might uh you know give a measured response, but actually one of them is not the truth, and one of them is the truth. So there's still work to be done, but I think it's with this approach that we can we can get there.


Randy Silver 31:26


Jonathan, so you you've had these vast range of experiences in how you build products. And that the Evans Foundation, it's small. I know you don't even have the full-time dev working with you on on this uh on this initiative. And at Google, you have, you know, at least from the outside, infinite resources, uh, infinite tokens, infinite people, we're at least incredible people in team sizes. But you've also been doing this for a number of years, and the way that tools come in, you know, we've got some people theorizing that the ideal version of a product development team is going down to a one-person team. Other people think it's requires a certain number of people for friction and discussion. Other people are keeping the the two pizza team, but increasing the size of the responsibility for the team. I'm just curious a little bit about your experience with this between Google and the Evans Foundation. What is a team look like these days? What is the difference in your approach to building? How where do you think this is all going?


Jonathan Evens 32:25


Yeah, I I think that's a good question. I'm not sure I have the, I don't think I have a definite answer, to be honest. And so maybe I'm just gonna be, you know, in French you say when you don't answer yes or no, you're you're from Normandy. It's uh that's the expression. So I'm kind of gonna give uh an answer from Normandy. Uh, but I actually think it's it's probably the truth. Uh that that would, I mean, that's probably that would be my bet, is that you'll have kind of a gradient, to be honest. I think there are some problems where you can build it all on your own. And if you're a developer already, and you or I should say a senior developer, so you you have more than you know five years of experience, and you know, you've you've worked in production systems before and so on. I think that you can get very far all by yourself, to be honest, on some problem. I think if you're trying to do the digital citizenry that we're doing at the Evans Foundation, I think it's gonna be a bit harder for you to do it all on your own. Um, and that's what we thought. And that's be because because you don't just need the the dev skills, you also need a principled approach. You need to think about what the product is gonna look like, you need AI research because it's an unsolved problem. Um, so there's there's a lot more iteration going on, but from our side, from the foundation's perspective, basically, so Matt, who is who is basically our our our tech and democracy lead, he's de facto acting as a as a PM with some you know some some claud magic dev skills, right? Or other large language models, and you know he's he's been able to do quite a bit, and that means that we only ask for AI research when we need the help because we've identified problems that we cannot solve ourselves.


Randy Silver 34:21


At a larger company, let's go stick with the the Google type thing. At larger companies, where do you think this is gonna go? For you said you you've been working on a range of initiatives over the last year. Have you seen a change in the way that team size is being approached or team makeup?


Jonathan Evens 34:37


Yeah, so the interesting thing I saw over the last, let's say, you know, two, three years is really the move from what was traditionally machine learning engineering, which is hyperparameter tuning, and yeah, basically, you know, really being a model, a model engineer, right? To a lot can be done with prompting. A lot is about you know, writing these iPython notebooks and you know, and just going through, you know, iterating on your prompt and going through the outputs and building the eval set and seeing where you're where you're going with that. And that's where I do agree that the the roles really blend basically, because engineers will handle you know an important part of the of the evaluations, but not all of the parts of the evaluations. And that's when we go back to this product principle approach, because the product principle approach is true product management, and that has to influence the evaluations, right? So maybe I I'll try to make it concrete. If I'm trying to build a generative media product, which you know, I won't get into the details of you know what it does, but let's say it's visual, you know, it's it's video based and so on. And some of the evaluations are going to be about, you know, how realistic is you know, is this part of the video, right? How how like is the image coherent, right? So those are things that are a bit more technical and that engineers can can really, I would say, handle. But then there are other parts which are more related to this product sense, which you know, which would be about what I'm just trying to think about something without, you know, without sharing too much confidential stuff, but it would be it would be more about uh, for example, you know, what kind of movements should we prioritize, right? Because those movements are key to those types of industries and can really create a difference for their use cases. And so that's where I think the PM product sense and product principles comes into play. Anyways, and and so going back to your question about the the team size, I I think there's a lot of overlap, and I think I think there's just gonna be a lot more prototyping going on. That's what I think.


Randy Silver 37:04


I'm looking forward to coming back to this this interview in a year and saying, oh, that's what he was talking about.


Lily Smith 37:12


Um Jonathan, this has been so good. Thank you so much for um sharing your time with us. I know you have a a household full of screaming children behind you, so I'm it's super impressed that you have not been distracted by them. Um they're probably all AI product managers in the making well.


Jonathan Evens 37:32


Machine learning product managers.


Lily Smith 37:33


Machine learning, yes, that's right. So I wasn't paying attention. I guess one more kind of quick question before we go. Uh, you know, for everyone sort of listening to this, what would be your ask of them when they're looking at uh building AI features into their products or you know, considering how they use AI in the future?


Jonathan Evens 37:56


Uh to my to my kids.


Lily Smith 37:59


Oh, no, not to the kids, but to the audience. But now I want to know what you would ask the kids too.


Jonathan Evens 38:05


Okay, maybe maybe I'll answer the first part to my kids and then to the to the audience. Um I think for my kids, I would say I think the two most important skills are critical thinking and um being able to adapt and learn quickly, I would say. Those are the two things that I think will not change. I mean, unless you think that the the jagged frontier of artificial general intelligence is going to eventually you know supersede us in in everything, which you know, in my personal opinion, is still an open question, to be honest. Um, I don't think anyone has the answer. Some people have believed. So I'd say those two things are are still um critical.


Lily Smith 38:49


And then another Normandy answer for that one then.


Jonathan Evens 38:54


Well, I I I I would be more on the I guess I I think there's more to I mean anyways, we we're not gonna get into the you know the whole philosophical debate, but maybe maybe I'll just I'll just I'll just state it like this, which is something I said actually two, three years ago, which is true intelligence is if you train the large language model or an AI on all the data up until you know 10,000 years ago, right? And you said it was super intelligent, and then you kill all the humans, right, from 10,000 years ago, and you let it run, and then you then have the real world where we didn't have AI, and then you look at both outcomes. Is our world much more exciting and interesting than the world where AI was alone? And I think the answer is always gonna be yes. That's what I think. So, in that sense, I think we'll never reach true complete superintelligence. That is the philosophical part of it.


Randy Silver 39:52


I'm just glad we're not gonna actually run that experiment.


Jonathan Evens 39:56


No, we're not, we're not, we're not. Um so so yeah, so and then for the for the audience, uh, could you could you remind me of the question again?


Lily Smith 40:05


Sorry, because we got into the philosophical Yeah, so you know, if they're if they're kind of working on AI features or you know, trying to develop their skills as a as a product manager that is having to work with AI in their products, like what's the sort of asking them? Like, what should they be thinking about considering to ensure that they're they're doing a good job, basically?


Jonathan Evens 40:27


Yeah, so I I would say continue to have user empathy and really focus on the these deep user problems and uh or deep user pain points, and to start with the problem, I would say if you're gonna be the the second type that I mentioned, which is the AI feature or AI product PM, I think this is essential. And that's basically what I you know was kind of doing from behind on the you know the Google search um AI mode personalization, for example. And I talked about this quite quite a bit in the talk, but I think it's it's I think it's really key. I think the flip side of this is you're gonna build a gimmick feature, and I don't want to do trash talk, but we're all gonna have to uh you know live with uh a meta-ai logo on on WhatsApp and never know why we should what to do with it and just be bothered by it. Uh so we just have to be, you know, so I think don't build that, you know, build something that you know is is actually solving your pain points, I would say. Um then I think we shouldn't forget about how like basically building a great product is about how elegantly does it really solve a real human problem and how does it solve, how does it eliminate the friction? I think that doesn't change. It's still the case. And the last one is you mentioned the maker manifesto. I would say, I think I would I would just follow the makers manifesto. I think the principles are are quite correct, you know, purpose uh over over possibility, um, you know, the human accountability over full automation. These kinds of principles I think are quite key, and a lot of people have been thinking about those. So I would follow the makers manifesto.


unknown 42:09


Yeah.


Lily Smith 42:10


Awesome. Uh Jonathan, thank you so much for joining us. I really, really appreciate it.


Jonathan Evens 42:16


My pleasure.


Lily Smith 42:17


The product experience hosts are me, Lily Smith, host by night, and chief product officer by day.


Randy Silver 42:24


And me, Randy Silver, also host by night. And I spend my days working with product and leadership teams, helping their teams to do amazing work.


Lily Smith 42:33


Luran Pratt is our producer, and Luke Smith is our editor.