Transcript

From Individual AI to Institutional AI

Event held on Sep 8–10th, 2026
Disclaimer: This transcript was created using AI
  • Julia Nimchinski:

    Thank you so much, Ishan, and thank you, Erik, and welcome to the show, Kady and Wade. How’s your… Beginning of the year.

    Wade Foster:

    I think it’s September already.

    Kady Srinivasan:

    It’s hard.

    Julia Nimchinski:

    How are you doing, Wade?

    Wade Foster:

    Good.

    Julia Nimchinski:

    Yeah.

    Kady Srinivasan:

    Hi, Wade, good to see you.

    Wade Foster:

    Good to see you too, Kady.

    Kady Srinivasan:

    It’s Kady, by the way, guys. Julie, I think I started with… it’s all good. Nice to see you, on this hard skill exchange, Wade. I am super excited for this Agentic Harness, topic that we are talking about. So I know we have only about 30 minutes, so maybe I thought, we could jump right in. I don’t think you need introduction. I think most people know who you are. So, if you don’t mind, I’m gonna jump in, because you have so many thoughts to share. I just want to get everybody, to hear them. So, let’s start with this, which is.

    I think Andreessen recently came out with this, whole article about intelligence and how it’s commoditizing super fast, and their thesis is that the whole durable mode kind of shifts to whoever owns that workflow, context, and distribution. And I think you’ve… your, point of view is kind of pretty much the same thing, right? That… that, that whoever owns the workflow and the context is the one that’s going to win in the future. So, where are you seeing companies kind of getting this backwards, locking themselves into one LLM’s model and Harness instead of kind of building that flexible layer.

    Wade Foster:

    Yeah, I think this is where benchmarks are really helpful at illustrating this point. So, you know, you can use any benchmark, but I’ll maybe just share my screen on, you know, one that we have at Zapier called Automation Bench. This is just a public benchmark, and what it does is it assesses all these models on particular tasks, and in our case, they’re like traditional automation workflows. So here’s an example. We just closed the Meridian Core platform deal. Market as one, and route the win notice to the right team, per our routing policy. Confirm the account tier from the account hierarchy spreadsheet, convert currencies if needed. and check for any open support escalations.

    So, inside this benchmark, it has you know, 600 of these tasks, and it’s, spec’d out across sales, marketing, operations, support, finance, HR, etc. And if you go look at the leaderboard, you can see how, how it shakes out. And so you can see, even the best model is only scoring at about 42%. And these models are changing all the time. So, GPT-6, this was just released last week. Before that, you had Cloud Fable 5.1, which was released, like, a week and a half ago. Then you had Gemini, you know, here, which is a much cheaper model, so, you know, still performs well, but it does at a fraction of the cost.

    But you could just sort of see, like, how fast… The model capabilities are changing, and also how fast they’re improving over time. And so what this tells you is that if you’re a company and you’re saying, hey, we’re a… you know, a Claude shop, or we’re a ChatGPT shop, or we’re a Gemini shop. You’re leaving a lot of, like, improvements on the table as these models get better and better. They’re also getting cheaper and cheaper, but the way in which they’re getting better, the way in which they’re getting cheaper is also changing. You know, you can go look at, for example, the by-domain ones.

    And, you know, the HR one is, for whatever reason, seems to be lagging behind. The models are not quite as good at this stuff. And so what you really want to be able to have inside your company is a harness, a toolkit that allows you to interchange whatever is state-of-the-art for you. And that’s, like, where I think a lot of companies are going wrong, is they’re sort of saying… they’re hooking their wagon to one particular solution instead of leaving their stack flexible, so that they can swap in capabilities as the market starts to offer new and newer stuff.

    Kady Srinivasan:

    Yeah, and you’ve talked about this before, too, this kind of the smartest companies you’ve seen, they have a Harness that allows them to optimize workflows for the jobs they have, right? Can you tell us a little bit more about what, in your opinion, has… what is working for the companies that are doing this well versus the ones that are not?

    Wade Foster:

    Yeah, so I think the things that we see the best companies doing is that they’re measuring success on these workflows. So they’re paying attention to what actually matters. So, this is, again, where the benchmark’s gonna help us. So if we reference back to Automation Bench, you can look across, that best model was only completing 41% of the tasks. So that still is quite low. So if the way you’re thinking about rolling out AI or rolling out automation inside your organization, you do need to be stepping back and saying. Where am I gonna use, you know, old-fashioned deterministic workflows or code?

    Where am I gonna use an agent? Where am I gonna have a human in the loop? And you have to start to measure the success rate of those things. And as you start to combine the capabilities, that’s what starts to give you higher success rate. And you start to see, okay, our completion rate on this particular workflow, or this particular task. can get a lot higher. And you can also measure the cost associated with it, you can measure the speed to complete it with it, and so the best companies are identifying those workflows, and they’re starting to build out this, like, it’s this, like, build, iteration, learn loop that helps them continually optimize this stuff.

    And as they get better and better, they’re learning how to hand off more and more to an agent or to a workflow, and take humans out of the loop, or use humans only in the places that are, like, most, crucial.

  • Kady Srinivasan:

    We’ll come back to the human in the loop, because I know you have a very, very interesting example from Zapier, with the email automation piece, but before we do that, when we think about an Agentic Harness, do you believe that it needs to be a horizontal one that goes across all of the domains within a company, or is it more of a vertical one? Or is that going to be a blend based on what an enterprise is? How do you think about that?

    Wade Foster:

    I suspect we’re gonna see a bit of a blend here. You know, I think… you know, you probably need to have, like, a set of shared capabilities that will be somewhat horizontal. This will be things like your permission structure, your context layer, your model routing, your observability, all these sorts of tools are probably Gonna be owned centrally.

    Kady Srinivasan:

    Hmm.

    Wade Foster:

    But then… whatever, like, you know, your marketing team, or your sales team, or your IT team, or your product development team needs, for a particular workflow. that’s where I think you’re gonna start to see a lot more divergence, and people will say, hey, this harness, this harness and tools, this harness and tools and XYZ other things can combine to get us a better outcome than we’re getting elsewhere. And so I do think that on the edges, you’re gonna see teams have a lot of freedom. to optimize for what they care about. But at the end of the day, you still want a shared set of building blocks, otherwise you’re gonna have a tough time, You know, servicing all the different needs inside of a company really well.

    Kady Srinivasan:

    Right, and that probably also brings us to that kind of question of context management, and who actually building the context, I feel, is starting to become more and more of a discipline or a capability within a company, and who gets to own that? Especially if it’s more of a horizontal Agentic Harness. Is it the IT team and enterprise IT team? Is it, business functions that is actually adding to the whole thing? How do you think about who needs to own the idea of maintaining the context graph?

    Wade Foster:

    Well… I don’t know that I’ve come across, like, a universal answer to this question yet.

    Kady Srinivasan:

    Yeah.

    Wade Foster:

    What I can say is how we’re going about it inside of Zapier. We have tiered our context into a couple different layers. So we have company context. We have team context, and we have individual context. Now, obviously, individual context, that’s curated by the individual themselves. Like, I maintain my own personal context, you know, it’s kind of a second brain for me, you know. as does many other folks inside the company. Now, at the company layer, that is me and a handful of support staff that are really helping me maintain the company-level, you know, critical information.

    This might be, like, your company strategy, your ICP, you know, if you have values, or, certain strategy docs, things like that. those are sort of curated and maintained there. Then at the team level, you know, each team sort of maintains their own set of context. And in each case. you’re often not asking humans alone to maintain this. In fact, I would say probably at Zapier, the vast majority of that context is kept up to date at this point in time by workflows and agents who are helping us perpetually keep that stuff, you know, sort of fresh, because one of the real challenges that you can, run into with this context system is stale context, you know, out-of-date context.

    Kady Srinivasan:

    Bye.

    Wade Foster:

    You know, if that sort of leaks into your prompts, or leaks into your workflows, you start to get wrong answers, or it’s referencing the, you know, like, something that might have been correct 3 months ago, but isn’t… it’s not actually right today, because companies are dynamic, they’re ever-evolving.

  • Kady Srinivasan:

    Yeah, I think that’s also an interesting one in terms of when we think about updating the context on a regular basis, you also run into this idea of tokenization, right? How much… how many tokens do you end up burning? Especially if what you’re trying to update is based on this deterministic workflow type of a thing. So how do you balance How much deterministic versus probabilistic workflows a company should have, and where does that kind of balance start?

    Wade Foster:

    Well, you know, I think for us, the way I… The thing that makes the most sense to me is you really want to put reasoning inside the rails. versus trying to rein in a free-range agent. And, you know, the best way you can do that is start to break down, like, the workflow. And oftentimes. You know, a workflow has, you know, much more, like, deterministic steps in it than you probably are thinking of. So, you know, in a lot of cases, it kicks off via, like, a trigger. So it’s like some event happened in the company.

    Maybe you got a new lead, or maybe, like, a project kicked off, or, you know, maybe someone churned. Like, there’s some event out there that kicks off, like, a critical workflow. And after that workflow kicks off, then there’s a set of context gathering. And so it’s like, okay, we need to go fetch this data from here, fetch data from there. None of those things are agents. Those are, like, you know, API calls, those are, like, things that you want hyperscripted, because you want them to work the same way every single time, you want them to run low cost, etc.

    Then, somewhere in the middle, is where you start to need reasoning. You need judgment. And that’s where you start to say, okay, we’re gonna kick off an agent, or we’re gonna kick off, like, an AI eval stat that’s gonna take all this context, take all this stuff come in, and it’s gonna, you know, have a set of rules, it’s gonna have a set of guidance, it’s gonna have some goals, and now you want to reason over that stuff. And after it does all that sort of… that work, then you’re saying, okay, I want you to kick all that stuff back out, and then go start to take action.

    And when you take action, you don’t want the agent to go, you know, sort of guess at what those actions are, you want it to assign those out to code, or to a workflow, and say, okay, your job is to go update the CRM, your job is to go, you know, message somebody in Slack, your job is to go send the email to the customer, etc. And, you know, maybe you’re putting a human in the loop somewhere in the process here, where you’re saying, hey, I want the human to evaluate the system and say, yes, that passes muster, or no, it actually doesn’t, we need to go evaluate, you know, the process upstream.

    And so… you know, I think that this, like, way of thinking allows you to make sure that, you know, whatever workflows you have in place tend to run way more accurately, they tend to run at a order of magnitude lower cost. And you have a lot more control over what the tool is actually doing at the end of the day. And this is, I think, for a lot of the mission critical, or the production use cases inside of an organization, this is kind of the process you want to go through. It might be different for sort of, like, low-stakes personal things.

    Where you’re like, I’m kind of okay for it to just, like, have an agent run, and it’s, like, a little easier to do, and, you know, the odds of it, like, the cost of making a mistake is low, or the odds of it, like, burning tons of tokens and running up my bill are, like, also relatively low. So you kind of have to think about the situation you’re in.

    Kady Srinivasan:

    Yeah, we… we did something similar on the marketing side at Freshworks, tried to create, like, a little context, graph of marketing brain for certain… certain aspects, and… but exactly to your point, we’ve had to figure out what is deterministic, what can… what should yield the exact same thing every single time that we run, and then what is probabilistic, what is agentic. And then learn from it. And on that note, I know when we talked earlier. You had walked me through this example of autonomous email for support, feature that you released, and the whole journey of you initially had human in the loop, and then you finally went to the other side of the spectrum, where you could take the human in the loop out of it, because then you’ve optimized that system.

    So, walk us through that whole thing. How did it work? What were you trying to optimize for? All this sort of good stuff.

    Wade Foster:

    Yeah, sure. So, let me actually… I have an old, or a somewhat recent slide I can show that sort of walks through this use case. So, like many companies, Zapier does email support for a lot of our customers. And one of the things we were asking was, hey. You know, what percentage of our, email could actually be handled by an agent? And this process, you know, started by us providing tools to our frontline support, where they could use, like, Zapier MCP to go look up information and, you know, act as an assistant for them to do faster support.

    But over time, we started to say, well, why don’t we actually see if we can invert their roles? Where the agent can actually be the one that does the support, and the human sort of assists in sort of helping it get to the finish line. And you can see here, you know, kind of one of the more, this is actually, like, the V-1 version of this, so, like, one version ago of this workflow, basically worked the following ways. One, you would get a ticket that would come in. And we would sort of… we had some sampling, where we were selecting from a certain group of tickets that we felt were, the right ones to run this experiment on.

    And then that would pass the context over to 5 parallel agents. So we had 5 independent agents across 5 different models kick off and start to do internal investigation, external evidence, and come up with a diagnosis. And once they all, finalized their diagnosis, it would go to this consensus gate. And this is an agent, and the agent would look at the 5 outputs, and if it saw that 4 of the 5 agreed on the diagnosis, it would pass this gate. And then, you know, we would go, look at some other checks and balances. So, you know, is there fabrication coming in here?

    Is there bad advice? Is there over-promising? Is there internal leaks? Just other things that sort of would help us minimize the risk. And then there would be human quality gate. And so this is where a human would come in and say. you know, yes, I’m gonna approve this, I think the agent did a great job, let’s go ahead and send it, or I’m gonna reject it, and I’m gonna supply a reason back to it. And the reason I’m gonna supply… the reason you supply this reason is that this goes and feeds back in to the overall architecture of the system, so that now, when these investigations kick off in the future, they have new evidence that tells them, like, oh, I gotta watch for this.

    And, you know, we ran this loop quite a bit, and, you know, you can see we’re now up to, I don’t know, almost a little over a thousand drafts a month. 75% of them are passing this human quality gate, approved. And the result is that the resolution time for these tickets is 27% faster. So we’re starting to get, like, a lot faster support kicking in here. And over time, we have started to build confidence where, in some cases, for certain eligible tickets, we’re comfortable removing the human quality 8, and that’s only because for a slice of these tickets, the approved unchanged number got really high.

    It was starting to kick into, like, 98, 99%, where we’re saying, well, you know, let’s go see what happens if we run those fully automated. And so this process, I think, is what you need to do for certain workflows, is to start breaking it down, understand, like, where do you have a human in the loop, what’s running deterministically, what’s not. You know, one of the things I’m not showing here is the cost of the system. You know, we’re measuring the cost of all of these things. You know, I think the… the consensus is, of course, that, like, well, AI is cheaper than humans, and yeah, I think that’s probably still true, but the costs are a lot closer than you’d think in some cases.

    And so you just need to have, like, a lot of data to manage, like, these more sophisticated systems that you’ve got running in your business.

    Kady Srinivasan:

    Yeah, amazing. It’s so good when you can show something and that comes to life, right? But it’s also one of those things where the UI looks so amazing and simple, but I’m sure there’s a lot of encoding of different things that happen in the background.

    Wade Foster:

    Well, I mean, most of this stuff is happening in the background, right? These are all just, like, little, you know, automations or little agents running, you know, on a computer in the cloud somewhere.

  • Kady Srinivasan:

    Right, right. The one thing… the other thing that kind of stuck with me when we spoke was how you talked about, marketing as a discipline. You have to be able to encode your judgment and your taste. and make it… bring it down into, like, principles that you can actually put in a prompt or whatever. Tell us a little bit more about how did that come to life in the example that you shared, or in general, what is your advice?

    Wade Foster:

    Yeah, well, that’s a great question. So, in this email example, you know, for any of you who have worked with these models and tried to get them to write, you probably know how bad they are at writing. They fall into these, you know, AI truisms, over and over again. You know, Claude will say, you know, this is the, you know, the sharp point all the time, or this is the honest truth, or this is the load-bearing thing. Like, it just has this, like, vernacular That it defaults to over and over and over again, and so it’s really hard to get them to write In a way that is honestly tolerable to a lot of readers.

    And, you know, in marketing, and certainly in our case, you know, we wanted it to write in the voice of tone of our company, like, the way we care about. We don’t want it to fall back on these, like, you know, AI idiosyncrasies. And, you know, I remember as we were working through this project, getting to a checkpoint and reviewing some of the outputs and kind of being blown away and thinking, like, wow, this doesn’t sound like any of the AI writing I’m used to seeing. And I had been building skills for myself to try and get these things to write well, and I’d made some headway, but I still found myself kind of constantly having to, like. add new rules, add new things, and still just not getting it to dialed in real quick.

    And so I asked the team, I was like, can you show me what you did here? And, turns out, there is a 200-line, like, prompt, or markdown file, effectively, that describes the voice and tone of a Zapier support email. And it is hyper-specific. You know, this is what the high name part should look like. This is what the first sentence in the email should look like, this is what the troubleshooting sentence in the email should look like. This is what the sign-off should look like. 200 lines just going into incredible depth. And that’s the level of work that was required to get this thing to write in a way that just didn’t sound so, like it was AI slop.

    And I think this is the exercise that marketers really, have to think through, is that it is very difficult to get these tools to sound like your company brand. And a lot of folks and companies don’t actually take the time to step back and say, this is what we sound like, this is what we want to go do, and do it in a level of detail that actually gets the models to, reflect that back out in their outputs. So I would encourage folks to put the effort in on this. It probably requires more, more than you think, I think is what I have learned.

    And, you know, maybe at some point in time, like. The models will get so good that you can sort of give it a couple front sentences, and they’ll sort of be off to the races, but that’s not been my experience with the latest crop of models.

    Kady Srinivasan:

    Yeah, yeah, completely. We, we did something similar, too. It took us about 75 different iterations to be able to come to a set of instructions that we would feel comfortable with, but even then, we had to have human in the loop to be able to massage all of the messaging, so it’s not an easy plug-and-play at this point. Not if you want to sound… you don’t want to sound like AI slop. We’re, I know we are coming up on time. What is the smallest Workflow that you would say to go-to-market leaders to start with, if they were to start to go on this journey of transformation and other things?

    Wade Foster:

    Yeah, I, you know… I think your question is right, like, find the smallest thing. I think what I see a lot of people make the mistake on is they’re… they’ll try and, you know, be like, how do I automate this entire thing, start to finish? And often, that’s just… you start by boiling the ocean. A simple use case that I have seen a lot of folks start with is meeting prep. you know, your sales team is meeting with prospects all the time, so build a simple agent that preps them. The reason this is a good one is because it’s fairly low stakes.

    You know, you can build meeting prep, you can hand it to an internal person, your sales rep. If it’s incorrect. or makes mistakes in some ways, it’s not a huge deal, because your AE is probably savvy enough to spot, like, hey, that smells wrong, that doesn’t seem right, etc. So you don’t have egg in front of your face in front of the prospect at the end of the day. So you kind of got this human in the loop that’s sort of helping you with that. And so that’s kind of a nice place to start. Then you, you know, once you tackle that, you can start to add more complexity to it.

    So you can start to say. you know, where am I willing to have the system start to take action in the real world, okay? Where am I going to start to have it hook up to tools? Okay, that’s interesting. Where will I start to bring evals into the loop? Where am I going to start to measure these things more concretely and scientific-backed? So, those are the types of, you know, capabilities you can start to add as your workflows get more and more sophisticated. But, you know, a simple meeting prep workflow is something anyone could just leave this call today and set up and build for themselves.

    And not… and honestly, ignore probably 99% of the stuff we’ve talked about today.

    Kady Srinivasan:

    Yeah. We did get a couple of questions from the audience, so I’m gonna talk about this. What does the hardness look like for marketing specifically, and any tips to share for other marketers? I think this could be both for you and me, Wade, but I’ll let you go first.

    Wade Foster:

    Well, you know, I think when folks talk about harnesses, like, we’re mostly just talking about a tool in which you can sort of access the models. So there’s a lot of third-party stuff off the shelf. You think of, like. you know, Claude Co-work, or ChatGPT work, or Cursor, like, these are all a Harness that allows you to access these AI models internally. So you can use any of those off the shelf. And have a fair amount of success. I’ve seen a lot of companies start to build their own internal harnesses that help them, you know, with automation as well, too.

    I certainly think… I have a bias myself, which is that I like to use harnesses that allow me to hook into different models. And so, that’s where I say, I’m not sure I want to just go adopt the Anthropic thing, or the OpenAI thing. I like using tools that give me access to other models, and that gives you more flexibility in what you’re building out.

    Kady Srinivasan:

    Yeah, I… the only thing I’d add in there, exactly like you said, is I think I also worry a little bit about tokenization, token costs as well, which makes it even more important, I think, to your point about picking the right models, because we have done this before, where we’ve… tried to create a marketing brain and dumped it all in Cloud, and it just burnt through so much more than I was willing to spend on that, whereas maybe we could have used a Gemini for something similar. So the… the creating the hardness, using the hardness, but keeping an eye on what task needs what kind of horsepower, if you will, from a model perspective is very important.

    And then, I think in advanced stages, like, the routing to the right models, right? I think there are a lot of companies who help do the right kind of routing based on the Context that you’re giving to this specific query. There’s a question for you, Wade. What does, or how does the mental model of building a workflow change when agents are doing the orchestration? That’s a really good one.

    Wade Foster:

    Well, I think in many cases, you probably don’t want the agent to be doing the orchestration, and so one of the things you can ask it is. hey, I want you to build a workflow, or I want you to write code for this. You know, if you install Zapier MCP, you will, Zapier will go, like, your agent, you know, Claude or whatever, will go build a Zapier workflow, and those things will run more cheaply, more deterministically over time, and it will… it will still use AI, it will still use an agent for the places you need reasoning.

    And that allows you to often get. better outcomes than simply saying, hey, I want to build an agent that’s going to handle the orchestration of all of this stuff, because the agent’s going to make mistakes. I mean, we saw that in, you know, the automation benchmark. Like, even the best model, the most expensive model, is only at 40% completing those tasks. But a workflow, with AI is likely gonna get you much, much closer to 100% completion on those tasks.

    Kady Srinivasan:

    Yeah, amazing. We’ve learned a lot about Agentic Harnesses, I think. Any parting words, any words of wisdom for the people closing in on this?

    Wade Foster:

    Yeah, I think the… I’d leave us with two. I think one, you know, if you haven’t really experimented with using different models for different, situations, you definitely should. And two, the second thing I would encourage you all is that, This stuff is… used to be impossible, and now it’s just hard. And so, yeah, there’s a lot of, like, narrative out there that would make you think that this kind of, like, snap your fingers easy. I wouldn’t say it’s quite that. There is a level of, like, elbow grease you have to put in to, like, get these, like, production-grade workflows really humming, but the good news is it’s possible.

    And, you know, in the past, a lot of these workflows were just darn near not even possible. Even if you had, like, really good technical teams and engineering teams, it was just really, really challenging. And now… You know, most marketing ops teams can do a lot of this work.

    Kady Srinivasan:

    Yeah, the one additional plug I’ll probably put in there is the kind of people you hire to be able to do all of this is also very important. I call them multi-threaded marketers because they have a lot of different threads of disciplines that they can bring together. It’s not as easy to create this, so you need to be able to understand the nuances of what is your demand gen, your content strategy, your brand strategy, positioning, and bring that together. the right people, and then use the right technology. To build this.

    Wade Foster:

    Totally. Get those systems thinkers, they’re great. Yeah.

    Kady Srinivasan:

    That’s right, that’s exactly right. All right, and then I think we are done, Julia, on to our next panel.

    Julia Nimchinski:

    Thank you so much, Kady and Wade. And yeah, no pattern interrupts this time.

Table of contents
Operationalize Agent-Native GTM
Work directly with leaders operationalizing Agent-Native GTM for B2B markets shaped by autonomous buyers, agent-native discovery, and machine-to-machine economics.

    Register now

    To attend our exclusive event, please fill out the details below.







    Subscribe me to future HSE AI events

    I agree to the HSE Privacy Policy and Terms of Use *