Transcript

From Business Processes to Business Harnesses

Event held on Sep 8–10th, 2026
Disclaimer: This transcript was created using AI
  • Julia Nimchinski:

    And we are about to transition to our all-star CXL panel. Mark, take it away.

    Mark Organ:

    Awesome! Yeah, very excited to have this all-star panel here, working on… I think this is the hottest area of… I mean, not just tech, it’s kind of cool that this is kind of infrastructure stuff, but, I mean, it really ties so much to business value. that I think this is why, for the first time, HSC is focused on, you know, infrastructure because of the business value, and that’s really what I want to get out of this panel. I want people to have some, things that they can do Monday morning, 8 o’clock. and, drive more value for their companies.

    But first, we need to talk about what a Harness actually is, because there’s a lot of different definitions out there. So I’d love to, get some… each member of the panel to… we’re gonna start with a couple of folks, and then, if there’s agreement, we’ll just say we agree. Or if there’s a disagreement, say, you know, what the area of disagreement is. But, what… what makes something a harness, rather than just an AI application, an agent, or a workflow?

    Olivia Nottebohm:

    I’m happy.

    Mark Organ:

    Okay, thank you.

    Olivia Nottebohm:

    Not a bomb, So for… for how we think about it here at Box is that it’s effectively wrapping a raw model in a secure harness, right? So it’s supplying it with a context layer, whether it’s your company’s PDFs, your contracts, your notes, all that unstructured data. Again, this is the Box perspective. It’s a set of permissions, right? So it’s ensuring that the AI is only seeing what the specific user has access to, and then it’s a set of tools, so the ability to, in our case, extract metadata, summarize folders, or trigger automated workflows. Those are all things that’s provided in that harness.

    Mark Organ:

    Got it. Anybody want to add? Is there something that Olivia’s missing, in her definition, or is there something that, is too much that’s in the harness?

    Arjun Pillai:

    Yeah, I mean, in my world, it’s whatever that allows the model and the agent to do the business jobs reliably. So that is instructions, then the company context, tools, permissions, memory. And the critiquing. You know, you can add more things, but whatever that allows you… allows the agent to do the job reliably, that is what is in the harness.

    Tooba Durraze:

    I would agree. I feel like it’s, like, everything that’s needed to do a job is, like, what I would think of, and that could contain technologies like different models, context, data, but basically, it’s defined by having a purpose or an objective as a consumer.

    Arjun Pillai:

    Marky, if I can give you an example. The model is the new hire, right? You hire it from the… from Stanford. That is the model. Agent is… you give the job description to the model, say, this is the job description. If it’s a salesperson, this is the job description, this is the quota, go sell, right? This is what you call as an agent. Now, Harness is the Salesforce access, the onboarding, the playbooks, the scripts, Gong access, you know, box access, Docket access, you know, all those things that you are giving to the person, so that that person can do the job reliably.

    That is Harness.

    Mark Organ:

    I got it. Now, where does the learning happen? So, now the sales rep, new hire. let’s say, makes a mistake, makes a mistake on a particular opportunity. What learns from that mistake, and then has it so that future Sales reps don’t make that same mistake.

    Olivia Nottebohm:

    I’m just gonna flag, I’m feeling a little nervous that we’re using an analogy as the human, as the LLM, and…

    Mark Organ:

    Yeah, it is a little weird, but I’ve kind of wrong with it.

    Olivia Nottebohm:

    Very quickly.

    Tooba Durraze:

    anchor, yeah.

    Olivia Nottebohm:

    Maybe we could choose a different, line of questioning on…

    Mark Organ:

    Yeah, we can’t… well, it was interesting, because my last one, it was, my interview, my interview, we also used analogies, and that’s how humans… humans best understand things through analogies, so I think it’s good to use analogies, but we need to use the right ones. But, you know, the thing that did come up, again, is this idea of learning, organizational learning. So, we can pick We could pick a different analogy here, but, Is it the agent that does that? Is it the harness that does that? Is there something that sits on top of the harness that does that?

    Where… where does… where does that, learning happen, and how does it get updated and pushed down to, to agents of the future?

    Tooba Durraze:

    I just, I, I think it depends on what the purpose of the harness is. Like, to Lydia’s point, like, the etymology of the word harness implies, like, it’s a container of some sort. At the end of the day, on a… there’s learning on the the Harness level. There’s an objective you’re trying to meet. If you’re not… like, there are decision loops, and things that are happening within that, that are maybe not getting you to your objective fast enough, or whatever it may be, those are traces, those could be… you could… have recursive feedback that exists within the Harness.

    You can have, like, human-in-the-loop feedback, whatever it may be. But I think that there is… they are, depending on the purpose of the Harness. is, like, where the learning can happen. It’s not any different than designing, like, a system, and then what would be the… what would be the checkpoints, whether it’s machine or human-led, that you check whether the system is performing as needed? I would decouple… Models from that a little bit, because there is, like. context coming out of the harness that is creating better learning, and understanding outside of the harness. There’s also, like, model learning and development that happens outside of the harness, so the harness in itself could be a data out for that.

    Mark Organ:

    David, what do you think?

    Olivia Nottebohm:

    I think, also, the harnesses can have different instructions, right? So, for an example, our promise is that our box harnesses ensure that no model that’s being used is training on the data that’s passing through it. Right? And that’s a set of instructions that we provide for our harnesses. And so, to Tooba’s point, it can take different forms.

    David Shim:

    Yeah, and I would say from the harness perspective, the way that we think about it is, like, just like Arjun said, if it’s a new employee, the models are really the brain, and then the harness actually feeds into that, and so… it really is kind of like that contextual knowledge that a new employee goes in and says, like, hey, I’m smart, I’ve got an MBA, I’ve had this history of experience, I’ve got everything that’s available on the internet, and I understand what’s going on. So you’ve got that as a starting point. Now give me context for my manager that is done this job for 10 years, that has a history of emails, messages, meetings that are available, and when you take all that information, you can start to train the models to customize it to you.

    And I think where the biggest value point, and Olivia talked a little bit about the permissions and who gets access to what, but when you do it in the right way, where you can actually stack different teams together, different people within an organization, and share that information, that Harness becomes so much more valuable, because you can append it to anybody that comes in. If you’ve got a different model that’s coming into play, where you’re like, yeah, I want to test out Cloud versus OpenAI versus DeepSeek versus something else. you’ve got that harness that you can go in and say, hey, I want to apply it to this model and see what the output looks like.

    Mark Organ:

    Yeah, that’s cool. Does anyone have it where the Harness can actually create new agents? Or is that not part of what they do?

    Tooba Durraze:

    Yeah, our harnesses, like, are on a macro level, can spawn universes, which is kind of, like, how we contain the world, many harnesses inside a universe. And then on a smaller scale, can spit out, like, even, agents and, like, potentially some structure around, like, a sub-harness that needs to be created, because again, maybe the objective got so broad. Maybe one part of the recursive feedback is that What you’re trying to contain in one harness is too much, right? So, like, then, through the feedback, it should sprout. Like, other harnesses or other agents to be able to achieve that.

    Arjun Pillai:

    Mark, if you think about it, we all are actually creating sub-agents in our harness. Cloud code or Claude Cowork is a Harness, right? ChatGPT or codecs is a harness. So when you give a task to a Claude Code, you usually see, many times spawning sub-agents, right? And if you read the thought chain, you will see that it’s spawning subagents. So, that Harness is actually generating sub-agents to get all those different tasks done. Right? And that’s very, very normal in the motion, depending on the harness. And the harness is defining that experience. Some harness… Claude, if you give the same task to Claude, sometimes Claude will create sub-agents, Claudex might not create sub-agents.

    And you… you… sometimes people respond, say, Claude feels more personable, or Codex feels… to the point, right? This is all that Harness playing its job. In… in the systems that we operate.

    Mark Organ:

    Got it.

    Tooba Durraze:

    Wait, hold on, I have an analogy for Olivia, actually. It was, like, occurring to me. Like, a car, like, if the model is basically your engine, and then the harness is the car, and then you are, like, the person constructing is, like, telling the car where to go, right? At the end of the day. It’s the humid thing I get stuck on, also, yeah.

    Olivia Nottebohm:

    Yeah.

    Mark Organ:

    Alright, where, where was I? So, yeah, I mean, it’s, who asked?

    David Shim:

    Actually…

  • Mark Organ:

    I was gonna move in a little bit of direction. Who in the company actually creates these harnesses? Because I can imagine where you could get some real sprawl. if the head of sales or the head of services, they… they want to… they have their own functions that they want to run, you know, is there… is this the IT department that creates these harnesses? Like, how… who makes these things, and how are they controlled and managed? What are their metrics?

    Olivia Nottebohm:

    So, I would bucket it into two areas. The one is what I call a productivity, like, you’re doing work for yourself, and people will create agents, they’ll give those agents instructions, but those agents are not being shared with other people, right? They, you know, they’ve given a set of instructions that basically helps the work that they do. And then, at the enterprise level, which is, I think, where you have the quantum impact of agents. You actually create agents by function. And they are agreed upon, optimized, and then put into production so that anyone in that same role would leverage that same agent. and benefit from that same harness.

    And there’s a… there’s an intentionality to that harness, because the type of output that’s needed for that role, for that action. is defined, and actually the repeatability and scalability of is very important, and then you can imagine why it’s important, because you then want to be able to link those agents together in workflows, right? Right. So, I think as people think about it, I find that it’s helpful to think about those two very different buckets.

    Mark Organ:

    Got it. Mom. Yeah, go ahead.

    Arjun Pillai:

    In my opinion, today, it’s actually the vendors who are creating a lot of the harnesses, you know, like…

    Mark Organ:

    Good job.

    Arjun Pillai:

    to create a harness for the marketer to put an agent on their website. David’s job is to create a harness that, you know, brings together all the conversational intelligence in that company together. So, we are actually generating the harness, and then, within those companies. they are actually improving the harness by telling it where to go, what not to do, what to do. They are improving the harness. Over a period of time.

    Mark Organ:

    Who does it, though? Is this… is this the… is this, like, the CIO’s organization? Is this ID department, or… or…

    Tooba Durraze:

    No.

    Mark Organ:

    Individual departments have their own…

    Tooba Durraze:

    Who has the job to be done is the person who’s making it, because at the end of the day, it’s like, you… it used to be whatever the rub with business intelligence and data scientists used to be, which is like, you’re not in the business, so you can write the best math possible. At the end of the day, the job to be done sits somewhere else. the architect, like, there’s a technical architect, potentially, that could be someone from IT or a central body, maybe, to Olivia’s point, if there are harnesses that can be used by different parts of the house. at the end of the day, it’s like, whoever designs the strategy or where the thing needs to go.

    So our viewpoint is the opposite, a little bit of Arjun, where I think right now, I agree, vendors are creating a lot of harnesses, because there are not… there’s not a lot of information around how they work, so there’s, like, lived experience in vendors. But primarily, the architect of the harness, like, someone needs to define, like, what are you trying to do? That strategy needs to come from the company.

    Mark Organ:

    Yeah? Right? Yeah. That makes sense.

    Olivia Nottebohm:

    And then to plus one, to Tooba’s point, the company, and then the functional leader who intimately knows the work to be done.

    Mark Organ:

    Right. Yeah, that makes sense. So it does suggest that, people in the departments need to be upgraded, need to learn about these things, and then I can imagine that in IT organizations, there are specialists that… that help, that can come in and help make that happen. But since… since harnesses are so much about business rules. it would make sense that you’d need somebody, to Tooba’s point, who has the job to be done, understands the business rules. It would make sense that those are the people that would be the ones that are building it.

    Olivia Nottebohm:

    The place where we find a really important partnership with our ID team is, integrations. So, most of these harnesses require that you’re accessing data, and you have to make sure that those access points are well thought through, they pass, at least because we’re an enterprise company and we’re dealing with other people’s data. you know, that they pass through security measures and all of that. I think maybe different companies have different stances on that, but at least at Box, we don’t let just anyone create a harness and access pockets of data throughout the company, right? And that’s where the IT comes into play, because they’re basically the gatekeeper of how those integrations get built and how the access points are created.

    David Shim:

    And I would say, like, that’s the case that we would see when it scales, but I think early on, what we’re finding on our end is, like, it’s an individual that has a problem solved, they’re going in, and all of a sudden, it’s like, oh, IT has to approve this, then they go through the steps to get it approved, and then they build something that’s really interesting, that saves a bunch of time, and then all of a sudden, someone’s like, how do I get access to that? And they’re like, hey, this is what you do, and they do it again, IT gets another request, and I think you’re gonna see the same thing here with AI Harnesses.

    It’s gonna start with the individual that has the time, that has the problem to solve, and they’re gonna force IT to go in and say, give me more permissions, give me more access, do the security reviews. But, that’s where we’re seeing a lot of that. Like, we work with a multinational brand right now, where it was one person on one team that turned it on, and now we’ve got over a thousand people that have actually built agents across each one of those brands, and it rolls up to the parent company, and these harnesses are all different, where each company has a different goal, but they also go back to the parent company to say, what is the overall mission, what is the overall, kind of, KPI that we’re trying to drive to?

    Tooba Durraze:

    But that QPI is set by the heads off, right? I think because, like, there is this rub, to your point, David, in markets right now, where, a lot of… similar to how vendors are creating harnesses, there are people on the ground floor creating a lot of harnesses that could be templatized. We’re trying to understand where the value would come from past. Right? Some people are just, like, faster to adoption. But I think… my lived experience is more like what Olivia was saying, which is. at the end of the day, like, in a business, when you’re trying to achieve a strategy which is set at kind of this level, then it trickles down that way, essentially.

    So the governance needs to be, like, at that level. You can’t have, like. Harnesses that are, like, kind of competing with harnesses, or, like, similar objectives, there’s an overlap, but they are kind of doing two very different things. So I think outside of the play and kind of, like, tinker mindset that we’re all in right now, it will probably become a little bit more centralized. Sadly…

    David Shim:

    Longer term, I agree. Longer term, I agree.

    Tooba Durraze:

    Yeah.

    Mark Organ:

    Yeah, yeah, I mean, that makes sense. It’s the same evolution that we’ve seen in a lot of other disruptive tech, right? But let’s try to make this practical for the audience. If you’re advising CEOs starting today, and of course the CEO’s job is to maximize the share price of the company, what business processes would you turn into a harness first? In order to help him or her achieve their mission.

    Olivia Nottebohm:

    I mean, I would say things that are highly repeatable. And, are… even though AI isn’t deterministic, that the results work with AI, right? Like, there are certain things that are highly repeated that you have to have be deterministic, because there’s, like, zero room for any variability.

    Mark Organ:

    What do you mean for the audience? What do you mean by deterministic?

    Olivia Nottebohm:

    Well, you can imagine there are, certain things in terms of… How something is treated in accounting. That’s not up for interpretation.

    Mark Organ:

    Oh, I see.

    Olivia Nottebohm:

    It has to always be treated the same way, right? You can’t have an agent that tells you 95% of the time it’s this way and 5% of the time it’s a different way, right? So I’m just trying to say, like, there are certain places where people understand that Absolutely, it just has to be almost, like, formulaic. But for everything else that is highly repeatable. and where AI is appropriate, then… then building out those agents first is… is incredibly valuable. If you look at… if you imagine, like, a 2×2 of… amount of effort?

    Mark Organ:

    Screen?

    Olivia Nottebohm:

    to require the harness, and then repeatability of the action being taken. If it takes a lot of effort to create that harness, and it’s… that action is only taken once.

    Mark Organ:

    Wow.

    Olivia Nottebohm:

    That action would have to be incredibly valuable to make sense for someone to start by building that harness. Like, typically…

    Mark Organ:

    But…

    Olivia Nottebohm:

    find in organizations is, like, it’s the lower lift ones, like David was describing, you know, someone’s playing around, they have a job to be done, right? But it’s also highly repeatable, because everyone’s saying, hey, can I… can I leverage that harness? Can I leverage that harness? And that’s where you put the time and effort into, you know, what I call hardening the harness, so that People can use it, but in a safe and secure way.

    Mark Organ:

    That’s a great, that’s great. I like it. Deterministic, highly repeatable. It makes sense why you see a lot of customer service type of implementations for this, because you…

    Tooba Durraze:

    I, I think, I think it should be the opposite, though, because it’s like…

    Mark Organ:

    Cool, I like that.

    Tooba Durraze:

    the… you should… if I’m this… if I’m a C-suite person, a CEO, and kind of barring the, the fact that, like, certain things are high left, or we… imagine the perception of that they’re high left. the highest leverage thing. You have a job to be done… there’s a job to be done at this level, at the C-suite level as well. Why wouldn’t you think about building a harness that is, like, highest leverage for you? Because no amount of… Harness is working really well, efficiency-wise, sometimes won’t balance out with the highest leverage thing. It’s like, to us, like, our equivalence is, like, if you had a genie with three wishes.

    Lowest hanging fruit is basically asking a genie to, like, bring you, like, a Snickers bar, essentially, right? It’s like, computationally. it’s such a big thing, that if I’m… if I’m personally, like, a C-suite person, this is how we advise people, is efficiency will always find a home. Find whatever is the highest leverage job we’ve done for you, and, like, create Okay, I…

    Mark Organ:

    I am resonating, you know, with what Olivia said about hardening the harness, like, you need to go down the experience curve, like, if you… so what I like about having something repeatable… so again, I think that’s why customer service implementations seem to be an area that is getting automated a lot with agents, and that’s because there’s a lot of opportunities for learning.

    Tooba Durraze:

    But it’s the lowest-hanging food problem, though, right? Fidelity-wise.

    Mark Organ:

    Okay, yeah.

    Arjun Pillai:

    And there’s a reason to that too, right, Mark? Adding on, plus warning, Olivia, and then adding on a little bit is… There is also the cycle time, right? If you look at a support ticket or, like, a coding problem, coding is probably the number one thing where it has broken out, right? Every company is using so much coding. That’s because to take… taking alleviasing and expanding a little bit, there’s enough volume to learn. about code. There’s a clear outcome. A code either compiles or it doesn’t. There is enough available context in the company. There are mistakes you can contain, because humans are basically, looking at the code if there is error.

    And then, you know, it’s a quick cycle time, right? You write a code, you compile a code, you write a code. Versus if you take, like, can I close a deal, and try to put a harness, if your deal cycle is, like, 180 days, it takes 180 days for it to cycle back, right? So don’t go and choose those kinds of things to harness out first. Just look at enough volume to learn a clear outcome, there is enough context, mistakes are okay, just go pick those as the Harness to build first. And then, plus one into Tooba’s point, where there could be that high leverage thing, which is one-off, but it’s still high leverage, you can still build it.

    But I’d keep that cycle time as a key criteria of defining what should be your first harness.

    Mark Organ:

    Right. David, what do you think?

    David Shim:

    Yeah, I would say from a… day-to-day perspective, for the harnesses to… and the AI agents to have really high value, it’s having the right ingredients that you put into it. So, every single part of this is the system of record that you put in. I will selfishly say, for meetings, it’s read.ai, and having that ephemeral data. of meetings where all those conversations occur for customer support, that all disappears today. So that’s like Snapchat today. So, if you think about Snapchat, you get a message, you see it, and it disappears within 24 hours. There’s no kind of harness that captures that information.

    And if you think about everybody on this call, they’re in a Zoom meeting right now, that time disappears if you actually don’t harness that meeting content, that context, etc. And for us, it’s like, that’s what you should be capturing right now, because you have a system record for email. You’re able to pull that down today. You’ve got HubSpot Salesforce, you’ve got Box for your files, you’ve got Docket. You’ve got all these different products that are in place today that you can pull in, but the one thing that is ephemeral, that disappears every single day is meetings.

    So I would go in and say, that is the easiest thing to bring in on top of your existing system of records.

    Mark Organ:

    Yeah, I got it. Yeah, the example I thought of may not be a good one. So, the one, and it’s… maybe it’s one that I incur, sorry, I find a lot in my work, which is that of pricing and packaging. So, CEOs might say, should we raise our prices next year? Who should we raise them for? How much and why? But, that’s not really a very repeatable kind of process, and I wonder if my example really isn’t a very good one there. That’s not the kind of thing that we should try to get a Harness and agents to do.

    Tooba Durraze:

    I had an example of that show up, actually, with an actual customer, where it didn’t… there wasn’t a harness around pricing and packaging, but the harness was around 3 highest average things that CC person can do today, like, 100-plus million in revenue size, company-wise. Three things that they can do today, in terms of, like, meeting their quota. And the thing that came up from it was. pricing pressure based on pricing and packaging data change for two of their competitors. So that would be an example of, like, you could spawn off. like, a harness or something, right?

    But there’s, like, the objective is not just general purpose, pricing and packaging changes, but, like, hey, like, materially should we change? Based on, kind of, what the market just shared, to meet our goal for the quarter, essentially.

    Mark Organ:

    Got it. So, I mean, in that example, what… what does the harness… what does the harness do? What do the… what do the agents do, and what does the model do in… in this… in that example that you gave?

    Tooba Durraze:

    So, agents are the worker bees, like, back to that kind of analogy, like, that I said, they’re doing, like, the micro jobs on top of the bigger job to be done, right? The Harness constraints are Pricing and packaging, not just general, but pricing and packaging in the context of this specific market pressure and the goal of meeting like, our revenue requirements for this quarter. And actually, in this, like, the kind of outputs that ended up being were, like. No amount of material change in pricing and packaging. Olivia’s gonna think this is commonsensical, but no amount of change in pricing and packaging right now is actually going to result in materially you impacting your revenue for this quarter, but keep it as an always-on so you can preemptively think about, like, when it needs to be changed.

    So, the CEO there has a lightbulb moment to say, I think we should think about pricing and packaging changes. Preemptively, can it be watching certain things? And, like, output, like, okay, now is the time to think about that, essentially.

    Arjun Pillai:

    Mark, on that particular question, agreeing to everything that has been discussed, but one thing that I would say is, in this case, the agent is actually… you are curtailing the agent from taking an action in terms of actually testing out the pricing. What that means is, the useful recommendation of that harness is, here is what you can test. In this segment, you can change it by 10%.

    Mark Organ:

    His name.

    Arjun Pillai:

    expected upside, and what would make you stop? What would make you go, right?

    Mark Organ:

    Yeah.

    Arjun Pillai:

    this is what the harness will give you, right? It’s not like the agent is gonna get on a call and say, we are increasing our price by 10. That’s not gonna happen, right? So…

    Olivia Nottebohm:

    Hopefully not. Hopefully not. Hopefully you’re doing some A-B testing into the market before.

    Arjun Pillai:

    Yes, hopefully we’ll build that agent that does it in the future, but at this point, and I want to call it as a distinction, right? We are talking about harness and everything that goes into the harness. It doesn’t mean that every harness will have every element. In this case, the action of the agent is actually curtailed. It is not gonna do it. The output of the agent is an experiment that Olivia is gonna give to the regional salesperson of the enterprise segment to test it out. Right? And then the test results will come back that you put back to the same Harness and say, here is what we found, what do you think?

    Right? So, I just wanted to call out that.

    Mark Organ:

    That’s cool. Very clear. Very clear. I really like that.

    Olivia Nottebohm:

    Mark, your original question was, what are some things people can go build? And when I think of go-to-market, because it’s kind of what I think about for Box, is if you go function by function, there are just highly repeatable actions that people take. That if you were to provide agents and harnesses.

    Mark Organ:

    He just didn’.

    Olivia Nottebohm:

    with a set of instructions for those people that they’re able to then elevate and do more complex work, right? So if you start with marketing, you can give a whole full set of instructions for, you know, what is the tone of, in our case. the box messaging, you can give it access to all of our write-ups of all of our products. I mean, you can really provide so much context into that harness. Right? And then you can give a set of instructions, and lo and behold, you never have to write your own blog again. Well, at least.

    Mark Organ:

    Yeah.

    Olivia Nottebohm:

    Right? And then you can go function by function, right? And for sales, you have SDR motions, where you’re sending outbound emails. And then for sellers, well, you have to prep for a customer meeting, and then for professional services, you have to write an SOW. You can literally create harnesses for each of those. We have, at Box, we have 63 harnesses. And you can then move to pricing, but if you ask me, like, where would I start, I don’t think I would start with pricing.

    Mark Organ:

    You wouldn’t start with pricing, but I like what Arjun… what Arjun had to say, because that really clicked for me. If you can find a way to make it repeatable, iterative and repeatable. So, even within the pricing thing, you can find high of all 10% of the customer base, let’s say, and I’m gonna go and try some things to see if I can upsell them, and I’m gonna go and put some guardrails around that. now making it iterative and repeatable, now you can actually get the return on it. But yeah, to your point, I came up with a very challenging, example.

    But yeah, it makes sense, it’s kind of like a style guide. I remember, You know, when I was running Influitive, and I drove the marketing department crazy, because I always took their press releases, and I’d mark it up full with my red pen all the time. And I remember my coach challenged me, he said. I want you to never have to mark up a press release again. How are you going to give instructions to the marketing organization so they just know how to write a press release properly the way that you like it? And so I had to really think about that, like, what’s the standard thing?

    I said, okay, no paragraph could be more than 7 sentences, no sentence could be more than this one. You know, all the quotes had to go and back up the story, and actually, that really became a harness. And I didn’t have to ever look at a press release again, I could trust my team to go and do it well. I don’t know if that… if that S was sort of coming up for me as I’m hearing you guys,

    Olivia Nottebohm:

    Yeah, and then there’s, you know, there’s a hack now, which is you would have just ingested all press releases that you’d approved, and it would have.

    Mark Organ:

    Oh, yeah.

    Olivia Nottebohm:

    Backed into what your set of instructions were.

    Mark Organ:

    Oh, that’s cool.

    Arjun Pillai:

    And Mark…

    Olivia Nottebohm:

    Right? We just say, make the tone of all of these documents, and then it figures out its own set of instructions.

    Mark Organ:

    And do those instructions then go back and repopulate the harness, then? Does the harness get now… Get smarter as well, that’s very cool.

    Arjun Pillai:

    Yeah, and this is why we call that Harness Will Compound, because, you know, you just created a harness that compounds, so that function, that particular task is yours now, right? Your company is yours now. Whatever models come, it’s only going to improve what you already have harnessed. If you can keep harnessing more and more things, this is how you end up building, like, a factory inside your company. Right? So, this is… this is the PR part of your marketing, right? If you do AEO like this, if you write, to Olivia’s point, content like this, if you do demand gen ads like this, that expands into a marketing factory within the company that is harnessed for your company, that is compounding all the time.

  • Mark Organ:

    Cool. So how do I know… so now I’m a CEO, I’ve got a bunch of harnesses and agents doing stuff and getting smarter over time. What metrics should I be tracking to know that I’m doing well? How do I benchmark my success in sort of AI-ifying my company against… I mean, obviously, I’ve got business metrics that I can track, but are there metrics that you’re using in your organization to track the sophistication of these processes that you’re automating?

    Tooba Durraze:

    Yeah?

    Arjun Pillai:

    I can keep… go ahead.

    Tooba Durraze:

    Sorry, sorry, let me go quickly, and then I’ll pass it through, because you have a more concrete answer. It’s both ways. There’s, like, if you’re a vendor who’s, like, creating and producing harnesses, at the end of the day, you have to give both on quality. I saw, a comment from, the audience on, like, okay, well, how do we know it’s performing well? So, there are some kind of performance metrics that are. core to, like, how you construct the harness, and when you’re constructing the harness that you need to have in place. The same way, when you’re creating, translating strategy into work, you’re gonna say, okay, I’m going to measure this work by this criteria of success.

    And then the flip side of it is also true, which is, if you are C-suite, hopefully all these harnesses that you’re creating, you are standardizing, like, the way that you will look at success, so you, at the C-suite, to your point, Mark, it needs to build up to a whole, right? You need to tell the picture of, like, these things are working well together. I don’t know if AI fine is, like, a quantifiable metric in that way, but…

    Mark Organ:

    Huh.

    Tooba Durraze:

    Essentially, like, how many of our business processes that are repeatable have been kind of taken over by Harnesses, allowing us to have the capacity, yeah.

    Mark Organ:

    Yeah. Oh, that’s a good one.

    Arjun Pillai:

    I can give two examples. One is my company, I went to my engineering leaders and say, hey, do you want to hire? They are immediately like, nope, more tokens, please. Right? So they don’t want to hire. And then I asked them, like, what does this really look like? And they gave me this. Whatever took 7 days to go from 0 to dev. now takes 2 days, right? From dev to production, it is still not as fast, but to dev, it’s, like, super fast, right? So there is a… there is an efficiency gain there. It’s roughly, whatever, 120%, 125%.

    So that’s an outcome that you can measure. But then there are leading indicators. I read this about Uber. Uber recently published that their coding factory, their software factory, has 3,600 skills that are getting fired 30,000 times a day. Right? So these 3,600 skills that they have given, harnessed within Uber for their coding purposes, and they are firing 30,000. That’s a leading indicator. Does that mean that.

    Mark Organ:

    Yeah.

    Arjun Pillai:

    Shipping everything? No, but it’s certainly the leading indicator. So, my first example is the actual indicator of business impact in terms of efficiency that you can track after the fact, and then the second is an example of a leading indicator that you can track to know whether you are AI-ifying your company in the right direction.

    Mark Organ:

    Yeah. That’s cool. Would you guys agree with, like, number of skills and sort of the number of times that they’re used? Are those good metrics?

    Olivia Nottebohm:

    I think it’s an interesting statistic, because it implies a certain amount of work has gone into it. You could potentially get the same work done with fewer skills. So I don’t know if it’s, like, a volume-based thing, you know, back to the time of token maxing. But I do think that the actions taken in business terms is what I would anchor most on, right?

    So, if your goal is to drive awareness. and you’re going from X percent unaided awareness to Y percent unaided awareness, and then you have a plan on how to drive awareness, and that includes Publishing blogs, getting editorials, etc, etc, etc, and you know you have to increase the surface area to get increased awareness, and you’re using agents to do it, and you’re not having to To do that, you know, having a scaling problem just with the team size you have. Then, you can very quickly track the business impact, right?

    David Shim:

    Is that where you’re… oh, sorry.

    Olivia Nottebohm:

    No, no, I mean, so I would focus, if you can, more on the business impact. Ultimately, the final lagging indicator, at least for a business, is revenue, but there’s a bunch of leading indicators in the business impact that you can and should be able to attribute to before you get to revenue.

    David Shim:

    And where I’d map it back to is, it’s a lot like online advertising right now, where it’s about the attribution. How do you attribute the value that you put in? And right now, when you start talking about, hey, the number of calls that were made. That’s like ad impressions. Hey, if I serve enough impressions and my revenue goes up, I think it’s a good thing, and that’s okay. Where we’re going towards is mapping it back to the actual conversion event. So if you’ve got Salesforce or HubSpot connected, what were the features that were pulled into that harness that said, this is what I recommended, the agent recommended you do, did it drive the ultimate conversion at the end of the day?

    And being able to attribute that is going to drive the adoption, and like. from a meetings perspective, when we first started, it was like, hey, we give you meeting notes, that’s great. People are like, that’s helpful. Then they’re like, hey, give me some recommendations, have your agents tell me what meetings to cancel. And so what we found was, like, we saw 20% fewer meetings across the entire org if they adopted the platform. So they were listening to the agents saying, like, cancel this meeting. Hey, this could be a follow-up email instead of a meeting about a meeting.

    And so when you start getting.

    Mark Organ:

    Yeah.

    David Shim:

    nudges, people are like, okay, this is great. Then they start to say, can I plug this into different platforms where they can access that data? And then you get to that final attribution piece, which is the revenue side.

    Tooba Durraze:

    I think in the absence of, like, historical data around these things performing, and mixed with, like, to Arjun’s point, decisions are, like, take longer times, right? Sometimes. you come up with, like, all these, like, proxy metrics. It’s the equivalent of, like, when we say an event is successful, you’re not gonna know if you got pipeline out of that on first day, or revenue. You might say an event is successful because this many people attended this year versus last year. Like, today, at the end of today, right? So, I think some proximity metrics like, the ones, like, we were mentioning, which is, like, okay, but it’s, like, running on these skills, or, like, it’s, like, took away, like, two to three hours a week per analyst, etc.

    Those help in the short term as you’re building towards, like, what the impact is actually fundamentally going to be.

    Mark Organ:

    Got it. Yeah, I got a question for the audience here. how do you ensure the model consistently follows Harness instructions and keeps its behavior within the boundaries you set? What practices or controls help you confirm that the model remains… remains aligned over time?

    Arjun Pillai:

    Good question. I can maybe start… instead of trying to answer that as a big question, you know, that’s a big loaded question, I would first break it down into specific questions. Did the agent… the answer that the agent gave within the harness and the model, was the answer supported? Was the action permitted? Did it follow the system policy, you know, basically the prompt. could you reconstruct what happened? Right? The Harness can help with all of this, but I’ll break it down to specific questions first. And then, even with all the Harness, everything that we do, we cannot promise that every answer will be correct, every permission.

    It will go, you know, left or right. Then, what you do is… First, make the failure visible by breaking it down and figuring out where exactly is the problem. And then look at how to solve this, right? Is it a prompt issue, or is it a model capability issue? Like, you know, then you kind of start solving it one at a time, instead of thinking that, okay, one agent doing this very hard, complex problem, did it align to everything, right? That’s how I would do it, but…

    Olivia Nottebohm:

    Yeah. I would also say it is rare that your first go at a set of instructions, or the harness, like that portion of what the harness represents, is great. Right? What we find is that the instructions set, you see the output, and then you realize, oh, I should tweak this, oh, I should tweak this, oh, I should tweak this. And so what you’re doing is you’re modifying the instructions. You’re not… like, at least at Box, we’re using other people’s models. Like, we’re not training our own models. Very few companies train their own models, actually, right?

    And so what you’re trying to do is really make sure your instruction set is dialed in so that the models are giving you the answer you would like. Now, you definitely do see different results with different models, right? So you have to take that into account, and as someone, you know, as a team that’s building harnesses for customers to use. what we’re trying to do is drive an efficiency curve, where it’s like, we’re getting the outcomes that the customer wants, but we’re not having to constantly use the latest and greatest frontier model, which isn’t really expensive and would cost them a lot of money.

    And so I think that’s also something that people building harnesses have to navigate.

    Tooba Durraze:

    I think telemetry, like, to Arjun’s point, the worst thing you can do is have silent failures, because if you don’t break it down into, like, smaller problems, the way you’re constructing a system. like, it will get unwildly, because you will build, and then you will build more, and then you will build more on top. So just break down the, like, how you measure things. Telemetry is, like, really important, even if you don’t look at it every day. And then the second thing I would say in the spirit of kind of tinkering is… Don’t aim for perfection, but build guardrails.

    So, don’t let it autonomously do anything yet. In the beginning, watch and observe, like, what it’s asking to do, and say approve or not approve. That’s how you’ll learn, do I need to instrument more ways to measure where the failures might be? Do I need, again, to Olivia’s point, like, better instructions? But the starting point should include, like, measurement or telemetry as a part of the architecture of the harness.

    Mark Organ:

    Yeah, no, that really speaks to the, you know, trust. And there’s… there’s a lot of mistrust around agents and models. You know, these days, we’ve heard a statistic that over 90% of AI pilots are failing in businesses. I can’t believe that, but whatever, I’ve heard that stat. We’ve all seen hallucination. You know, so, what does trust actually mean in a business harness? You know, does it mean it just won’t do something dangerous, it just won’t hallucinate, it has permission to do this? You know, which of these sorts of things are actually harness problems?

    David Shim:

    I think trust is less of an issue, sorry, not trust, hallucinations are less of an issue. I think we all know the model to hallucinate. To Tooba’s point, like, you want to break it down, you want to have checkpoints where you say, I approve or disapprove of this action. I think trust is the most important thing right now, because if you’re going to go in as an employee, if I’m going to put all my data into a harness, I want to know that I’m not going to get screwed over down the road because something was used against me. or I didn’t know that this data was being in because I was in a one-on-one, and I said something, and it got pulled into the harness, and now my boss sees it, the entire company sees it.

    So, it’s important to have that level of trust to put that data in, and a lot of it is about being transparent with your team, where it’s not a big brother-type situation, where it’s like, everyone has to use it, but it’s like, hey, when it comes to meetings, emails, messages, Salesforce, all that information is going into the harness. be wise in terms of what you want to be able to measure within that system. And then people can go and say, I forgot to turn that off, that’s on me, I didn’t like what the outcome was because of it, but I had control because the company said, this is what’s going to happen.

    And once you have that level of trust. people will start to share more and more things, and I think the best example is social networks. Initially, people were like, hey, I’m never gonna post pictures, people are gonna see my family, they’re gonna see my kids, all this stuff. Now people are going and saying, like, okay, there’s a level of trust where it’s like, okay, if I decide to post, I can make it all my friends and family, my close friends and family, or just to a single person. And I think you’re gonna see that with AI harnesses, where people are gonna control who gets access to what.

    And the enterprises are going to enable that.

    Arjun Pillai:

    And, Mark.

    Olivia Nottebohm:

    One of the key components of trust is that the system has authorization to act, right? This security element, that it’s actually grounded in the right data or information, and then you can actually explain it if something goes wrong, right? Something will go wrong, but is it clear enough that you can explain why it went wrong?

    Mark Organ:

    Right.

    Olivia Nottebohm:

    You even have that when, like, as Box, we offer out metadata extraction as a harness, right? That’s an Agentic harness that we offer. And in that scenario, we offer confidence scores, right? So you then also are asked, sometimes in a production environment to not only, like, gain trust, but you’re actually putting numbers on the level of trust, quote-unquote, of the veracity of the outcomes that’s coming out of that harness. And when you’re running enterprise processes, that also is really important.

    Mark Organ:

    That’s huge. Yeah, Arjun, you wanted to comment.

    Arjun Pillai:

    Yeah, Mark, my agent is slightly different in the sense that we are one of the agents that our customers are directly putting to speak with their customers. Most of the agents that we see, we use internally, versus, like, it has to go externally. So trust is, like, a really, really big deal out there. One thing that I have noticed with customers is They… they want to know that it’s working the way that they intended, right? And trust has got two pieces to it. As an example, one customer came to me and said that, hey, your agent is answering pricing questions.

    And when we looked at it, yes, it was answering pricing questions, but it was answering from a blog post that it was written, like, 3 years ago. So there is always that thing where, is the agent going wrong, or is the data that went into the agent that is going wrong, and then that’s creating the trust issues? One thing that we recommend to our customers is, hey, expand the autonomy one action at a time. Don’t take an agent… like, when I am starting to use my agents to write email, right, I give the read access first.

    Then I gave the triaging access to my email. Now I have given the send access. I did not give all the access on day one, right? Yeah. So you observe first, recommend second, then execute with approval, so you keep expanding the autonomy of your agent, one action at a time. I… to cross that trust factor, right? To David’s earlier point, you have to feel comfortable to extend that autonomy to the agent. That’s how we have… we help our customers. If you don’t feel that the agent should charge their credit cards, don’t give it just yet.

    Let’s keep it running for 3 months. Once you have the trust, let’s give that actionability.

    Tooba Durraze:

    It’s like the neighbor that moved in next to you that’s brand new, after you start hanging out with them, they will be trust. Like, on the technical side, to Arjun and Olivia’s points, like, yeah, you need, like, logs, traces, you need to, like, be able to see evidence, because that’s one part of it. The psychology part of it is, like, basically, like. introducing broccoli into people’s diet, but, like, disguising it as chicken nuggets. It’s frequency of you.

    Mark Organ:

    Genius?

    Tooba Durraze:

    and it coming kind of in your realm will start to build trust, because at the end of the day, like, even to, like, David’s saying, like, we don’t, we know hallucinations. Now, look at how far we’ve kind of come where hallucinations to freak everyone out. So it’s like, if you understand it, there’s nuance around it, you understand how to, like, take something from it. Versus, like… and Arjun, actually, this doesn’t apply to you, because you’re right. In your point, like, outward-facing agents and harnesses have a very strict criteria of trust versus a lot more leeway of internal-facing stuff.

    But I think, yeah, I think it’s a matter of, like, treat psychology by recency frequency, and technically have, like, evidence trails and all of that, so you can go back…

  • Mark Organ:

    Yeah, that’s cool. I think what I take from it is just trust is extremely important. It has to be engineered into the system. And yeah, I think it’s like everything, or like anything. Trust is critical. But, that leads to, kind of, where… what do the humans actually do in this world? I mean, today, and then also looking out maybe a couple of years. What do the people do versus what the harnesses do, and the models and the agents do? What do they do today, and what are they going to be doing in the next 2 or 3 years?

    David Shim:

    I think from a today’s standpoint, we’re more manage… we’re getting to become more managers, where the agents are going in, taking the harness, making decisions, and then checking in with us to say, like. can I move forward with this? Can I not move forward with this? And they’re taking action against it. I think the best example of this would be, if you look at all the hedge funds in the world today, they would argue that they had AI for a really long time, and they would go in and say, the models say take this action based on this, this variable, but there was always a person at the end of the line that said, like, okay, go make this transaction, don’t make this transaction, outside of the high-frequency trading stuff.

    But, like, there was, hey, I want to deploy $2 billion against this. someone has to go in and sign off on it, and I think that is the value proposition where you can’t really fire an agent, but you can fire a person, and by being able to fire a person, it gives someone a responsibility to say, I need to look at this, I need to make sure that this will not have a negative impact, but a positive impact on the business. And so I think that is our role today. I think as we start to go into the future, I’ll leave it up to other folks, but I think it is going to be more autonomous at a certain point.

    They’re gonna go more run, and now you’re more observing, and you’re deciding when you need to step in.

    Arjun Pillai:

    Yeah, 100%. Humans will spend way more time deciding what good looks like. You know, our jobs are…

    Mark Organ:

    quickly.

    Arjun Pillai:

    evolving into what good looks like and improving how it gets done, right? So, this includes setting of objectives. I own a function in my company, so I set the objective for my agents. then I handle unusual situations where it is… the agent is, like, confused, conflicted. I handle the situations. I correct the system on a daily basis. It’s some way of strengthening the harness, right? Correcting the system. And then, obviously, maintaining the relationships with other head of departments and all of that, right? So that’s what humans are actually basically evolving to, like, the managerial thing that David is saying.

    In AI, we talk about taste, right? still subjective and it is human. And what is taste? Taste is a cumulative set of decisions that you give to the agent. Hey, use Lato as a font and not Roboto. That is part of the taste. juice gradient in the background and not solid color, that is a taste. So you keep… this is a design agent, this example is of a design agent, but you keep improving that agent to a point where that agent’s work reflects who you are and what you want, and you have given the taste to the agent.

    That job still is with the humans.

    Tooba Durraze:

    I say you’re in the car, you have to tell the carburetor to go. You can have the car all you want. If you’re not telling the carburetor to go, someone has to think about the gas in the car, maybe at some point like that.

    Mark Organ:

    Hmm.

    Tooba Durraze:

    kind of automates, but at the end of the day, there is, like, if there is no job to be done, like, then it’s like, what is the involvement, right? So, you need the car to do something, you still have to tell it where to go.

    Mark Organ:

    Yeah.

    Tooba Durraze:

    machinery kind of drive it. And to Arjun’s point, some people might take the scenic route, some people might take, like, the fastest route. That could be, like, your preference, your taste, but if you don’t have a destination, you sitting in a car is not going to do anything.

    Mark Organ:

    Yeah, no, that makes sense. Like, it’s the destination, it’s, edge cases, so if you think about driving a car, most times you’re on autopilot, and eventually the car’s on autopilot all the time, but you’re on autopilot, but sometimes there’s special situations where you have to Take over, essentially, and think things through. And so that would make sense, along with just making sure the car is actually working properly. So, if the car’s got problems, you gotta take it in and get fixed. So, how’s that as an analogy? Driver of the car.

    Olivia Nottebohm:

    Love it.

    Mark Organ:

    Alright. Yeah, so, imagine that we… Come back to the summit in 2030. What does an extremely effective, highly agentic company actually look like? Does it have, like, thousands of harnesses? Does it have, like, one mega… You know, overarching harness over everything. How many, you know, how many people are in there? What are they… You know, what are… what are the, what are the, how many humans are actually working inside traditional workflows versus supervising and improving systems? Take me through what the company of the future actually looks like here.

    Tooba Durraze:

    Can I just throw a joke in there?

    Mark Organ:

    Yeah, that’s true.

    Tooba Durraze:

    By going from a car to a school bus, get more people in the same vehicle. headed in the right… in the same direction, I think is an important part of it, right? All cars must… like, at the end of the day, like, it’s like a caravan. There’s one destination for the. Which is revenue to make the company grow. In my future world, we’re in a party bus, a school bus, more people, less harnesses, more people in, like, one universal harness, multiple universal harnesses, headed the same direction.

    Arjun Pillai:

    I disagree to some extent. I don’t think it’ll be thousands or ten thousands of harnesses. Mark, I think the best way to think about it is to expand what we already know, right? If you look at all the departments in a company, certainly product engineering is the one that has made the farthest progress in terms of AI transformation. 90% of all the processes have changed. So for us to see what would the change look like in other departments, the best way is to look at what has happened in software coding, and then try to kind of simulate that to different things.

    So I am… I’m fairly certain that in the future, what would happen is there are several major business systems, and when I say business system, I mean there’s a software factory for coding. There’ll be a marketing factory for marketing, there’ll be a sales factory for sales, customer support, so on and so forth, and those systems are responsible for the outcomes. And inside these systems, there are many specialized agents within them, and the number will obviously depend on… Box will have very different agents than Docket, right, to be fair. And these will share definitions, identity permissions, all of these things.

    But I think these business systems will be there, and humans will be out of the loop, and working with these systems in terms of giving it direction, taste, and all of that. So that is the future. The one that I really like, right, Elon Musk talks about this. Elon talks about, I build the factories, so factories is my product. and the factory builds Tesla the cars, right? So that is the analogy that I like. So I will build the factory which will build my product, right? So my job is to be out of the system, take care of the factory, and the factory’s gonna build my marketing, build my sales, build my factory, with… with different, different, involvement, from the humans, depending on the departments.

    Olivia Nottebohm:

    My view is that… well, first of all, you know, we… I’m sure you’re contemplating, right, vastly different industries. Like, running a hospital is obviously very different than running a software company, et cetera, et cetera. But, I’ll stay within software, since it’s what I know well. And for that, I actually think that We will have different job descriptions for humans. And those humans will be leveraging various harnesses. But it will have to have shared security, and you still have to get the information that the harness uses correct, right? So you do have that, shared consistency across any harness that gets built in an enterprise. that is trying to run business functions, or coding, or, you know, whatever activities it’s running.

    And I actually, yes, of course, believe that people will be managing agents, but from a growth mindset perspective, I don’t view it to reduce down to, you know, 5 humans and a sea of ages. There is a vast set of things that humans still need to be doing and need to be interacting with other humans on. And that, I think, is, ongoing, and the need to interact with the outside world and with customers and with partners and all of that, I think is very real. So, yes, I think that humans will be buoyed and can do far more, and frankly, will probably be able to do it at a higher quality level. by 20…

    Mark Organ:

    funny.

    Olivia Nottebohm:

    And that is exciting to me because of what we can arrive at, but I don’t view it as a negative doomsday situation.

    Mark Organ:

    Got it. Yeah. Well, that has been, I mean, that has been the history of disruptive technology, is that there have been more and more people employed, in different functions that didn’t exist before. And you’re right, delivering higher levels of quality and value. So it would make sense that, that would more likely continue than, than not continue. So that is cool. Alright, we got a few minutes left. Lightning round, finish the sentence, I want each of you. The most valuable part of an enterprise AI system five years from now will be the blank. Model, the data, context, Harness.

    Human beings, something else. We’ll start with you, Olivia.

    Olivia Nottebohm:

    Oh, for sure, context. 100%. I think model.

    Mark Organ:

    That’s right.

    Olivia Nottebohm:

    Commoditized, but all that will matter will be making sure the correct context is being provided.

    Mark Organ:

    Got it. David?

    David Shim:

    Context, I think it’s going to be really important what systems of record that you start to build on today, because that’s what you’re going to be investing in building your models, the harnesses you’re going to get trained on.

    Mark Organ:

    Okay, Tooba?

    Tooba Durraze:

    I have a slightly different opinion. I think it’s going to be the direction, whoever sets the direction, because you… in our experience, the amount of harnesses that have failed from a lack of people being able to, like, describe the thing you’re trying to make it do. has, like, a really big impact, despite having the best context-backed models and all of that. Know how to use a machine, and what you’re using it for.

    Mark Organ:

    Hey, Arjun?

    Arjun Pillai:

    I think it’s the same. Accumulated judgment of the company, what it learned from the decisions and outcomes, consistently being fed back to the system. Over a period of time, that’s going to be the most valuable thing for the enterprise companies.

    Mark Organ:

    Right on. And I wonder who’s gonna be in charge of that, function in companies. That’s fascinating to think about. But we are, we are at the end of our time. Julie, I’d like to hand it back to you, and for any sort of closing thoughts.

    Julia Nimchinski:

    Thank you so much. Such an insightful panel. Our audience is writing that it could last 2 hours. Easily, so…

Table of contents
Operationalize Agent-Native GTM
Work directly with leaders operationalizing Agent-Native GTM for B2B markets shaped by autonomous buyers, agent-native discovery, and machine-to-machine economics.

    Register now

    To attend our exclusive event, please fill out the details below.







    Subscribe me to future HSE AI events

    I agree to the HSE Privacy Policy and Terms of Use *