Introducing the Universal Managed Agents API.Read the announcement
Panel

How to build for agents that act in the real world?

Brainbase Team6 min read
  • Gokhan Egri,
    CEO at Brainbase

    Gokhan Egri, CEO at Brainbase
  • Thais Castello Branco,
    CEO at Taste Labs

    Thais Castello Branco, CEO at Taste Labs
  • Lauren Dai,
    Agentic Commerce GTM at Stripe

    Lauren Dai, Agentic Commerce GTM at Stripe
  • Jon Bunting,
    Senior AI Solutions Architect at Cloudflare

    Jon Bunting, Senior AI Solutions Architect at Cloudflare
Table of Contents

On September 28 we sat four companies down in front of a room of builders at Cloudflare's San Francisco office, in the middle of the Startup Speedrun Hackathon, and asked one question: how do you build for agents that act in the real world? Not agents that answer, but agents that buy things, run workflows and make decisions on someone's behalf.

Each panelist works on a different layer of that problem. Thais Castello Branco, founder and CEO of Taste Labs and early at Exa before that, works on the data and judgment agents need to produce things that are good rather than average. Lauren Dai runs go-to-market for agentic commerce at Stripe, where Link Agent Wallet gives an agent a way to pay. Jon Bunting, Senior AI Solutions Architect at Cloudflare, came in through Replicate and spends his days with teams building on that infrastructure. Gokhan Egri is our founder and CEO. Here is what they talked about.

Good before personal

Thais opened on what models are bad at. They have become very good wherever there is one right answer, and much weaker wherever there are many: design, writing, anything where people disagree about what good looks like. A model that averages over all of it lands in the middle.

Thais Castello Branco
“If you produce the average answer, you're likely creating something that's not that inspired.”
Thais Castello Branco,
CEO at Taste Labs

Her fix separates two things people conflate. Two people with different tastes can still agree a design is well made, so the first job is getting over that threshold of quality; only then do you decide which dimensions should bend to the person. Taste Labs found the web was already homogenising before AI arrived, and the answer it argues for is an inspiration layer: references agents can draw on, the way a designer keeps a wall of them.

Agents buy on a budget

Lauren's advice to anyone building tools agents will use was the most concrete of the afternoon.

Lauren Dai
“Agents just don't buy the way that humans do.”
Lauren Dai,
Agentic Commerce GTM at Stripe

An agent is handed a budget for a task, five hundred dollars to plan a birthday or to run a sales campaign, and it optimises inside that budget. It will not spend a fifth of it on a monthly subscription. So sell per task, and take out everything that assumes a person on the other end: the account, the sign-up, the API key to paste. That is what protocols like the machine payments protocol are for, an agent meeting a paywall, paying, and carrying on.

Human in the loop, as policy

On whether agents should check in with people, the panel agreed on the goal and split on the mechanism. Jon's view is that asking is the point, not a limitation: a model trained on human feedback has a human in its objective.

Jon Bunting
“Human in the loop is not just a feature. It is the intended purpose of the models.”
Jon Bunting,
Senior AI Solutions Architect at Cloudflare

Lauren's concern is that synchronous approval blocks the workflow, so the pattern that scales is approving a policy once: a budget and a threshold under which the agent does not ask. Gokhan added that inside a company, most of the time the agent will not ask a person at all. It will ask another agent that has the right context, the way an employee asks accounting.

Internal agents are their own problem

Gokhan drew the line Brainbase is built on. External agents, the Harveys and Legoras, are high-IQ specialists: proprietary data, an A-team, post-training, all to be the best at one thing. Internal agents are the opposite. They are mid-IQ generalists running HR, accounting or KYC, where nobody has the data and nobody will staff a team to build it.

Gokhan Egri
“In 10 years, companies are going to be spending two, three times more on agents they have inside than the agents they buy from the outside.”
Gokhan Egri,
CEO at Brainbase

What makes those agents work is not a bigger model but the right harness, context and ontology, and a benchmark built with the customer before anything changes, so everyone agrees what success means.

What comes next

The panel closed with a hot take each. Jon expects multi-agent systems that negotiate against each other rather than collaborate. Lauren expects that within six months you will be able to tell an agent to build, run and grow a business from one prompt and check in once a week, because the step changes in models and their cost keep arriving faster than anyone predicts.

Thais has become, in her words, more evals-pilled over time: agent capabilities follow our ability to measure them.

Thais Castello Branco
“The ability to measure, to me, follows the ability to define.”
Thais Castello Branco,
CEO at Taste Labs

So deterministic problems get solved fast, and what is left is whatever is hard to define. The work that matters is deciding what good means and how to evaluate it. She added a smaller take: chat has won as one interface, but the format of an agent's output, structured for other agents and visual for people, is still unsolved.

Gokhan's take is the one Brainbase started on. In 2024 the market treated work as the last thing frontier models would automate; he expected it to be one of the first, with the frontier racing past it toward science. Most workflows are now easy for agents; what they lack is the primitives, like payments and runtimes, not intelligence. Automating labor has become a deployment problem rather than an intelligence problem. He expects frontier labs to drift toward government work, and automating work to feel less like a god in the machine and more like picking the right instance size, harness and context.

Gokhan Egri
“I do not need to use the Millennium Prize model to do my HR.”
Gokhan Egri,
CEO at Brainbase

Thank you to Thais, Lauren and Jon for the conversation, to Cloudflare for hosting, and to everyone who built through the day. The full recording is on YouTube.

Related posts