How to build for agents that act in the real world?
Gokhan Egri,
CEO at Brainbase
Thais Castello Branco,
CEO at Taste Labs
Lauren Dai,
Agentic Commerce GTM at Stripe
Jon Bunting,
Senior AI Solutions Architect at Cloudflare
Table of Contents
On September 28 we sat four companies down in front of a room of builders at Cloudflare's San Francisco office, in the middle of the Startup Speedrun Hackathon, and asked one question: how do you build for agents that act in the real world? Not agents that answer, but agents that buy things, run workflows and make decisions on someone's behalf.
Each panelist works on a different layer of that problem. Thais Castello Branco, founder and CEO of Taste Labs and early at Exa before that, works on the data and judgment agents need to produce things that are good rather than average. Lauren Dai runs go-to-market for agentic commerce at Stripe, where Link Agent Wallet gives an agent a way to pay. Jon Bunting, Senior AI Solutions Architect at Cloudflare, came in through Replicate and spends his days with teams building on that infrastructure. Gokhan Egri is our founder and CEO. Here is what they talked about.
Good before personal
Thais opened on what models are bad at. They have become very good wherever there is one right answer, and much weaker wherever there are many: design, writing, anything where people disagree about what good looks like. A model that averages over all of it lands in the middle.

“If you produce the average answer, you're likely creating something that's not that inspired.”
CEO at Taste Labs
Her fix separates two things people conflate. Two people with different tastes can still agree a design is well made, so the first job is getting over that threshold of quality; only then do you decide which dimensions should bend to the person. Taste Labs found the web was already homogenising before AI arrived, and the answer it argues for is an inspiration layer: references agents can draw on, the way a designer keeps a wall of them.
Agents buy on a budget
Lauren's advice to anyone building tools agents will use was the most concrete of the afternoon.

“Agents just don't buy the way that humans do.”
Agentic Commerce GTM at Stripe
An agent is handed a budget for a task, five hundred dollars to plan a birthday or to run a sales campaign, and it optimises inside that budget. It will not spend a fifth of it on a monthly subscription. So sell per task, and take out everything that assumes a person on the other end: the account, the sign-up, the API key to paste. That is what protocols like the machine payments protocol are for, an agent meeting a paywall, paying, and carrying on.
Human in the loop, as policy
On whether agents should check in with people, the panel agreed on the goal and split on the mechanism. Jon's view is that asking is the point, not a limitation: a model trained on human feedback has a human in its objective.

“Human in the loop is not just a feature. It is the intended purpose of the models.”
Senior AI Solutions Architect at Cloudflare
Lauren's concern is that synchronous approval blocks the workflow, so the pattern that scales is approving a policy once: a budget and a threshold under which the agent does not ask. Gokhan added that inside a company, most of the time the agent will not ask a person at all. It will ask another agent that has the right context, the way an employee asks accounting.
Internal agents are their own problem
Gokhan drew the line Brainbase is built on. External agents, the Harveys and Legoras, are high-IQ specialists: proprietary data, an A-team, post-training, all to be the best at one thing. Internal agents are the opposite. They are mid-IQ generalists running HR, accounting or KYC, where nobody has the data and nobody will staff a team to build it.

“In 10 years, companies are going to be spending two, three times more on agents they have inside than the agents they buy from the outside.”
CEO at Brainbase
What makes those agents work is not a bigger model but the right harness, context and ontology, and a benchmark built with the customer before anything changes, so everyone agrees what success means.
What comes next
The panel closed with a hot take each. Jon expects multi-agent systems that negotiate against each other rather than collaborate. Lauren expects that within six months you will be able to tell an agent to build, run and grow a business from one prompt and check in once a week, because the step changes in models and their cost keep arriving faster than anyone predicts.
Thais has become, in her words, more evals-pilled over time: agent capabilities follow our ability to measure them.

“The ability to measure, to me, follows the ability to define.”
CEO at Taste Labs
So deterministic problems get solved fast, and what is left is whatever is hard to define. The work that matters is deciding what good means and how to evaluate it. She added a smaller take: chat has won as one interface, but the format of an agent's output, structured for other agents and visual for people, is still unsolved.
Gokhan's take is the one Brainbase started on. In 2024 the market treated work as the last thing frontier models would automate; he expected it to be one of the first, with the frontier racing past it toward science. Most workflows are now easy for agents; what they lack is the primitives, like payments and runtimes, not intelligence. Automating labor has become a deployment problem rather than an intelligence problem. He expects frontier labs to drift toward government work, and automating work to feel less like a god in the machine and more like picking the right instance size, harness and context.

“I do not need to use the Millennium Prize model to do my HR.”
CEO at Brainbase
Thank you to Thais, Lauren and Jon for the conversation, to Cloudflare for hosting, and to everyone who built through the day. The full recording is on YouTube.
Related posts
@anthropic/claude-opus-5.5 is live on Brainbase
@openai/gpt-6-sol and @openai/gpt-6-luna are live on Brainbase
@spacexai/grok-4.7 is live on Brainbase
Automating Domino's local competitor pricing research
@openai/gpt-6-astra is live on Brainbase
@meta/muse-spark-1.3 is live on Brainbase
@google/gemini-3.8-flash is live on Brainbase
@anthropic/claude-fable-5.1 is live on Brainbase
@zai/glm-5.3-flash is live on Brainbase
Introducing Universal Managed Agents API
@spacexai/grok-4.6 is live on Brainbase
@meta/muse-spark is live on Brainbase
@thinking-machines/inkling is live on Brainbase
Brainbase is now generally available


