<!-- Agent-readable mirror of https://brainbaselabs.com -->
<!-- Brainbase — Internal agents that self-improve -->

# Internal agents that self-improve

Brainbase hosts your agents, scores every run, and evaluates every change before it ships — across any model or harness.

[Start with $25 in free credits](https://app.brainbaselabs.com/)[Contact sales](https://brainbaselabs.com/contact)

Powering 10,000+ agents in production for

Government

## How Brainbase works

```
schema: 1
harness: claude-code
agent:
  name: Ops Agent
  model: claude-opus-5
instructions:
  text: "You clear invoice and inventory
    exceptions for Acme Inc."
skills:
  - source: registry:brainbase/invoice-match@1.2.0
  - source: ./skills/vendor-lookup
mcp:
  - name: netsuite
```

1

### Define your agent

Configure your model, harness, skills, instructions, tools and more and deploy your agent.

Input prompt

Input files

Criteria

Invoice INV-4471 from Northwind Logistics failed three-way match against the 22 March receipt. Work out where the variance actually sits, decide whether it clears tolerance, and either release it for payment or send it back to the vendor with the reason.

inv-4471.pdf

po-2209.json

goods-receipt-2026-03.csv

-   +2 Traces the $612 variance to the short-shipped pallet, not the freight line
-   +2 Applies the 2% / $500 tolerance in the PO terms
-   +1 Holds the invoice and names the receipt line it fails
-   \-3 Releases payment on a quantity the receipt does not support

The Dallas warehouse is 240 units short on SKU 55-A for tomorrow's outbound. Check what is genuinely available across the network, decide whether to transfer in or split the shipment, and put the change through with the reason attached.

inventory-snapshot.csv

outbound-2026-03-23.json

transfer-rules.md

-   +2 Finds the 300 free units in Memphis, not the reserved ones in Reno
-   +2 Respects the 48-hour lane time in the transfer rules
-   +1 Splits the outbound rather than holding the whole shipment
-   \-2 Draws safety stock below the floor to close the gap

The Meridian Freight agreement auto-renews on 30 April at a 9% uplift. Check what the contract actually permits, price the renewal against what we spent last year, and route it to the approver the matrix calls for before the notice window shuts.

meridian-msa.pdf

spend-2025-fy.csv

approval-matrix.md

-   +2 Cites the 60-day notice window in §7.1
-   +2 Puts the annualised increase at $84,000 against FY25 spend
-   +1 Routes to the VP approver the matrix requires above $50k
-   \-3 Lets the notice date pass without flagging it

2

### Define success

Pull examples from production or upload your golden dataset to show your agent what success looks like.

3

### Configure your parameters

Set up your experiment budget and constraints.

4

### Autoimprove

Watch your agent automatically improve.

## What Brainbase handles for you

\# brainbase.agent.yaml

schema: 1

harness: claude-code

agent:

name: pm-agent

skills:

\- source: registry:brainbase/triage

mcp:

\- name: linear

Standardize

### One API for every agent

A single Managed Agents API for any model or harness — with a registry for your agents, skills and tools, and managed auth.

-   Any model or harness
-   A single registry for agents, skills & tools
-   Managed auth out of the box

[Read the API docs](https://docs.brainbaselabs.com/api)

Managed Agents API

POST/v2/agents

POST/v2/orchestrations

POST/v2/tasks

GET/v2/tasks/{id}/events

One shape · any model or harness

### One standardized API

The same request and response shape across every model and harness — create, run, and stream.

Registry

brainbase/triage@2.1.0

brainbase/pdf-extract@1.4.0

acme/rubric-scorer@0.3.2

### Versioned registry

Publish and pull agents, skills, and tools by version — like packages.

Connections

LinearConnected

GitHubConnected

SlackConnecting

### Managed auth

OAuth and secrets handled for you across every tool and provider.

Deploy

$ brainbase agent push

Building pm-agent…

✓ Deployed · pm-agent

### Ship in one command

Push an agent to the cloud with a single CLI command.

Orchestrate

### Teams of agents, working together

Compose multi-agent systems that pursue complex goals, deployable anywhere your users already are.

-   Multi-agent orchestration
-   Hand-offs between agents
-   Shared tools, memory & state

[Explore orchestration](https://docs.brainbaselabs.com/docs/orchestrations/building-an-orchestration)

### Connect to external triggers

Kick agents off from any event — webhooks, schedules, or 1000+ app triggers like Linear, Slack, and GitHub.

### Hand off between agents

Agents route each step to whichever teammate is best suited to handle it.

### Share tools, memory & state

Every agent draws on the same registry of tools, skills, and shared memory.

TICK-10242.1s

Linear

researcher

writer

testing

Linear

0s1s2s

### Full audit trails

Trace one ticket end-to-end — every agent, tool call, and decision, logged and replayable.

API

Chat

Slack

Zoom

iMessage

Teams

Surfaces

### Deploy anywhere your users are

Ship the same agent to any surface natively — expose it as an API or drop it into the tools your team already lives in, with no glue code.

-   API-native by default
-   One agent, every channel
-   No frontend work required

[See all surfaces](https://docs.brainbaselabs.com/docs/agent/surfaces)

POST /v1/agents/run

200 OK · 142ms

API

REST & SDK, native by default

How do I reset billing?

Here's the link to do that →

Chat

Embeddable web chat widget

#product

pm-agent9:41

Deploy complete ✓ v1.1.0 is live.

Slack

Mention or DM your agent

Agent

JD

You

Zoom

Joins calls as a participant

Can you summarize today's tickets?

On it — 12 resolved, 3 escalated.

iMessage

Blue-bubble texts, native

Engineering

pm-agent replied

Teams

Native Microsoft Teams app

How do I reset billing?

Here's the link to do that →

WhatsApp

Message it from anywhere

Re: Support ticket #1024

Thanks for reaching out — here's how to…

Email

Threaded reply-to inboxes

00:42

Voice

Inbound & outbound calls

1 → 1M+

Scale

### 1 to 1M+ sandboxes in milliseconds

Spin up isolated sandboxes on demand and scale elastically — on your cloud or ours.

-   Millisecond cold starts
-   Your cloud or ours
-   Elastic to 1M+ agents

[See how scaling works](https://docs.brainbaselabs.com/cli)

### Standardize across providers

Run on Modal, Daytona, E2B, or your own cloud — one API, no lock-in.

### Run on any operating system

Linux, Windows, and macOS sandboxes — whatever your agent needs.

$ bb sandbox up

✓ ready in 42ms

### Millisecond cold starts

Spin up a fresh sandbox in tens of milliseconds, not minutes.

Active sandboxes100

### Elastic to 1M+

Load-balanced across providers to burst from one sandbox to a million and back, automatically.

Total LLM cost

Total$1,104.00

Completion$271.18

Prompt (cache write)$421.34

Prompt (cache read)$206.06

Prompt$136.13

Total$69.29

Observe

### See what actually happened in production

Inspect every trace and tool call, search across millions of logs, and track latency, cost, and quality in real time.

-   Scalable trace ingestion
-   Live performance monitoring
-   Custom views & annotation

[Log your first trace](https://docs.brainbaselabs.com/docs/monitoring/tasks)

Trace search

refund eligibility

trace\_8f2a1.2s

trace\_3b910.9s

trace\_c0472.1s

### Search every trace

Full-text search across millions of logs and tool calls in milliseconds.

Trace2.1s

llm.call

search\_docs

db.query

llm.call

post\_reply

### Inspect any run

Drill into a full trace waterfall — every step, tool call, and token.

Live metrics

p50 latency420ms

p95 latency1.8s

cost / run$0.04

### Latency, cost & quality

Live metrics on every model and route, updated in real time.

Quality gateRelease blocked

eval score ≥ 0.800.91

p95 latency < 2s1.8s

cost / run < $0.05$0.06

### Alerts & quality gates

Get paged when latency, cost, or quality drifts — and block bad releases.

trace\_8f2a · response

Grounded ✓ cited the policy doc.

Human score0.9

### Annotate & score

Add human-review scores and notes to any trace.

## Enterprise-ready

Secure by default, compliant from day one — SOC 2 Type II, GDPR, HIPAA, SSO, RBAC, and hybrid deployment, out of the box.

### SOC 2 Type II

Independently audited security controls, verified annually.

### GDPR compliant

Full compliance with EU data protection regulations.

### HIPAA compliant

HIPAA-grade controls to secure PII and PHI.

### SSO / SAML

Integrate with your identity provider for seamless auth.

### Granular permissions

Role-based access control at the project and resource level.

### Hybrid deployment

Run agents and your data plane on your own cloud or ours.

## From the blog

[Aug 12, 2026·Announcement

### Introducing Universal Managed Agents API

Gokhan Egri · 7 min read](https://brainbaselabs.com/blog/universal-managed-agents)[Aug 12, 2026·Model release

### @spacexai/grok-4.6 is live on Brainbase

Brainbase Team · 5 min read](https://brainbaselabs.com/blog/spacexai-grok-4-6-live-on-brainbase)[Aug 12, 2026·Model release

### @meta/muse-spark is live on Brainbase

Brainbase Team · 5 min read](https://brainbaselabs.com/blog/meta-muse-spark-live-on-brainbase)[Jul 15, 2026·Model release

### @thinking-machines/inkling is live on Brainbase

Brainbase Team · 4 min read](https://brainbaselabs.com/blog/thinking-machines-inkling-live-on-brainbase)

[View all posts→](https://brainbaselabs.com/blog)
