Introducing the Universal Managed Agents API.Read the announcement
Research

Choosing a System One model for your workflow

  • Joshua Okolo1,2
  • Gokhan Egri2

1Harvard University2Brainbase Labs

· 14 min read

TL;DR

  • We compiled τ²-bench customer-service agents into workflows and sent their typed decisions to fast System One models instead of a frontier LLM.
  • We built an adapter that turns any open chat model into a typed decision engine, then tested djev and five Qwen models from 0.8B to 35B.
  • The best model for each domain cut cost per conversation by 56 to 68% against the agent, and each decision ran about 5 to 7 times faster than an LLM call.
  • Qwen3.5-4B was the pick for retail and telecom and djev for airline, so the model is chosen per workload.
lower cost than the agent in retail, with the 4B
68%
lower cost than the agent in retail, with the 4B
lower cost in telecom, with more tasks passed
63%
lower cost in telecom, with more tasks passed
per typed decision on the 4B, against 1.2 to 1.9 s for an LLM
275 ms
per typed decision on the 4B, against 1.2 to 1.9 s for an LLM
4B agreement with LLM decisions on held-out tasks
0.927
4B agreement with LLM decisions on held-out tasks

1Small decisions, big models

TypeSafe introduced System One models with the release of Jev, its decision model. A System One model takes a state and a closed question, and returns a probability for every possible answer in a few hundred milliseconds.

TypeSafe argues that general capability will come faster from fast decision models with strong typings than from agents that reason for longer. Slow reasoning still has a place, and it runs only where it is needed.

An LLM agent makes dozens of small decisions in a single customer-service conversation. It has to work out which order the customer means, whether an item belongs in the return and whether the customer has confirmed.

Each of those decisions goes through a frontier model that takes more than a second to answer and bills for every token. Most of them are closed questions with a handful of possible answers.

We wanted to test that on real multi-turn workflows. If it holds, a small typed model should answer most of an agent's decisions, and the agent should get cheaper and faster without getting worse.

2The setup

τ²-bench simulates customer-service conversations in retail, airline and telecom. An LLM plays the customer, and the agent has to follow a policy and use tools to resolve the request.

Each conversation is graded on the final database state, the actions taken and, for some tasks, natural-language assertions. We used Claude Sonnet 4.6 as the judge for those assertions in every arm, so our scores are for comparing the arms with each other and are not leaderboard numbers.

We compiled each domain's agent into a workflow. Code handles the tool calls and lookups, a small Claude Haiku 4.5 model pulls details such as names and order numbers out of customer messages and writes the replies, and every branch point becomes a typed question for a decision engine.

When a conversation goes somewhere the workflow does not cover, it hands off to the full agent. Every arm runs the same frozen workflow on the same 40 held-out tasks, twice each, and only the decision engine changes.

Here is one held-out retail conversation run with Qwen3.5-4B. The customer asked to change the colour of the wireless earbuds in an order to blue, and the workflow asked six typed questions along the way.

  1. 1

    Route the intent

    What is the customer's current request? 8 options

    modify order

    p 0.83 · 313 ms

  2. 2

    What to modify

    Which part of the pending order does the customer want to change? 3 options

    items

    p 0.998 · 270 ms

  3. 3

    Which item

    Which item does the customer want to modify? 4 items on the order

    wireless earbuds, white

    p 0.997 · 230 ms

  4. 4

    New variant

    Which new variant does the customer want? 7 variants

    blue, 8 hour battery

    p 0.998 · 327 ms

  5. 5

    Payment

    Which payment method covers the price difference? 2 options

    PayPal

    p 0.731 · 575 ms

  6. 6

    Confirmation

    Did the customer explicitly confirm the change? 2 options

    yes

    p 0.99 · 324 ms

Code turned those six answers into a single call to modify the order, and the conversation passed. Each answer came back in a few hundred milliseconds with a probability attached.

3Starting with diffusion

Our first engine was djev, a Jev adapter on Google's DiffusionGemma. A diffusion model denoises every output position at once, so a set of short typed answers costs about one pass.

djev speaks the System One API. The workflow sends the state and a set of questions with named choices, and gets back a probability for every choice.

POST /v1/systemone
{"state": {"context": "...conversation + tool results..."},
 "questions": {"route_intent": {"type": "choice", "instructions": "...",
                                "criteria": {"return": "...", "exchange": "...", "info_lookup": "..."}}}}
-> {"route_intent": {"return": 0.94, "exchange": 0.05, "info_lookup": 0.01}}

We replayed every logged retail decision against djev, and it agreed with the LLM on 89% of them. Its confidence meant something too, since the decisions it was surest of agreed 98% of the time.

That made a confidence gate possible, where djev answers when it is sure and hands the question to an LLM when it is not. With the gate at 0.97, the hybrid matched the LLM workflow end to end at 25% lower cost.

4Adapters for open models

djev is one model, and anyone building a system has to choose which model goes in the slot. So we built the same interface on top of open models served with vLLM.

The adapter turns each typed question into a prompt with the state, the question and options lettered A, B and C. The model answers with a single letter, with thinking turned off.

We read the probabilities of that one token and keep only the valid letters. The probability of option k is a softmax over the logits z of the option letters alone.

p(k∣x)=exp⁡(zk)∑j∈optionsexp⁡(zj)p(k \mid x) = \frac{\exp(z_k)}{\sum_{j \in \text{options}} \exp(z_j)}
(1)

One forward pass gives a probability for every choice, computed the same way at every model size. The same adapter serves Qwen3.5-0.8B, 2B and 4B, Qwen3.8-27B and Qwen3.6-35B-A3B, so every model sees the identical question in the identical format.

5Choosing the model

We did not want one model forced onto every workflow, so we made the choice itself something to optimise. The rule looks at how many typed decisions a workflow makes per conversation and at the success rate it needs.

It then picks the cheapest model and gate threshold that should reach that rate. Agreement comes from replaying the workflow's own decision log through each candidate.

S^(m,t)=SLLM−k⋅d⋅(1−am(t))\hat{S}(m, t) = S_{\text{LLM}} - k \cdot d \cdot \big(1 - a_m(t)\big)
(2)
(m∗,t∗)=arg⁡min⁡m, t  cost(m,t)subject toS^(m,t)≥Starget(m^*, t^*) = \arg\min_{m,\,t} \; \text{cost}(m, t) \quad \text{subject to} \quad \hat{S}(m, t) \ge S_{\text{target}}
(3)

Here dd is the number of typed decisions per conversation and am(t)a_m(t) is model mm's agreement with the LLM when the gate escalates below threshold tt. SLLMS_{\text{LLM}} is the pass rate with LLM decisions, and kk is fitted on the djev arms.

The obvious model, agreement raised to the number of decisions, was far too pessimistic. Many wrong decisions never change the outcome or get corrected later in the conversation, and the linear penalty above predicted end-to-end pass rates to within 0.06 on average.

We chose the model and threshold for each domain on the dev split, before running anything on the held-out tasks.

6How the models compare

To compare the models as decision engines, we replayed 4,355 logged decisions through each one with identical state and questions. That separates the model from the noise of a live conversation.

Figure 1. Agreement with the LLM's decisions on held-out ledgers, by model size and domain.
0.40.50.60.70.80.91.00.8B2B4B27B35B-A3BModel sizedjev (dev ledgers) 0.93
  • Retail
  • Airline
  • Telecom
  • All domains

Agreement with LLM decisions

Agreement jumps between 2B and 4B and then flattens. The 0.8B and 2B pick the first option far too often, while the 4B, 27B and 35B-A3B land within a few points of each other.

The 27B leads by about two points and takes 3.3 times longer per decision than the 4B. The 35B-A3B scores the same as the 4B and runs slower.

Figure 2. Left, calibration of the top-choice probability. Right, the confidence gate, trading decisions sent to an LLM against agreement.

Calibration, held-out, all domains

0.00.20.40.60.81.00.30.40.50.60.70.80.91.0Model confidence (top choice)

Agreement with LLM

Confidence gate: escalate when unsure

0.40.50.60.70.80.91.00.00.20.40.60.8Share of decisions escalated to the LLM

Agreement of the hybrid

  • Qwen3.5-0.8B · ECE 0.138
  • Qwen3.5-2B · ECE 0.083
  • Qwen3.5-4B · ECE 0.036
  • Qwen3.8-27B · ECE 0.026
  • Qwen3.6-35B · ECE 0.021

Confidence splits the same way. The small models' probabilities say little about whether they are right, and the 4B and larger are well calibrated.

A well-calibrated 4B also leaves the gate little to do. At 93% agreement, sending its unsure decisions to an LLM made no difference we could measure.

6.1Which decisions are hard

Averages hide where the mistakes are. Broken down by decision type, the 4B agrees with the LLM on nearly every yes-or-no and state-reading question, and most of its disagreements sit in one place.

Figure 3. Qwen3.5-4B agreement with the LLM by decision type on held-out ledgers, for types with enough decisions to measure. Dot size grows with the number of decisions, and routing is highlighted.
  • Telecom · VPN state
    100%
  • Telecom · data saver state
    100%
  • Telecom · network mode state
    100%
  • Telecom · mobile data state
    100%
  • Telecom · SIM state
    100%
  • Telecom · airplane mode state
    98%
  • Telecom · issue type
    97%
  • Telecom · route the intent
    75%
  • Airline · more than one request
    97%
  • Airline · route the intent
    89%
  • Retail · which order
    100%
  • Retail · item in the return
    100%
  • Retail · what to modify
    97%
  • Retail · item in the exchange
    95%
  • Retail · what to look up
    94%
  • Retail · route the intent
    73%
50%60%70%80%90%100%

4B agreement with the LLM

Routing the customer's request is the hardest of the common decisions in every domain. The 4B agreed with the LLM on 73% of retail routing decisions and 75% in telecom, against 95 to 100% on whether an item belongs in a return or exchange.

A customer's message can often be read as more than one kind of request, and routing has to commit to one. It usually comes first in a conversation, so it is the natural place to spend a gate or a bigger model.

The smallest models fail in a different way. In airline the 0.8B and 2B chose the first option on about two thirds of decisions where the LLM did so 18% of the time, while the 4B and larger matched the LLM's rate.

7Results by domain

7.1Retail

A retail conversation takes about 7 typed decisions. The workflow routes the intent, picks the order, decides which items a return or exchange covers and confirms.

Figure 4. Retail, held-out tasks. The pick is highlighted.

Held-out pass^1 · bar is the pass rate · right is Anthropic $ per conversation

  • Sonnet agent (no workflow)$0.228 / conversation

    86.3%
  • Workflow · LLM decisions$0.153 / conversation

    68.8%
  • Workflow · djev$0.078 / conversation

    56.2%
  • Workflow · djev + gate$0.113 / conversation

    70.0%
  • Workflow · Qwen3.5-4B$0.073 / conversation

    68.8%
  • Workflow · Qwen3.5-4B + gate$0.082 / conversation

    62.5%
  • Workflow · Qwen3.8-27B$0.088 / conversation

    70.0%
  • Workflow · Qwen3.5-2B$0.078 / conversation

    37.5%

Qwen3.5-4B matched the LLM workflow at half its cost, and came in 68% below the agent. djev on its own passed fewer tasks, and its hybrid won them back at a higher price.

The remaining gap to the agent comes from what the workflow covers. The LLM workflow shows the same gap.

7.2Airline

Airline is the deepest workflow, with about 11 decisions per conversation. Many of them are judgements about eligibility and fare rules.

Figure 5. Airline, held-out tasks. The pick is highlighted.

Held-out pass^1 · bar is the pass rate · right is Anthropic $ per conversation

  • Sonnet agent (no workflow)$0.200 / conversation

    92.5%
  • Workflow · LLM decisions$0.140 / conversation

    70.0%
  • Workflow · djev$0.087 / conversation

    70.0%
  • Workflow · djev + gate$0.106 / conversation

    70.0%
  • Workflow · Qwen3.5-4B$0.091 / conversation

    60.0%
  • Workflow · Qwen3.5-4B + gate$0.091 / conversation

    62.5%

djev is the pick here. It matched the LLM workflow for 38% less, while the 4B passed fewer tasks.

The 4B agrees with the LLM on 94% of airline decisions, but its confidence barely separates its right answers from its wrong ones, so the gate cannot catch its mistakes. The best model for retail is the wrong model for airline.

7.3Telecom

Telecom troubleshooting has fewer closed decisions per turn. Its conversations run long and expensive, and the agent itself struggles.

Figure 6. Telecom, held-out tasks. The pick is highlighted.

Held-out pass^1 · bar is the pass rate · right is Anthropic $ per conversation

  • Sonnet agent (no workflow)$0.390 / conversation

    80.0%
  • Workflow · LLM decisions$0.168 / conversation

    91.2%
  • Workflow · djev$0.125 / conversation

    85.0%
  • Workflow · djev + gate$0.124 / conversation

    86.3%
  • Workflow · Qwen3.5-4B$0.145 / conversation

    90.0%
  • Workflow · Qwen3.5-4B + gate$0.137 / conversation

    87.5%

Here the workflow beat the agent, and the 4B passed more tasks than the agent at 63% lower cost.

The LLM workflow beat the agent as well, so the gain comes from the workflow's fixed troubleshooting order. The 4B keeps that gain for less.

8Cost and speed

Figure 7. Every arm on held-out tasks, pass rate against Anthropic spend per conversation. User-simulator cost is the same for every arm and left out. Up and to the left is better.

Retail

0.30.50.70.90.00.10.20.3$ per conversationSonnet agent (no workflow): 86.3%, $0.228Workflow · LLM decisions: 68.8%, $0.153Workflow · djev + gate: 70.0%, $0.113Workflow · djev: 56.2%, $0.078Workflow · Qwen3.5-2B: 37.5%, $0.078Workflow · Qwen3.5-4B: 68.8%, $0.073Workflow · Qwen3.5-4B + gate: 62.5%, $0.082Workflow · Qwen3.8-27B: 70.0%, $0.088

pass^1 (held-out)

Airline

0.30.50.70.90.00.10.20.3$ per conversationSonnet agent (no workflow): 92.5%, $0.200Workflow · LLM decisions: 70.0%, $0.140Workflow · djev + gate: 70.0%, $0.106Workflow · djev: 70.0%, $0.087Workflow · Qwen3.5-4B: 60.0%, $0.091Workflow · Qwen3.5-4B + gate: 62.5%, $0.091

pass^1 (held-out)

Telecom

0.30.50.70.90.00.20.4$ per conversationSonnet agent (no workflow): 80.0%, $0.390Workflow · LLM decisions: 91.2%, $0.168Workflow · djev + gate: 86.3%, $0.124Workflow · djev: 85.0%, $0.125Workflow · Qwen3.5-4B: 90.0%, $0.145Workflow · Qwen3.5-4B + gate: 87.5%, $0.137

pass^1 (held-out)

  • Sonnet agent (no workflow)
  • Workflow · LLM decisions
  • Workflow · djev
  • Workflow · djev + gate
  • Workflow · Qwen3.5-4B
  • Workflow · Qwen3.5-4B + gate
  • Workflow · Qwen3.8-27B
  • Workflow · Qwen3.5-2B

Per decision, the 4B answered in 275 ms against 1.2 to 1.9 seconds for Sonnet, and cost about 100 times less.

Per conversation the saving is smaller, because escalations to the full agent cost what they always did. Whole-conversation wall time barely changed, since the user simulator and the escalations take most of it.

What remains after the swap is almost entirely the full agent. Conversations handed off to it account for 94% of the 4B arm's remaining cost in retail and 99% in telecom, and the Haiku calls make up the rest.

So the next saving comes from widening what the workflow covers. A conversation the workflow finishes on its own costs a fraction of one it hands to the agent.

Figure 8. Pass rate, cost per conversation and decision latency for every arm and domain. Drag to rotate. Arms relaunched at 40 to 80 concurrent conversations have inflated latency and are marked on hover.
$0.08$0.12$0.16200100050000.400.600.80$ / conversationDecision latency (ms, log)pass^1
  • Retail
  • Airline
  • Telecom
  • Drag to rotate

9Which workloads suit a System One model

The saving grows with the share of an agent's work that is closed decisions. In τ² that share is high, with 27 to 41% of agent turns and 7 to 12 typed decisions per conversation.

Policy-driven service flows fit well, with their routing, eligibility checks, confirmations and structured troubleshooting. Open-ended drafting and analysis fit poorly, because there the output is the work itself.

In every domain the hardest decisions were routing an ambiguous request and working out which order or item the customer meant. Those are where a gate and a bigger model earn their cost.

10Applying this to your own agent

Everything here can be repeated on another agent. Start by logging its decisions and finding the ones that are closed questions with a fixed set of answers.

Compile those into a workflow, with code for the plumbing and a typed question at every branch. Leave anything open-ended to the full agent.

Then replay the logged decisions through each candidate model. Agreement shows which models can do the job, calibration shows whether a gate can catch their mistakes, and the budget picks among the rest.

Repeat the replay when the workflow or the models change. A new open model is one replay away from a place in the slot.

11Conclusions

System One models make an agent's closed decisions zero-shot, in a few hundred milliseconds each, with no model trained for the workflow. Adapting to a new workflow takes a replay of its decisions against the candidates, where a special-purpose model would take a dataset and a training run.

In our tests that kept the LLM workflow's accuracy and cut cost per conversation by 56 to 68% against the agent. The largest model never won, with a 4B matching models seven to nine times its size in retail and telecom and djev taking airline.

Any team running an agent already has what this takes, a log of its decisions and a shelf of open models. Replaying one against the other shows how much of the agent a System One model can take over.

Appendix

A.1Agreement by model

Table 1. All held-out decision ledgers, replayed with identical state and questions.
ModelAgreement95% CIECEAUROC
Qwen3.5-0.8B0.612[0.57, 0.65]0.1380.67
Qwen3.5-2B0.709[0.66, 0.75]0.0830.80
Qwen3.5-4B0.927[0.91, 0.94]0.0360.82
Qwen3.8-27B0.948[0.94, 0.96]0.0260.89
Qwen3.6-35B-A3B-FP80.931[0.92, 0.95]0.0210.85

A.2End-to-end results

Held-out tasks, two trials each. Escalated is the share of tasks handed to the full agent.

Table 2. Retail.
Armpass^1pass^2$ / conversationEscalated
Sonnet agent (no workflow)0.8630.800$0.228–
Workflow · LLM decisions0.6880.575$0.15343.8%
Workflow · djev0.5620.475$0.07838.8%
Workflow · djev + gate0.7000.550$0.11342.5%
Workflow · Qwen3.5-4B0.6880.600$0.07337.5%
Workflow · Qwen3.5-4B + gate0.6250.550$0.08237.5%
Workflow · Qwen3.8-27B0.7000.600$0.08840.0%
Workflow · Qwen3.5-2B0.3750.275$0.07832.5%
Table 3. Airline.
Armpass^1pass^2$ / conversationEscalated
Sonnet agent (no workflow)0.9250.900$0.200–
Workflow · LLM decisions0.7000.650$0.14047.5%
Workflow · djev0.7000.650$0.08740.0%
Workflow · djev + gate0.7000.650$0.10645.0%
Workflow · Qwen3.5-4B0.6000.550$0.09139.5%
Workflow · Qwen3.5-4B + gate0.6250.550$0.09136.4%
Table 4. Telecom.
Armpass^1pass^2$ / conversationEscalated
Sonnet agent (no workflow)0.8000.700$0.390–
Workflow · LLM decisions0.9120.850$0.16885.0%
Workflow · djev0.8500.725$0.12583.8%
Workflow · djev + gate0.8630.775$0.12482.5%
Workflow · Qwen3.5-4B0.9000.825$0.14597.5%
Workflow · Qwen3.5-4B + gate0.8750.850$0.13797.5%

A.3Cost per decision

Table 5. Per typed decision.
EngineLatency p50Cost
LLM (Sonnet 4.6, constrained JSON)1.2 to 1.9 s$0.0045 to 0.0104
djev0.14 to 0.35 sself-hosted
Qwen3.5-4B, cold / warm prefix cache (A100)275 / 108 ms$0.00007

A.4Notes

The adapter records how much probability the model put on any valid letter before renormalising, and for the 4B that averaged 0.999. Every run checks that decisions were logged and that the engine reported no errors, after an early djev smoke test passed on escalations alone while every call had timed out.

Each arm has 80 conversations per domain, which gives intervals of about ten points either way.

References

  1. [1]Barres, V., Dong, H., Ray, S., Si, X., Narasimhan, K.. τ²-Bench: Evaluating conversational agents in a dual-control environment. arXiv preprint arXiv:2506.07982, 2025.