Choosing a System One model for your workflow
- Joshua Okolo1,2
- Gokhan Egri2
1Harvard University2Brainbase Labs
· 14 min read
TL;DR
- We compiled τ²-bench customer-service agents into workflows and sent their typed decisions to fast System One models instead of a frontier LLM.
- We built an adapter that turns any open chat model into a typed decision engine, then tested djev and five Qwen models from 0.8B to 35B.
- The best model for each domain cut cost per conversation by 56 to 68% against the agent, and each decision ran about 5 to 7 times faster than an LLM call.
- Qwen3.5-4B was the pick for retail and telecom and djev for airline, so the model is chosen per workload.
- lower cost than the agent in retail, with the 4B
- 68%
- lower cost than the agent in retail, with the 4B
- lower cost in telecom, with more tasks passed
- 63%
- lower cost in telecom, with more tasks passed
- per typed decision on the 4B, against 1.2 to 1.9 s for an LLM
- 275 ms
- per typed decision on the 4B, against 1.2 to 1.9 s for an LLM
- 4B agreement with LLM decisions on held-out tasks
- 0.927
- 4B agreement with LLM decisions on held-out tasks
1Small decisions, big models
TypeSafe introduced System One models with the release of Jev, its decision model. A System One model takes a state and a closed question, and returns a probability for every possible answer in a few hundred milliseconds.
TypeSafe argues that general capability will come faster from fast decision models with strong typings than from agents that reason for longer. Slow reasoning still has a place, and it runs only where it is needed.
An LLM agent makes dozens of small decisions in a single customer-service conversation. It has to work out which order the customer means, whether an item belongs in the return and whether the customer has confirmed.
Each of those decisions goes through a frontier model that takes more than a second to answer and bills for every token. Most of them are closed questions with a handful of possible answers.
We wanted to test that on real multi-turn workflows. If it holds, a small typed model should answer most of an agent's decisions, and the agent should get cheaper and faster without getting worse.
2The setup
τ²-bench simulates customer-service conversations in retail, airline and telecom. An LLM plays the customer, and the agent has to follow a policy and use tools to resolve the request.
Each conversation is graded on the final database state, the actions taken and, for some tasks, natural-language assertions. We used Claude Sonnet 4.6 as the judge for those assertions in every arm, so our scores are for comparing the arms with each other and are not leaderboard numbers.
We compiled each domain's agent into a workflow. Code handles the tool calls and lookups, a small Claude Haiku 4.5 model pulls details such as names and order numbers out of customer messages and writes the replies, and every branch point becomes a typed question for a decision engine.
When a conversation goes somewhere the workflow does not cover, it hands off to the full agent. Every arm runs the same frozen workflow on the same 40 held-out tasks, twice each, and only the decision engine changes.
Here is one held-out retail conversation run with Qwen3.5-4B. The customer asked to change the colour of the wireless earbuds in an order to blue, and the workflow asked six typed questions along the way.
- 1
Route the intent
What is the customer's current request? 8 options
modify order
p 0.83 · 313 ms
- 2
What to modify
Which part of the pending order does the customer want to change? 3 options
items
p 0.998 · 270 ms
- 3
Which item
Which item does the customer want to modify? 4 items on the order
wireless earbuds, white
p 0.997 · 230 ms
- 4
New variant
Which new variant does the customer want? 7 variants
blue, 8 hour battery
p 0.998 · 327 ms
- 5
Payment
Which payment method covers the price difference? 2 options
PayPal
p 0.731 · 575 ms
- 6
Confirmation
Did the customer explicitly confirm the change? 2 options
yes
p 0.99 · 324 ms
Code turned those six answers into a single call to modify the order, and the conversation passed. Each answer came back in a few hundred milliseconds with a probability attached.
3Starting with diffusion
Our first engine was djev, a Jev adapter on Google's DiffusionGemma. A diffusion model denoises every output position at once, so a set of short typed answers costs about one pass.
djev speaks the System One API. The workflow sends the state and a set of questions with named choices, and gets back a probability for every choice.
POST /v1/systemone
{"state": {"context": "...conversation + tool results..."},
"questions": {"route_intent": {"type": "choice", "instructions": "...",
"criteria": {"return": "...", "exchange": "...", "info_lookup": "..."}}}}
-> {"route_intent": {"return": 0.94, "exchange": 0.05, "info_lookup": 0.01}}We replayed every logged retail decision against djev, and it agreed with the LLM on 89% of them. Its confidence meant something too, since the decisions it was surest of agreed 98% of the time.
That made a confidence gate possible, where djev answers when it is sure and hands the question to an LLM when it is not. With the gate at 0.97, the hybrid matched the LLM workflow end to end at 25% lower cost.
4Adapters for open models
djev is one model, and anyone building a system has to choose which model goes in the slot. So we built the same interface on top of open models served with vLLM.
The adapter turns each typed question into a prompt with the state, the question and options lettered A, B and C. The model answers with a single letter, with thinking turned off.
We read the probabilities of that one token and keep only the valid letters. The probability of option k is a softmax over the logits z of the option letters alone.
One forward pass gives a probability for every choice, computed the same way at every model size. The same adapter serves Qwen3.5-0.8B, 2B and 4B, Qwen3.8-27B and Qwen3.6-35B-A3B, so every model sees the identical question in the identical format.
5Choosing the model
We did not want one model forced onto every workflow, so we made the choice itself something to optimise. The rule looks at how many typed decisions a workflow makes per conversation and at the success rate it needs.
It then picks the cheapest model and gate threshold that should reach that rate. Agreement comes from replaying the workflow's own decision log through each candidate.
Here is the number of typed decisions per conversation and is model 's agreement with the LLM when the gate escalates below threshold . is the pass rate with LLM decisions, and is fitted on the djev arms.
The obvious model, agreement raised to the number of decisions, was far too pessimistic. Many wrong decisions never change the outcome or get corrected later in the conversation, and the linear penalty above predicted end-to-end pass rates to within 0.06 on average.
We chose the model and threshold for each domain on the dev split, before running anything on the held-out tasks.
6How the models compare
To compare the models as decision engines, we replayed 4,355 logged decisions through each one with identical state and questions. That separates the model from the noise of a live conversation.
- Retail
- Airline
- Telecom
- All domains
Agreement with LLM decisions
Agreement jumps between 2B and 4B and then flattens. The 0.8B and 2B pick the first option far too often, while the 4B, 27B and 35B-A3B land within a few points of each other.
The 27B leads by about two points and takes 3.3 times longer per decision than the 4B. The 35B-A3B scores the same as the 4B and runs slower.
Calibration, held-out, all domains
Agreement with LLM
Confidence gate: escalate when unsure
Agreement of the hybrid
- Qwen3.5-0.8B · ECE 0.138
- Qwen3.5-2B · ECE 0.083
- Qwen3.5-4B · ECE 0.036
- Qwen3.8-27B · ECE 0.026
- Qwen3.6-35B · ECE 0.021
Confidence splits the same way. The small models' probabilities say little about whether they are right, and the 4B and larger are well calibrated.
A well-calibrated 4B also leaves the gate little to do. At 93% agreement, sending its unsure decisions to an LLM made no difference we could measure.
6.1Which decisions are hard
Averages hide where the mistakes are. Broken down by decision type, the 4B agrees with the LLM on nearly every yes-or-no and state-reading question, and most of its disagreements sit in one place.
- Telecom · VPN state100%
- Telecom · data saver state100%
- Telecom · network mode state100%
- Telecom · mobile data state100%
- Telecom · SIM state100%
- Telecom · airplane mode state98%
- Telecom · issue type97%
- Telecom · route the intent75%
- Airline · more than one request97%
- Airline · route the intent89%
- Retail · which order100%
- Retail · item in the return100%
- Retail · what to modify97%
- Retail · item in the exchange95%
- Retail · what to look up94%
- Retail · route the intent73%
4B agreement with the LLM
Routing the customer's request is the hardest of the common decisions in every domain. The 4B agreed with the LLM on 73% of retail routing decisions and 75% in telecom, against 95 to 100% on whether an item belongs in a return or exchange.
A customer's message can often be read as more than one kind of request, and routing has to commit to one. It usually comes first in a conversation, so it is the natural place to spend a gate or a bigger model.
The smallest models fail in a different way. In airline the 0.8B and 2B chose the first option on about two thirds of decisions where the LLM did so 18% of the time, while the 4B and larger matched the LLM's rate.
7Results by domain
7.1Retail
A retail conversation takes about 7 typed decisions. The workflow routes the intent, picks the order, decides which items a return or exchange covers and confirms.
Held-out pass^1 · bar is the pass rate · right is Anthropic $ per conversation
Sonnet agent (no workflow)$0.228 / conversation
86.3%Workflow · LLM decisions$0.153 / conversation
68.8%Workflow · djev$0.078 / conversation
56.2%Workflow · djev + gate$0.113 / conversation
70.0%Workflow · Qwen3.5-4B$0.073 / conversation
68.8%Workflow · Qwen3.5-4B + gate$0.082 / conversation
62.5%Workflow · Qwen3.8-27B$0.088 / conversation
70.0%Workflow · Qwen3.5-2B$0.078 / conversation
37.5%
Qwen3.5-4B matched the LLM workflow at half its cost, and came in 68% below the agent. djev on its own passed fewer tasks, and its hybrid won them back at a higher price.
The remaining gap to the agent comes from what the workflow covers. The LLM workflow shows the same gap.
7.2Airline
Airline is the deepest workflow, with about 11 decisions per conversation. Many of them are judgements about eligibility and fare rules.
Held-out pass^1 · bar is the pass rate · right is Anthropic $ per conversation
Sonnet agent (no workflow)$0.200 / conversation
92.5%Workflow · LLM decisions$0.140 / conversation
70.0%Workflow · djev$0.087 / conversation
70.0%Workflow · djev + gate$0.106 / conversation
70.0%Workflow · Qwen3.5-4B$0.091 / conversation
60.0%Workflow · Qwen3.5-4B + gate$0.091 / conversation
62.5%
djev is the pick here. It matched the LLM workflow for 38% less, while the 4B passed fewer tasks.
The 4B agrees with the LLM on 94% of airline decisions, but its confidence barely separates its right answers from its wrong ones, so the gate cannot catch its mistakes. The best model for retail is the wrong model for airline.
7.3Telecom
Telecom troubleshooting has fewer closed decisions per turn. Its conversations run long and expensive, and the agent itself struggles.
Held-out pass^1 · bar is the pass rate · right is Anthropic $ per conversation
Sonnet agent (no workflow)$0.390 / conversation
80.0%Workflow · LLM decisions$0.168 / conversation
91.2%Workflow · djev$0.125 / conversation
85.0%Workflow · djev + gate$0.124 / conversation
86.3%Workflow · Qwen3.5-4B$0.145 / conversation
90.0%Workflow · Qwen3.5-4B + gate$0.137 / conversation
87.5%
Here the workflow beat the agent, and the 4B passed more tasks than the agent at 63% lower cost.
The LLM workflow beat the agent as well, so the gain comes from the workflow's fixed troubleshooting order. The 4B keeps that gain for less.
8Cost and speed
Retail
pass^1 (held-out)
Airline
pass^1 (held-out)
Telecom
pass^1 (held-out)
- Sonnet agent (no workflow)
- Workflow · LLM decisions
- Workflow · djev
- Workflow · djev + gate
- Workflow · Qwen3.5-4B
- Workflow · Qwen3.5-4B + gate
- Workflow · Qwen3.8-27B
- Workflow · Qwen3.5-2B
Per decision, the 4B answered in 275 ms against 1.2 to 1.9 seconds for Sonnet, and cost about 100 times less.
Per conversation the saving is smaller, because escalations to the full agent cost what they always did. Whole-conversation wall time barely changed, since the user simulator and the escalations take most of it.
What remains after the swap is almost entirely the full agent. Conversations handed off to it account for 94% of the 4B arm's remaining cost in retail and 99% in telecom, and the Haiku calls make up the rest.
So the next saving comes from widening what the workflow covers. A conversation the workflow finishes on its own costs a fraction of one it hands to the agent.
- Retail
- Airline
- Telecom
- Drag to rotate
9Which workloads suit a System One model
The saving grows with the share of an agent's work that is closed decisions. In τ² that share is high, with 27 to 41% of agent turns and 7 to 12 typed decisions per conversation.
Policy-driven service flows fit well, with their routing, eligibility checks, confirmations and structured troubleshooting. Open-ended drafting and analysis fit poorly, because there the output is the work itself.
In every domain the hardest decisions were routing an ambiguous request and working out which order or item the customer meant. Those are where a gate and a bigger model earn their cost.
10Applying this to your own agent
Everything here can be repeated on another agent. Start by logging its decisions and finding the ones that are closed questions with a fixed set of answers.
Compile those into a workflow, with code for the plumbing and a typed question at every branch. Leave anything open-ended to the full agent.
Then replay the logged decisions through each candidate model. Agreement shows which models can do the job, calibration shows whether a gate can catch their mistakes, and the budget picks among the rest.
Repeat the replay when the workflow or the models change. A new open model is one replay away from a place in the slot.
11Conclusions
System One models make an agent's closed decisions zero-shot, in a few hundred milliseconds each, with no model trained for the workflow. Adapting to a new workflow takes a replay of its decisions against the candidates, where a special-purpose model would take a dataset and a training run.
In our tests that kept the LLM workflow's accuracy and cut cost per conversation by 56 to 68% against the agent. The largest model never won, with a 4B matching models seven to nine times its size in retail and telecom and djev taking airline.
Any team running an agent already has what this takes, a log of its decisions and a shelf of open models. Replaying one against the other shows how much of the agent a System One model can take over.
Appendix
A.1Agreement by model
| Model | Agreement | 95% CI | ECE | AUROC |
|---|---|---|---|---|
| Qwen3.5-0.8B | 0.612 | [0.57, 0.65] | 0.138 | 0.67 |
| Qwen3.5-2B | 0.709 | [0.66, 0.75] | 0.083 | 0.80 |
| Qwen3.5-4B | 0.927 | [0.91, 0.94] | 0.036 | 0.82 |
| Qwen3.8-27B | 0.948 | [0.94, 0.96] | 0.026 | 0.89 |
| Qwen3.6-35B-A3B-FP8 | 0.931 | [0.92, 0.95] | 0.021 | 0.85 |
A.2End-to-end results
Held-out tasks, two trials each. Escalated is the share of tasks handed to the full agent.
| Arm | pass^1 | pass^2 | $ / conversation | Escalated |
|---|---|---|---|---|
| Sonnet agent (no workflow) | 0.863 | 0.800 | $0.228 | – |
| Workflow · LLM decisions | 0.688 | 0.575 | $0.153 | 43.8% |
| Workflow · djev | 0.562 | 0.475 | $0.078 | 38.8% |
| Workflow · djev + gate | 0.700 | 0.550 | $0.113 | 42.5% |
| Workflow · Qwen3.5-4B | 0.688 | 0.600 | $0.073 | 37.5% |
| Workflow · Qwen3.5-4B + gate | 0.625 | 0.550 | $0.082 | 37.5% |
| Workflow · Qwen3.8-27B | 0.700 | 0.600 | $0.088 | 40.0% |
| Workflow · Qwen3.5-2B | 0.375 | 0.275 | $0.078 | 32.5% |
| Arm | pass^1 | pass^2 | $ / conversation | Escalated |
|---|---|---|---|---|
| Sonnet agent (no workflow) | 0.925 | 0.900 | $0.200 | – |
| Workflow · LLM decisions | 0.700 | 0.650 | $0.140 | 47.5% |
| Workflow · djev | 0.700 | 0.650 | $0.087 | 40.0% |
| Workflow · djev + gate | 0.700 | 0.650 | $0.106 | 45.0% |
| Workflow · Qwen3.5-4B | 0.600 | 0.550 | $0.091 | 39.5% |
| Workflow · Qwen3.5-4B + gate | 0.625 | 0.550 | $0.091 | 36.4% |
| Arm | pass^1 | pass^2 | $ / conversation | Escalated |
|---|---|---|---|---|
| Sonnet agent (no workflow) | 0.800 | 0.700 | $0.390 | – |
| Workflow · LLM decisions | 0.912 | 0.850 | $0.168 | 85.0% |
| Workflow · djev | 0.850 | 0.725 | $0.125 | 83.8% |
| Workflow · djev + gate | 0.863 | 0.775 | $0.124 | 82.5% |
| Workflow · Qwen3.5-4B | 0.900 | 0.825 | $0.145 | 97.5% |
| Workflow · Qwen3.5-4B + gate | 0.875 | 0.850 | $0.137 | 97.5% |
A.3Cost per decision
| Engine | Latency p50 | Cost |
|---|---|---|
| LLM (Sonnet 4.6, constrained JSON) | 1.2 to 1.9 s | $0.0045 to 0.0104 |
| djev | 0.14 to 0.35 s | self-hosted |
| Qwen3.5-4B, cold / warm prefix cache (A100) | 275 / 108 ms | $0.00007 |
A.4Notes
The adapter records how much probability the model put on any valid letter before renormalising, and for the 4B that averaged 0.999. Every run checks that decisions were logged and that the engine reported no errors, after an early djev smoke test passed on escalations alone while every call had timed out.
Each arm has 80 conversations per domain, which gives intervals of about ten points either way.
References
- [1]Barres, V., Dong, H., Ray, S., Si, X., Narasimhan, K.. τ²-Bench: Evaluating conversational agents in a dual-control environment. arXiv preprint arXiv:2506.07982, 2025.