Introducing vidjev: applying System One to video and live feeds
- Joshua Okolo1,2
- Gokhan Egri2
1Harvard University2Brainbase Labs
· 14 min read
Abstract
1Decisions at the speed of the world
In our last post we took customer-service agents on τ²-bench [1] and compiled them into workflows. A small System One model made the typed decisions, and cost per conversation fell by more than half.
A support conversation can wait a second for each answer. A lot of the decisions people want to automate cannot.
Cars move quickly, and so do drones, players on a pitch and people in front of a security camera. A decision about any of them has to land in a fraction of a second, and then land again for the next frame.
Robotics and computer vision are the first places this kind of fast, frequent decision shows up. They are also where teams usually train a dedicated model for every new task.
System One models let you skip the training. They work zero-shot, so a new application means choosing a model and writing the questions, with no dataset to collect and no training run to wait for.
That makes a new application quick to build and quick to improve. When a better open model comes out, you swap it in and measure again.
vidjev is how we brought System One to images and video. To put it to work we asked two practical questions, which model is best at following a car from a drone, and which is best at watching CCTV footage.
- drone following on 3.5-minute drives with two frames of memory, up from 40% with one
- 90%
- drone following on 3.5-minute drives with two frames of memory, up from 40% with one
- following on short drives from a single frame, Qwen3.6-35B-A3B
- 98%
- following on short drives from a single frame, Qwen3.6-35B-A3B
- zero-shot frame AUC on UCF-Crime, level with a detector trained on it
- 0.844
- zero-shot frame AUC on UCF-Crime, level with a detector trained on it
- CCTV cameras per GPU at one frame per second
- 28
- CCTV cameras per GPU at one frame per second
2How vidjev works
Every decision in vidjev is a typed question about a frame. The model answers with a bounding box around something it was asked to find, or with a choice from a fixed list of options.
The model only says what it sees and where it is. Code does everything after that, from turning a box into a position on the road to deciding which clips a person should review.
We never ask the model for a distance, an angle, a count or an action. Those come from plain geometry and bookkeeping, which is cheaper and easier to trust.
Typed answers are also what keep the models fast. A box is about twenty tokens and a choice is a single token, so the model spends its time looking at the image instead of writing.
2.1Confidence from the model itself
A language model does not produce one answer. At every step it assigns a probability to every token it could write next.
In free text that probability is spread across countless ways of saying the same thing. A typed answer narrows each decision point to a handful of tokens, so the model's token probabilities become its probabilities over the answers.
For a closed question we label each option with a letter and decode exactly one token. The probability of option k is a softmax over the logits z of the option letters alone.
That gives a full distribution over the options from one forward pass, with no sampling and no second call. Every CCTV check in this post is scored this way.
Boxes get their confidence the same way. When the model opens its box it could instead write an empty one, either as a single token or as an opening bracket followed by a closing one.
So we never have to ask the model whether the target is in view. The answer comes from how likely it was to write nothing.
2.2The candidates
We tested five open Qwen models, Qwen3.5-0.8B, 2B and 4B, Qwen3.8-27B and Qwen3.6-35B-A3B. The last is a mixture of experts that uses about 3B of its parameters for each token.
We also tested djev, a Jev adapter on DiffusionGemma. Every model received the same questions and ran on a single NVIDIA B300.
2.3Choosing a model
We run every candidate on the real task and log each decision with its probabilities and timing. Then we take the cheapest configuration that clears the bar the task needs.
A configuration covers more than the model. It includes how many frames the model sees and how the question is worded, and in our tests both mattered as much as model size.
3Following a car from a drone
3.1The setup
We built the drone test in CARLA [2], an open-source driving simulator with detailed cities, traffic and pedestrians. A drone flies above the street and has to stay behind one target car as it drives through town.
Four times a second the drone's camera frame goes to the model with one request, to put a box around the target car or report that it is not there. Code turns the box into a spot on the road, keeps track of where the car is heading and flies the drone to a point 12 metres behind it.
We score a run by how often the drone is following, meaning the target car is in the camera's view at about the right distance. Two reference policies bracket the models, one that is handed the true box and one that hovers in place.
3.2Short drives
We started with one-minute drives through a downtown map with 50 other cars, showing each model only the current frame. Qwen3.6-35B-A3B followed the target car 98% of the time, against 100% for the policy that is handed the true box.
Following (%), one-minute drives, one frame
The largest difference between models was whether they would admit the car was gone. The 35B drew a box on none of the frames where the car was hidden behind a building or a bus.
Qwen3.8-27B drew boxes as tight as the 35B's but kept drawing them on 76% of the hidden frames. Each of those boxes pulled the drone toward an empty patch of road, and it followed 56% of the time.
djev followed 77% of the time. The two smallest Qwen models put a box on almost every frame whatever it showed, and followed 14% and 38% of the time.
3.3Look-alike cars
Next we parked four other red cars around the target. That turns the job from finding a car into finding the right car.
We tried two ways of asking. One asks for a box around only the target car, and the other asks for a box around every red car with a flag on the one that is the target, which code then follows.
- Box only the target car
- Box every red car, flag the target
Following with look-alike cars (%)
Boxing every car helped the 4B and the 35B. The 35B went from 38% to 51% and the 4B from 11% to 36%.
The same change broke the 27B, which often answered the list in prose and fell from 15% to 3%. The wording of the question is a setting you choose for each model, the same way you choose the model.
3.4Long drives and memory
For the full test we planned 3.5-minute routes through three different cities, each with 10 to 22 turns and no street driven twice. Every route ran with 50 other cars on the road.
From a single frame, most models lost the car early and never found it again. A car that slips behind a truck for a second leaves a one-frame model with nothing to go on.
So we gave the models memory, showing them the last 2, 4 or 8 frames along with the current one. The current frame always goes first, because Qwen models place their box in whichever image comes first.
- Qwen3.5-0.8B
- Qwen3.5-2B
- Qwen3.5-4B
- Qwen3.8-27B
- Qwen3.6-35B-A3B
- djev (one image)
Following on long drives (%)
Memory helped Qwen3.8-27B most. Two frames took it from 40% to 90%, and eight frames reached 92%, the best result in the study.
Qwen3.6-35B-A3B improved as well, from 9% with one frame to 53% with two. Qwen3.5-4B went the other way and stopped boxing anything once it saw more than one image.
On some routes the best configurations beat the policy that is handed the true box. That policy only gets a box when enough of the car is visible, while a model with memory keeps tracking a car that is half hidden under trees.
3.5Fast enough to fly
The drone's camera delivers four frames a second, so a model that keeps up has to make four decisions a second. We timed every decision on the short drives, with each drone getting a GPU to itself.
- Box plus identity check
- Box only
Decisions per second, one drone per GPU
djev and the Qwen models up to 4B all box the car faster than the camera delivers frames, with djev at 5.4 decisions a second and the 4B at 4.3. Qwen3.6-35B-A3B manages 3.1 a second for the box alone and 2.1 with the identity check.
Qwen3.8-27B, the most accurate model once it has memory, decides 1.8 times a second from a single frame, and each extra frame of memory is one more image to read. It suits a drone that can decide less than twice a second, or a target that moves more slowly than a car.
Several drones can share a GPU by batching their frames into one call. Four drones sharing one GPU got through 5.8 decisions a second in total with the 27B and 18 with the 0.8B, although each drone still waits for the whole batch.
For a drone that has to act on every frame, djev is the pick. If three decisions a second are enough, Qwen3.6-35B-A3B is more accurate, and when holding the car over long drives matters more than rate, Qwen3.8-27B with two frames of memory is the pick.
4Watching CCTV footage
4.1The setup
Surveillance turns the drone's problem around. Instead of one stream that needs fast answers, there are hundreds of cameras that each need a cheap one.
We used UCF-Crime [3], a public collection of 1,900 real surveillance videos covering 13 kinds of incident, from fighting and robbery to arson and road accidents. We ran all 290 of its test videos through every model at one frame per second, about 37,000 frames each.
Each frame got the questions a person at a monitor would ask, whether anyone is fighting, whether there is fire or smoke, whether a vehicle has crashed, whether someone is lying on the ground and whether a weapon is visible. The model also picked between normal and the 13 incident types.
The anomaly score for a frame is one minus the probability the model gave to normal. None of the models had seen a UCF-Crime label before this test.
4.2Accuracy
Qwen3.6-35B-A3B reached a frame-level AUC of 0.844. Detectors trained on UCF-Crime itself report between 0.754 and 0.870, so a model that never saw the dataset matches the trained RTFM detector [4] at 0.843.
Earlier zero-shot methods on this benchmark land between 0.58 and 0.66.
Frame-level AUC on the 290 UCF-Crime test videos
Qwen3.8-27B came second at 0.803, with djev at 0.797 and Qwen3.5-4B at 0.791. Qwen3.5-0.8B trailed at 0.717.
4.3Typed checks
The individual checks were sharper still. Every Qwen model's fighting check picked out fighting, assault and abuse with an AUC of at least 0.94, and the collision check found road accidents at 0.96 to 0.98.
| Fighting check on fighting, assault, abuse | Collision check on road accidents | Fire check on arson, explosions | |
|---|---|---|---|
| Qwen3.5-0.8B | 0.96 | 0.96 | 0.80 |
| Qwen3.5-2B | 0.94 | 0.97 | 0.93 |
| Qwen3.5-4B | 0.97 | 0.98 | 0.95 |
| Qwen3.8-27B | 0.97 | 0.98 | 0.94 |
| Qwen3.6-35B-A3B | 0.96 | 0.98 | 0.91 |
| djev | 0.79 | 0.67 | 0.76 |
Because every flag comes from a named check, a reviewer can see why a clip was flagged. A fire and a fight arrive labelled as a fire and a fight.
4.4Cost
The most accurate model was also close to the cheapest. With about 3B active parameters, Qwen3.6-35B-A3B answers every question for a frame in 36 ms, enough for 28 cameras on one B300 at one frame per second.
Frame-level AUC
Qwen3.8-27B was slower and less accurate, at 19 cameras per GPU. We also tried a cascade in which a small model screens every frame and sends suspicious ones to a large model, and the small models flagged so many normal frames that it saved almost nothing.
For CCTV the choice is Qwen3.6-35B-A3B, on accuracy and on cost.
5Choosing the model
The two workloads picked different winners from the same set of candidates. The drone needed a model that could use memory, and CCTV needed one that could answer several questions per frame cheaply.
| Drone following | CCTV AUC | |
|---|---|---|
| Qwen3.5-0.8B | 10% | 0.717 |
| Qwen3.5-2B | 46% | 0.785 |
| Qwen3.5-4B | 33% | 0.791 |
| Qwen3.8-27B | 92% | 0.803 |
| Qwen3.6-35B-A3B | 53% | 0.844 |
| djev | 38% | 0.797 |
Model size would not have predicted either winner. The 4B handled single frames and fell apart with memory, while the 35B, with only 3B active parameters, beat the 27B on surveillance.
The same procedure found both winners. We ran the candidates on the real task, measured accuracy and speed, and took the cheapest configuration that met the bar.
6Conclusions
System One models are a new paradigm for building decision systems. Both applications in this post ran zero-shot, with no dataset collected and no model trained for the task.
Building a new application becomes two steps. You construct the task as typed questions, then benchmark the candidate models on it and choose the best one.
We showed that this works for high-frequency video, from a drone deciding several times a second to dozens of camera streams checked every second on a single GPU. The same framework already ran business workflows in our τ²-bench post.
High-frequency video is at the core of physical AI, where machines have to see and act in real time. System One models could be instrumental in this next wave of physical intelligence.
Appendix
A.1More videos
A.2Full results
| Policy | Following | With look-alike cars | Boxes while hidden | ms per decision |
|---|---|---|---|---|
| Qwen3.5-0.8B | 14% | 5% | 100% | 240 |
| Qwen3.5-2B | 38% | 21% | 89% | 248 |
| Qwen3.5-4B | 36% | 11% | 73% | 322 |
| Qwen3.8-27B | 56% | 15% | 76% | 641 |
| Qwen3.6-35B-A3B | 98% | 38% | 0% | 471 |
| djev | 77% | 20% | 27% | 268 |
| Given the true box | 100% | 99% | – | – |
| Hovering | 13% | 14% | – | – |
| Model | 1 frame | 2 frames | 4 frames | 8 frames |
|---|---|---|---|---|
| Qwen3.5-0.8B | 10% | 2% | 2% | 9% |
| Qwen3.5-2B | 19% | 21% | 46% | 33% |
| Qwen3.5-4B | 33% | 3% | 2% | 3% |
| Qwen3.8-27B | 40% | 90% | 78% | 92% |
| Qwen3.6-35B-A3B | 9% | 53% | 49% | 52% |
| djev | 38% | – | – | – |
| Given the true box | 85% | – | – | – |
| Model | Frame AUC | 95% interval | Fighting | Collision | Fire | ms per frame | Cameras per B300 |
|---|---|---|---|---|---|---|---|
| Qwen3.5-0.8B | 0.717 | 0.61 to 0.82 | 0.96 | 0.96 | 0.80 | 24 | 42 |
| Qwen3.5-2B | 0.785 | 0.72 to 0.84 | 0.94 | 0.97 | 0.93 | 35 | 28 |
| Qwen3.5-4B | 0.791 | 0.73 to 0.86 | 0.97 | 0.98 | 0.95 | 27 | 37 |
| Qwen3.8-27B | 0.803 | 0.75 to 0.86 | 0.97 | 0.98 | 0.94 | 53 | 19 |
| Qwen3.6-35B-A3B | 0.844 | 0.80 to 0.89 | 0.96 | 0.98 | 0.91 | 36 | 28 |
| djev | 0.797 | 0.74 to 0.86 | 0.79 | 0.67 | 0.76 | 231 | 4 |
A.3Prompts
Drone, one frame.
Locate the red sports car in this drone camera image. Output only JSON {"bbox_2d": [x1, y1, x2, y2]} with coordinates normalised to 0-1000 (x right, y down). If it is not visible output {"bbox_2d": []}.Drone, with memory.
The FIRST image is the current frame from a drone camera; the following 3 images are earlier frames (most recent first, 0.25 s apart). Locate the red Ford Mustang in the FIRST image. Output only JSON {"bbox_2d": [x1, y1, x2, y2]} normalised to 0-1000 (x right, y down). If it is not visible in the first image output {"bbox_2d": []}.Drone, box every red car and flag the target.
Follow only the RED Ford Mustang. Other cars, including other Ford Mustangs in different colours (blue, white, black, yellow), may be in view; ignore them. Detect every red vehicle in this drone camera image. Output only a JSON list, one entry per red vehicle: [{"bbox_2d": [x1, y1, x2, y2], "is_target": true|false}] normalised to 0-1000, with "is_target": true only for the red Ford Mustang (at most one). Output [] if there is no red vehicle.CCTV questions, each scored as a single-token choice.
This is a CCTV surveillance frame. What is happening? (14 options, normal plus the 13 UCF-Crime classes) Is a person lying on the ground? Are people fighting or physically attacking each other? Is there fire or smoke? Is there a vehicle collision or crash? Is a weapon (gun, knife, bat) visible?
A.4Setup details
The drone flies 14 metres up with a 70 degree camera pitched 35 degrees down, and each frame is 448 by 448 pixels. It is limited to 5 m/s² of acceleration and 18 m/s, and downward ray casts keep it 4 metres above anything below.
A step counts as following when the target car is visible in the camera and 6 to 18 metres from the drone. Ground truth comes from CARLA's instance segmentation camera, and the policy given the true box receives it whenever at least 20 pixels of the car are visible.
Short drives last 60 seconds in Town10HD with 45 other cars. Long drives last 3.5 minutes on Town03, Town05 and Town10HD, with 10 to 22 turns and 50 other cars. The simulator waits for each decision, so every model faces the same situations.
CCTV frames are sampled at one per second from the 290 standard test videos, 140 with an incident and 150 normal. Frame AUC follows the standard protocol, with scores interpolated to every source frame. Cameras per GPU is 1000 divided by the milliseconds per frame for all six questions. See Table 3.
All Qwen models run under vLLM on one NVIDIA B300. djev is served through its own API.
References
- [1]Barres, V., Dong, H., Ray, S., Si, X., Narasimhan, K.. τ²-Bench: Evaluating conversational agents in a dual-control environment. arXiv preprint arXiv:2506.07982, 2025.
- [2]Dosovitskiy, A., Ros, G., Codevilla, F., López, A., Koltun, V.. CARLA: An open urban driving simulator. Conference on Robot Learning, 2017.
- [3]Sultani, W., Chen, C., Shah, M.. Real-world anomaly detection in surveillance videos. IEEE Conference on Computer Vision and Pattern Recognition, 2018.
- [4]Tian, Y., Pang, G., Chen, Y., Singh, R., Verjans, J. W., Carneiro, G.. Weakly-supervised video anomaly detection with robust temporal feature magnitude learning. IEEE International Conference on Computer Vision, 2021.