Introducing the Universal Managed Agents API.Read the announcement
Research

Introducing vidjev: applying System One to video and live feeds

  • Joshua Okolo1,2
  • Gokhan Egri2

1Harvard University2Brainbase Labs

· 14 min read

Code

Nine feeds with the model's typed answers drawn over them. The drone tiles fly in CARLA, and the CCTV tiles are UCF-Crime test videos with faces pixelated.

Abstract

vidjev is how we brought System One to images and video. To put it to work we asked two practical questions, which model is best at following a car from a drone, and which is best at watching CCTV footage. Both run zero-shot, with no dataset collected and no model trained for the task.

1Decisions at the speed of the world

In our last post we took customer-service agents on τ²-bench [1] and compiled them into workflows. A small System One model made the typed decisions, and cost per conversation fell by more than half.

A support conversation can wait a second for each answer. A lot of the decisions people want to automate cannot.

Cars move quickly, and so do drones, players on a pitch and people in front of a security camera. A decision about any of them has to land in a fraction of a second, and then land again for the next frame.

Robotics and computer vision are the first places this kind of fast, frequent decision shows up. They are also where teams usually train a dedicated model for every new task.

System One models let you skip the training. They work zero-shot, so a new application means choosing a model and writing the questions, with no dataset to collect and no training run to wait for.

That makes a new application quick to build and quick to improve. When a better open model comes out, you swap it in and measure again.

vidjev is how we brought System One to images and video. To put it to work we asked two practical questions, which model is best at following a car from a drone, and which is best at watching CCTV footage.

drone following on 3.5-minute drives with two frames of memory, up from 40% with one
90%
drone following on 3.5-minute drives with two frames of memory, up from 40% with one
following on short drives from a single frame, Qwen3.6-35B-A3B
98%
following on short drives from a single frame, Qwen3.6-35B-A3B
zero-shot frame AUC on UCF-Crime, level with a detector trained on it
0.844
zero-shot frame AUC on UCF-Crime, level with a detector trained on it
CCTV cameras per GPU at one frame per second
28
CCTV cameras per GPU at one frame per second

2How vidjev works

Every decision in vidjev is a typed question about a frame. The model answers with a bounding box around something it was asked to find, or with a choice from a fixed list of options.

Figure 1. One decision in each of our two applications. The model answers a typed question about the frame, and ordinary code turns that answer into an action.
FRAMETYPED QUESTIONMODELTYPED ANSWERCODEACTIONdrone · 4 per secondbox the target car,or [] if not visibleSystem Onemodel[456, 556, 544, 703]P(visible) 0.99project to roadtrack the carplan the pathfly 12 mbehind the carcctv · 1 per secondis anyone fighting,normal or incidentSystem OnemodelP(fighting) 0.88P(normal) 0.02score the frame1 − P(normal)flag the clipfor review

The model only says what it sees and where it is. Code does everything after that, from turning a box into a position on the road to deciding which clips a person should review.

We never ask the model for a distance, an angle, a count or an action. Those come from plain geometry and bookkeeping, which is cheaper and easier to trust.

Typed answers are also what keep the models fast. A box is about twenty tokens and a choice is a single token, so the model spends its time looking at the image instead of writing.

2.1Confidence from the model itself

A language model does not produce one answer. At every step it assigns a probability to every token it could write next.

In free text that probability is spread across countless ways of saying the same thing. A typed answer narrows each decision point to a handful of tokens, so the model's token probabilities become its probabilities over the answers.

For a closed question we label each option with a letter and decode exactly one token. The probability of option k is a softmax over the logits z of the option letters alone.

p(k∣x)=exp⁡(zk)∑j∈optionsexp⁡(zj)p(k \mid x) = \frac{\exp(z_k)}{\sum_{j \in \text{options}} \exp(z_j)}
(1)

That gives a full distribution over the options from one forward pass, with no sampling and no second call. Every CCTV check in this post is scored this way.

Boxes get their confidence the same way. When the model opens its box it could instead write an empty one, either as a single token or as an opening bracket followed by a closing one.

P(visible)=1−[ P([])+P([)⋅P(]∣[) ]P(\text{visible}) = 1 - \Big[\, P(\texttt{[]}) + P(\texttt{[}) \cdot P(\texttt{]} \mid \texttt{[}) \,\Big]
(2)

So we never have to ask the model whether the target is in view. The answer comes from how likely it was to write nothing.

2.2The candidates

We tested five open Qwen models, Qwen3.5-0.8B, 2B and 4B, Qwen3.8-27B and Qwen3.6-35B-A3B. The last is a mixture of experts that uses about 3B of its parameters for each token.

We also tested djev, a Jev adapter on DiffusionGemma. Every model received the same questions and ran on a single NVIDIA B300.

2.3Choosing a model

We run every candidate on the real task and log each decision with its probabilities and timing. Then we take the cheapest configuration that clears the bar the task needs.

A configuration covers more than the model. It includes how many frames the model sees and how the question is worded, and in our tests both mattered as much as model size.

3Following a car from a drone

3.1The setup

We built the drone test in CARLA [2], an open-source driving simulator with detailed cities, traffic and pedestrians. A drone flies above the street and has to stay behind one target car as it drives through town.

Four times a second the drone's camera frame goes to the model with one request, to put a box around the target car or report that it is not there. Code turns the box into a spot on the road, keeps track of where the car is heading and flies the drone to a point 12 metres behind it.

Figure 2. Qwen3.8-27B with memory of its last three frames, the first 90 seconds of a 3.5-minute drive with 50 other cars. The card shows the model's box beside the true box, and the ribbon at the bottom marks every decision of the run.

We score a run by how often the drone is following, meaning the target car is in the camera's view at about the right distance. Two reference policies bracket the models, one that is handed the true box and one that hovers in place.

3.2Short drives

We started with one-minute drives through a downtown map with 50 other cars, showing each model only the current frame. Qwen3.6-35B-A3B followed the target car 98% of the time, against 100% for the policy that is handed the true box.

Figure 3. Following in traffic on one-minute drives with a single frame. The dotted lines are the policy handed the true box and a drone that hovers in place.
0%20%40%60%80%100%
14%
38%
36%
56%
98%
77%
given the true box
hovering in place
0.8B2B4B27B35B-A3Bdjev

Following (%), one-minute drives, one frame

The largest difference between models was whether they would admit the car was gone. The 35B drew a box on none of the frames where the car was hidden behind a building or a bus.

Qwen3.8-27B drew boxes as tight as the 35B's but kept drawing them on 76% of the hidden frames. Each of those boxes pulled the drone toward an empty patch of road, and it followed 56% of the time.

djev followed 77% of the time. The two smallest Qwen models put a box on almost every frame whatever it showed, and followed 14% and 38% of the time.

3.3Look-alike cars

Next we parked four other red cars around the target. That turns the job from finding a car into finding the right car.

We tried two ways of asking. One asks for a box around only the target car, and the other asks for a box around every red car with a flag on the one that is the target, which code then follows.

Figure 4. Following with four other red cars parked around the target, for the two ways of asking.
  • Box only the target car
  • Box every red car, flag the target
0%10%20%30%40%50%60%
11%
36%
15%
3%
38%
51%
Qwen3.5-4BQwen3.8-27BQwen3.6-35B-A3B

Following with look-alike cars (%)

Boxing every car helped the 4B and the 35B. The 35B went from 38% to 51% and the 4B from 11% to 36%.

The same change broke the 27B, which often answered the list in prose and fell from 15% to 3%. The wording of the question is a setting you choose for each model, the same way you choose the model.

3.4Long drives and memory

For the full test we planned 3.5-minute routes through three different cities, each with 10 to 22 turns and no street driven twice. Every route ran with 50 other cars on the road.

From a single frame, most models lost the car early and never found it again. A car that slips behind a truck for a second leaves a one-frame model with nothing to go on.

So we gave the models memory, showing them the last 2, 4 or 8 frames along with the current one. The current frame always goes first, because Qwen models place their box in whichever image comes first.

Figure 5. Following on 3.5-minute drives against the number of frames the model sees, with the best way of marking earlier boxes at each point.
0204060801001248Frames the model sees (current plus recent)given the true box
  • Qwen3.5-0.8B
  • Qwen3.5-2B
  • Qwen3.5-4B
  • Qwen3.8-27B
  • Qwen3.6-35B-A3B
  • djev (one image)

Following on long drives (%)

Memory helped Qwen3.8-27B most. Two frames took it from 40% to 90%, and eight frames reached 92%, the best result in the study.

Figure 6. The same model on the same route. On the left Qwen3.8-27B sees only the current frame and loses the car for good at about 27 seconds, and on the right it also sees its last three frames and holds on.

Qwen3.6-35B-A3B improved as well, from 9% with one frame to 53% with two. Qwen3.5-4B went the other way and stopped boxing anything once it saw more than one image.

On some routes the best configurations beat the policy that is handed the true box. That policy only gets a box when enough of the car is visible, while a model with memory keeps tracking a car that is half hidden under trees.

3.5Fast enough to fly

The drone's camera delivers four frames a second, so a model that keeps up has to make four decisions a second. We timed every decision on the short drives, with each drone getting a GPU to itself.

Figure 7. Decisions per second with one drone per GPU, measured over 1,200 decisions per model on the short drives with a single frame.
  • Box plus identity check
  • Box only
0123456
4.2
5.2
4.0
5.0
3.1
4.3
1.6
1.8
2.1
3.1
3.7
5.4
four a second, the drone's frame rate
0.8B2B4B27B35B-A3Bdjev

Decisions per second, one drone per GPU

djev and the Qwen models up to 4B all box the car faster than the camera delivers frames, with djev at 5.4 decisions a second and the 4B at 4.3. Qwen3.6-35B-A3B manages 3.1 a second for the box alone and 2.1 with the identity check.

Qwen3.8-27B, the most accurate model once it has memory, decides 1.8 times a second from a single frame, and each extra frame of memory is one more image to read. It suits a drone that can decide less than twice a second, or a target that moves more slowly than a car.

Several drones can share a GPU by batching their frames into one call. Four drones sharing one GPU got through 5.8 decisions a second in total with the 27B and 18 with the 0.8B, although each drone still waits for the whole batch.

For a drone that has to act on every frame, djev is the pick. If three decisions a second are enough, Qwen3.6-35B-A3B is more accurate, and when holding the car over long drives matters more than rate, Qwen3.8-27B with two frames of memory is the pick.

4Watching CCTV footage

4.1The setup

Surveillance turns the drone's problem around. Instead of one stream that needs fast answers, there are hundreds of cameras that each need a cheap one.

We used UCF-Crime [3], a public collection of 1,900 real surveillance videos covering 13 kinds of incident, from fighting and robbery to arson and road accidents. We ran all 290 of its test videos through every model at one frame per second, about 37,000 frames each.

Each frame got the questions a person at a monitor would ask, whether anyone is fighting, whether there is fire or smoke, whether a vehicle has crashed, whether someone is lying on the ground and whether a weapon is visible. The model also picked between normal and the 13 incident types.

The anomaly score for a frame is one minus the probability the model gave to normal. None of the models had seen a UCF-Crime label before this test.

Figure 8. Fighting042 from the UCF-Crime test set with faces pixelated. The fighting check and the anomaly score rise as the fight starts.

4.2Accuracy

Qwen3.6-35B-A3B reached a frame-level AUC of 0.844. Detectors trained on UCF-Crime itself report between 0.754 and 0.870, so a model that never saw the dataset matches the trained RTFM detector [4] at 0.843.

Earlier zero-shot methods on this benchmark land between 0.58 and 0.66.

Figure 9. Frame-level AUC on the 290 UCF-Crime test videos. Dotted lines are detectors trained on UCF-Crime, and the shaded band is earlier zero-shot methods.
0.50.60.70.80.9
earlier zero-shot methods
0.717
0.785
0.791
0.803
0.844
0.797
MGFN, trained
RTFM, trained
Sultani et al., trained
0.8B2B4B27B35B-A3Bdjev

Frame-level AUC on the 290 UCF-Crime test videos

Qwen3.8-27B came second at 0.803, with djev at 0.797 and Qwen3.5-4B at 0.791. Qwen3.5-0.8B trailed at 0.717.

4.3Typed checks

The individual checks were sharper still. Every Qwen model's fighting check picked out fighting, assault and abuse with an AUC of at least 0.94, and the collision check found road accidents at 0.96 to 0.98.

Figure 10. AUC of each yes-or-no check on the incident types it is meant to catch.
Fighting check
on fighting, assault, abuse
Collision check
on road accidents
Fire check
on arson, explosions
Qwen3.5-0.8B0.960.960.80
Qwen3.5-2B0.940.970.93
Qwen3.5-4B0.970.980.95
Qwen3.8-27B0.970.980.94
Qwen3.6-35B-A3B0.960.980.91
djev0.790.670.76

Because every flag comes from a named check, a reviewer can see why a clip was flagged. A fire and a fight arrive labelled as a fire and a fight.

Figure 11. Arson016. The fire check climbs with the flames.

4.4Cost

The most accurate model was also close to the cheapest. With about 3B active parameters, Qwen3.6-35B-A3B answers every question for a frame in 36 ms, enough for 28 cameras on one B300 at one frame per second.

Figure 12. Frame-level AUC against how many cameras one B300 can watch at one frame per second.
0.720.760.800.844102040Cameras per B300 at one frame per second (log scale)Qwen3.5-0.8BQwen3.5-0.8BQwen3.5-2BQwen3.5-2BQwen3.5-4BQwen3.5-4BQwen3.8-27BQwen3.8-27BQwen3.6-35B-A3BQwen3.6-35B-A3Bdjevdjev

Frame-level AUC

Qwen3.8-27B was slower and less accurate, at 19 cameras per GPU. We also tried a cascade in which a small model screens every frame and sends suspicious ones to a large model, and the small models flagged so many normal frames that it saved almost nothing.

For CCTV the choice is Qwen3.6-35B-A3B, on accuracy and on cost.

5Choosing the model

The two workloads picked different winners from the same set of candidates. The drone needed a model that could use memory, and CCTV needed one that could answer several questions per frame cheaply.

Figure 13. Each model's best result on the two workloads. Darker cells rank higher within a column.
Drone followingCCTV AUC
Qwen3.5-0.8B10%0.717
Qwen3.5-2B46%0.785
Qwen3.5-4B33%0.791
Qwen3.8-27B92%0.803
Qwen3.6-35B-A3B53%0.844
djev38%0.797

Model size would not have predicted either winner. The 4B handled single frames and fell apart with memory, while the 35B, with only 3B active parameters, beat the 27B on surveillance.

The same procedure found both winners. We ran the candidates on the real task, measured accuracy and speed, and took the cheapest configuration that met the bar.

6Conclusions

System One models are a new paradigm for building decision systems. Both applications in this post ran zero-shot, with no dataset collected and no model trained for the task.

Building a new application becomes two steps. You construct the task as typed questions, then benchmark the candidate models on it and choose the best one.

We showed that this works for high-frequency video, from a drone deciding several times a second to dozens of camera streams checked every second on a single GPU. The same framework already ran business workflows in our τ²-bench post.

High-frequency video is at the core of physical AI, where machines have to see and act in real time. System One models could be instrumental in this next wave of physical intelligence.

Appendix

A.1More videos

Qwen3.6-35B-A3B with memory of its last three frames on the same route as the hero clip.
Qwen3.5-4B with memory. It stops returning boxes once it sees several images, and the drone loses the car.
RoadAccidents127. The score peaks at the collision.
Normal_Videos_014, an ordinary clip where the score stays low.

A.2Full results

Table 1. Short drives, one frame, four seeds per model.
PolicyFollowingWith look-alike carsBoxes while hiddenms per decision
Qwen3.5-0.8B14%5%100%240
Qwen3.5-2B38%21%89%248
Qwen3.5-4B36%11%73%322
Qwen3.8-27B56%15%76%641
Qwen3.6-35B-A3B98%38%0%471
djev77%20%27%268
Given the true box100%99%––
Hovering13%14%––
Table 2. Long drives with 50 other cars, best overlay setting per cell, four routes each.
Model1 frame2 frames4 frames8 frames
Qwen3.5-0.8B10%2%2%9%
Qwen3.5-2B19%21%46%33%
Qwen3.5-4B33%3%2%3%
Qwen3.8-27B40%90%78%92%
Qwen3.6-35B-A3B9%53%49%52%
djev38%–––
Given the true box85%–––
Table 3. CCTV, all 290 UCF-Crime test videos at one frame per second.
ModelFrame AUC95% intervalFightingCollisionFirems per frameCameras per B300
Qwen3.5-0.8B0.7170.61 to 0.820.960.960.802442
Qwen3.5-2B0.7850.72 to 0.840.940.970.933528
Qwen3.5-4B0.7910.73 to 0.860.970.980.952737
Qwen3.8-27B0.8030.75 to 0.860.970.980.945319
Qwen3.6-35B-A3B0.8440.80 to 0.890.960.980.913628
djev0.7970.74 to 0.860.790.670.762314

A.3Prompts

Drone, one frame.

Locate the red sports car in this drone camera image. Output only JSON {"bbox_2d": [x1, y1, x2, y2]} with coordinates normalised to 0-1000 (x right, y down). If it is not visible output {"bbox_2d": []}.

Drone, with memory.

The FIRST image is the current frame from a drone camera; the following 3 images are earlier frames (most recent first, 0.25 s apart). Locate the red Ford Mustang in the FIRST image. Output only JSON {"bbox_2d": [x1, y1, x2, y2]} normalised to 0-1000 (x right, y down). If it is not visible in the first image output {"bbox_2d": []}.

Drone, box every red car and flag the target.

Follow only the RED Ford Mustang. Other cars, including other Ford Mustangs in different colours (blue, white, black, yellow), may be in view; ignore them. Detect every red vehicle in this drone camera image. Output only a JSON list, one entry per red vehicle: [{"bbox_2d": [x1, y1, x2, y2], "is_target": true|false}] normalised to 0-1000, with "is_target": true only for the red Ford Mustang (at most one). Output [] if there is no red vehicle.

CCTV questions, each scored as a single-token choice.

This is a CCTV surveillance frame. What is happening? (14 options, normal plus the 13 UCF-Crime classes)
Is a person lying on the ground?
Are people fighting or physically attacking each other?
Is there fire or smoke?
Is there a vehicle collision or crash?
Is a weapon (gun, knife, bat) visible?

A.4Setup details

The drone flies 14 metres up with a 70 degree camera pitched 35 degrees down, and each frame is 448 by 448 pixels. It is limited to 5 m/s² of acceleration and 18 m/s, and downward ray casts keep it 4 metres above anything below.

A step counts as following when the target car is visible in the camera and 6 to 18 metres from the drone. Ground truth comes from CARLA's instance segmentation camera, and the policy given the true box receives it whenever at least 20 pixels of the car are visible.

Short drives last 60 seconds in Town10HD with 45 other cars. Long drives last 3.5 minutes on Town03, Town05 and Town10HD, with 10 to 22 turns and 50 other cars. The simulator waits for each decision, so every model faces the same situations.

CCTV frames are sampled at one per second from the 290 standard test videos, 140 with an incident and 150 normal. Frame AUC follows the standard protocol, with scores interpolated to every source frame. Cameras per GPU is 1000 divided by the milliseconds per frame for all six questions. See Table 3.

All Qwen models run under vLLM on one NVIDIA B300. djev is served through its own API.

References

  1. [1]Barres, V., Dong, H., Ray, S., Si, X., Narasimhan, K.. τ²-Bench: Evaluating conversational agents in a dual-control environment. arXiv preprint arXiv:2506.07982, 2025.
  2. [2]Dosovitskiy, A., Ros, G., Codevilla, F., López, A., Koltun, V.. CARLA: An open urban driving simulator. Conference on Robot Learning, 2017.
  3. [3]Sultani, W., Chen, C., Shah, M.. Real-world anomaly detection in surveillance videos. IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  4. [4]Tian, Y., Pang, G., Chen, Y., Singh, R., Verjans, J. W., Carneiro, G.. Weakly-supervised video anomaly detection with robust temporal feature magnitude learning. IEEE International Conference on Computer Vision, 2021.