System1 Decisions from Multimodal Inputs
Research
5min Read
From classifying MNIST handwritten digits to playing Atari Q*bert: what gets faster when an open vision-language model scores answers instead of writing them?

Over the past year, we’ve been deploying video AI agents into workflows at some of the world’s largest media enterprises.
A recurring challenge in these deployments is answering several questions about the same footage: Is this live play or a replay? Which person is visible? Does this segment need a closer review? The footage may be complex, but each question often needs just one answer from a small set of choices.
Vision-language models (VLMs) typically generate those answers token by token. For workflows that ask many such questions, we wanted a faster way to return the required labels—and a way to reuse the processed footage when new questions follow.
Work on Jev, SemIf and CLM offered a starting point: models can make fixed-answer decisions from text by scoring the allowed answers directly.
We adapted that approach to open VLMs, using frames and video clips as evidence and returning probabilities without retraining the models or adding a classification head.* We also reused cached visual context across calls, so follow-up questions could build on footage the model had already processed.
* At the time of writing, OpenAI had announced a limited-preview Decisions API that uses Luna for finite-answer questions with text or image context. The announcement does not say whether Luna was fine-tuned for it.
How we adapted the VLM
We found that adapting a VLM for these decisions takes two changes at inference time.
First, inspired by SemIf‘s work with open language models, we read the logits for allowed answers and turn them into probabilities instead of asking the model to write each answer. That lets us score many questions together from the same visual context, following the parallel-decision idea Jev describes.
Second, we cache the VLM’s processed input. The system instructions, task hint and video frames form a shared prefix; the questions come last.
For a new batch about the same frame or clip, the server reuses that prefix and swaps Questions A for Questions B.
If a third batch depends on those answers, the application can append them as text after the cached frame, then ask Questions C. The model processes only that new text before scoring the new answers.
We forked llama.cpp, a widely used open-source inference engine, to add System1 answer scoring using its prompt cache. Our fork also includes scripts to reproduce the experiments below.
But does it work? VLM vs. VLM in System1 mode
Putting a VLM in System1 mode should make classification faster without sacrificing accuracy. We started with MNIST handwritten digits so we could measure both on a simple image task before moving to video.
In normal structured-output mode, Gemma 4 writes a schema-constrained JSON digit for each image, token by token.
In System1 mode, the same VLM scores the allowed labels 0–9 directly, sharing one input across 16 questions instead of writing 16 JSON fields. The benchmark covers all 10,000 MNIST test images in 16-image batches on an A100 GPU.
System1 classified 9,501/10,000 digits correctly (95.01%) in both runs; normal structured output got 9,480/10,000 (94.80%). On this test set, System1 recorded 0.21 percentage points higher accuracy while making the answer step much faster.
Answer step per image: 34.3× faster
Full request per image: input processing + answers
From raw game frames to actions
Jev showed Doom; CLM tested T-Rex and reported Super Mario. Their decision models received text descriptions of game state. Games are visual tasks, so we wanted to see what happens when the agent must read the actual game frames to decide what to do.
On an A100, we gave the quantized Qwen3.5 9B one unmodified Q*bert frame at each decision.
It made two sequential calls: first, it turned the frame into a predicted text state describing the cube colors, Q*bert, enemies, score and lives; then, given that predicted state and the frame, it chose the next move.
The native route wrote structured JSON answers. The System1 route scored the allowed answers directly.
Native structured output
65 decisions. The episode ended after Q*bert lost all lives, before completing level one.
System1 answer scoring
42 decisions. The episode reached the level-one transition.†
Mean per decision | Native | System1 |
|---|---|---|
Frame → state: answer step (input processing time excluded) | 2,736 ms | 306 ms |
State → action: answer step (input processing time excluded) | 136 ms | 17 ms |
Full frame → action (includes input processing time) | 5,592 ms | 3,047 ms |
Answer step means count JSON generation or System1 scoring after input processing. Full time includes both calls, image processing, input prefill and answer scoring.
Each column averages its own episode, so these are descriptive timings rather than paired measurements on identical frames. Video playback time is accelerated and should not be used to estimate inference latency.
Where this fits into media workflows
Many media workflows ask several small questions about the same footage. Together, the answers determine what gets tagged, reviewed or sent to another system. Fixed-answer scoring could make those repeated decisions faster, while caching could avoid processing the same visual input for each new set of questions.
The following are applications to evaluate beyond the experiments above. Each starts with a specific task and an output the team can use.
Help review teams find segments that need attention
A content-review workflow could begin with checks for visible material such as cigarettes or nudity in defined video segments. Several checks could share the same visual context. Flagged segments would then go to a reviewer or a model that can examine the surrounding context and explain the concern.
The output would be a review queue linked to the footage. The team would define the checks and assess missed material, false flags and review effort. A visual flag alone would not establish whether a programme meets a policy.
Turn training footage into useful catalogue metadata
A fitness team could use several questions to tag a training clip: which equipment is in use, whether the trainer is seated or standing, and whether they are pedalling. A later call could reuse the clip for another set of labels.
Those answers could populate search filters and catalogue fields, using the team’s definitions and review rules. Producing precise equipment-control instructions would require additional timing and validation work.
Make appearances searchable in broadcast archives
For a programme with a known presenter and guest roster, sampled frames could be checked against that roster, including an “unknown” or “none” option. Repeating the checks across predefined segments could produce an appearance index linked to the source footage.
The index would help archive users find candidate footage. Identifying who is speaking or what they are discussing would require additional evidence. The workflow would also need evaluation for missed appearances and incorrect matches.
Narrow sports footage before detailed analysis
Short sports segments could be checked for live play versus replay or classified against a defined set of events. Those results could narrow the footage passed to a more expensive model for detailed description or precise clip boundaries.
The output would be candidate segments for the next stage of the workflow. Classifying a segment as containing a goal would not locate the exact second it occurred or produce a finished highlight. Evaluation would need to account for relevant footage filtered out before deeper analysis.
Fixed-answer scoring applies when the output can be expressed as a predefined set of labels. Tasks that require descriptions or explanations still need generative output.
Evaluate classification accuracy and end-to-end latency on representative footage, including input processing and downstream steps. Prefix caching reduces repeated input processing only when calls share an unchanged prefix; new footage must still be processed.
What’s next
We are applying this approach to training and sports footage: tagging exercises and movement phases, marking plays and athlete states, and adding search labels to video libraries.
Our open-source implementation includes the inference changes and scripts to reproduce the MNIST and Q*bert experiments on your own hardware.











