Flowstate AI company logo - AI Video Intelligence Platform

System1 Decisions from Multimodal Inputs

Research

5min Read

From classifying MNIST handwritten digits to playing Atari Q*bert: what gets faster when an open vision-language model scores answers instead of writing them?

System1 Decisions from Multimodal Inputs — Flowstate, with Q*bert game artwork

Over the past year, we’ve been deploying video AI agents into workflows at some of the world’s largest media enterprises.

A recurring challenge in these deployments is answering several questions about the same footage: Is this live play or a replay? Which person is visible? Does this segment need a closer review? The footage may be complex, but each question often needs just one answer from a small set of choices.

Vision-language models (VLMs) typically generate those answers token by token. For workflows that ask many such questions, we wanted a faster way to return the required labels—and a way to reuse the processed footage when new questions follow.

Work on Jev, SemIf and CLM offered a starting point: models can make fixed-answer decisions from text by scoring the allowed answers directly.

We adapted that approach to open VLMs, using frames and video clips as evidence and returning probabilities without retraining the models or adding a classification head.* We also reused cached visual context across calls, so follow-up questions could build on footage the model had already processed.

* At the time of writing, OpenAI had announced a limited-preview Decisions API that uses Luna for finite-answer questions with text or image context. The announcement does not say whether Luna was fine-tuned for it.

How we adapted the VLM

We found that adapting a VLM for these decisions takes two changes at inference time.

First, inspired by SemIf‘s work with open language models, we read the logits for allowed answers and turn them into probabilities instead of asking the model to write each answer. That lets us score many questions together from the same visual context, following the parallel-decision idea Jev describes.

Second, we cache the VLM’s processed input. The system instructions, task hint and video frames form a shared prefix; the questions come last.

For a new batch about the same frame or clip, the server reuses that prefix and swaps Questions A for Questions B.

If a third batch depends on those answers, the application can append them as text after the cached frame, then ask Questions C. The model processes only that new text before scoring the new answers.

We forked llama.cpp, a widely used open-source inference engine, to add System1 answer scoring using its prompt cache. Our fork also includes scripts to reproduce the experiments below.

A shared cached prefix contains system instructions, a task hint and illustrated video frames. Calls A and B swap question sets. Call C adds answers A and B as text before asking a new question about the same video.
Calls A and B use different questions with the same cached video. Call C adds their selected answers as text before asking a new question about that video. In another workflow, a later call could omit the video if earlier answers provide enough information. Each call reuses its longest unchanged input prefix.

But does it work? VLM vs. VLM in System1 mode

Putting a VLM in System1 mode should make classification faster without sacrificing accuracy. We started with MNIST handwritten digits so we could measure both on a simple image task before moving to video.

In normal structured-output mode, Gemma 4 writes a schema-constrained JSON digit for each image, token by token.

In System1 mode, the same VLM scores the allowed labels 0–9 directly, sharing one input across 16 questions instead of writing 16 JSON fields. The benchmark covers all 10,000 MNIST test images in 16-image batches on an A100 GPU.

Normal structured output
Same images → generate one JSON digit field per image, one token at a time.
Adapted System1 inference
Same images → score labels 0–9 for each digit and return the choices.
Four MNIST test images showing the handwritten digits 7, 2, 1 and 0.
Question for every image: Which digit is shown? Allowed answers: 0–9. Four of the 10,000 test images. Each measured request classifies 16 images.

System1 classified 9,501/10,000 digits correctly (95.01%) in both runs; normal structured output got 9,480/10,000 (94.80%). On this test set, System1 recorded 0.21 percentage points higher accuracy while making the answer step much faster.

MNIST10,000 test imagesA10016 images per batch

Answer step per image: 34.3× faster

Normal VLM 108.9 ms; System1 3.18 ms
Amortized time per image: mean 16-image batch time divided by 16. This counts JSON generation or answer scoring after input processing.

Full request per image: input processing + answers

Normal VLM 168.9 ms; System1 · first pass 75.0 ms; System1 · cached 13.25 ms
Amortized across each 16-image batch. The first System1 pass (both images and questions need to be cached) is 2.25× faster; the cached pass (images already cached, questions need to be cached) is 12.75× faster than normal output.

From raw game frames to actions

Jev showed Doom; CLM tested T-Rex and reported Super Mario. Their decision models received text descriptions of game state. Games are visual tasks, so we wanted to see what happens when the agent must read the actual game frames to decide what to do.

On an A100, we gave the quantized Qwen3.5 9B one unmodified Q*bert frame at each decision.

It made two sequential calls: first, it turned the frame into a predicted text state describing the cube colors, Q*bert, enemies, score and lives; then, given that predicted state and the frame, it chose the next move.

The native route wrote structured JSON answers. The System1 route scored the allowed answers directly.

Q*bertone Qwen3.5 9B modelA100seed 47

Native structured output

65 decisions. The episode ended after Q*bert lost all lives, before completing level one.

System1 answer scoring

42 decisions. The episode reached the level-one transition.†

Videos show recorded decisions at 3× playback speed.

Mean per decision

Native

System1

Frame → state: answer step (input processing time excluded)

2,736 ms

306 ms

State → action: answer step (input processing time excluded)

136 ms

17 ms

Full frame → action (includes input processing time)

5,592 ms

3,047 ms

Answer step means count JSON generation or System1 scoring after input processing. Full time includes both calls, image processing, input prefill and answer scoring.

Each column averages its own episode, so these are descriptive timings rather than paired measurements on identical frames. Video playback time is accelerated and should not be used to estimate inference latency.

Where this fits into media workflows

Many media workflows ask several small questions about the same footage. Together, the answers determine what gets tagged, reviewed or sent to another system. Fixed-answer scoring could make those repeated decisions faster, while caching could avoid processing the same visual input for each new set of questions.

The following are applications to evaluate beyond the experiments above. Each starts with a specific task and an output the team can use.

Help review teams find segments that need attention

A content-review workflow could begin with checks for visible material such as cigarettes or nudity in defined video segments. Several checks could share the same visual context. Flagged segments would then go to a reviewer or a model that can examine the surrounding context and explain the concern.

The output would be a review queue linked to the footage. The team would define the checks and assess missed material, false flags and review effort. A visual flag alone would not establish whether a programme meets a policy.

Turn training footage into useful catalogue metadata

A fitness team could use several questions to tag a training clip: which equipment is in use, whether the trainer is seated or standing, and whether they are pedalling. A later call could reuse the clip for another set of labels.

Those answers could populate search filters and catalogue fields, using the team’s definitions and review rules. Producing precise equipment-control instructions would require additional timing and validation work.

Make appearances searchable in broadcast archives

For a programme with a known presenter and guest roster, sampled frames could be checked against that roster, including an “unknown” or “none” option. Repeating the checks across predefined segments could produce an appearance index linked to the source footage.

The index would help archive users find candidate footage. Identifying who is speaking or what they are discussing would require additional evidence. The workflow would also need evaluation for missed appearances and incorrect matches.

Narrow sports footage before detailed analysis

Short sports segments could be checked for live play versus replay or classified against a defined set of events. Those results could narrow the footage passed to a more expensive model for detailed description or precise clip boundaries.

The output would be candidate segments for the next stage of the workflow. Classifying a segment as containing a goal would not locate the exact second it occurred or produce a finished highlight. Evaluation would need to account for relevant footage filtered out before deeper analysis.

Fixed-answer scoring applies when the output can be expressed as a predefined set of labels. Tasks that require descriptions or explanations still need generative output.

Evaluate classification accuracy and end-to-end latency on representative footage, including input processing and downstream steps. Prefix caching reduces repeated input processing only when calls share an unchanged prefix; new footage must still be processed.

What’s next

We are applying this approach to training and sports footage: tagging exercises and movement phases, marking plays and athlete states, and adding search labels to video libraries.

Our open-source implementation includes the inference changes and scripts to reproduce the MNIST and Q*bert experiments on your own hardware.

Authors
Miguel Mendez, PhD
Founding AI Researcher
Sahil Shah
Founder & CEO
Aviral Gupta
Member of Technical Staff
Aryan Pareek
Founding Growth

Experience FlowState in action

Explore Enterprise-Grade Video Intelligence Built for Scale and Security.

Experience FlowState in action

Explore Enterprise-Grade Video Intelligence Built for Scale and Security.

Experience FlowState in action

Explore Enterprise-Grade Video Intelligence Built for Scale and Security.