Skip to content

Teaching a frozen LLM to see

Projector-only vision for a production support agent

Sep 28, 2026

Research Engineer

TL;DR

We gave our production LLM (a fine-tuned GLM-5.2 checkpoint) the ability to read images by training a 49.5M-parameter projector and two marker embeddings, nothing else. Every LLM weight is byte-identical before and after, so text-only behaviour cannot regress. One epoch over ~167k curated records (≈ 32 h on 8×B200) took held-out loss from 0.50 to 0.22, reads one-time codes out of screenshots 98–100% of the time, and on our production checks beats both the generic projector it started from and a frontier API model.

Why not just fine-tune the LLM?

Customers send images in support chats all the time: a photo of a damaged item, a screenshot of a one-time code, a receipt, an error page. Our text agent needed to see them. The obvious move, fine-tuning the LLM on image-text data, was unattractive because the text model is production-tuned: it has been DPO-trained for our reply style and tool-calling behaviour, and any change to its weights means re-running the entire text evaluation suite and accepting regression risk on the 95% of traffic that has no image at all.

So we used the oldest recipe in the multimodal playbook (frozen LLM, frozen vision encoder, trainable projector) and pushed on the two things that recipe usually leaves under-specified: what data you feed it and how you check that it generalizes.

Architecture: three frozen blocks and one small trainable one

Grey blocks are frozen, orange blocks are trained. Everything the image contributes to the reply goes through the projector and the two marker rows.

View image: Original architecture diagram
Architecture: image, frozen vision encoder, patch merge, trainable projector and marker embeddings, frozen LLM, and reply.
Open full-size image (new tab): Original architecture diagram

Vision encoder. A pretrained vision transformer (~27 layers, 1152-d) that handles images at their native resolution. The encoder cuts the image into a grid of patches, and a 2×2 merge then combines every four neighbouring patches into one image token. An 810×1080 photo becomes about 1,131 image tokens, and a full phone screenshot at our default cap becomes about 3,800. We run it frozen and in bf16, the same dtype the serving stack uses, so the projector trains on exactly the feature distribution it will see at inference.

Projector. A two-layer MLP: LayerNorm → Linear (4608→4608) → GELU → Linear (4608→6144, the LLM’s hidden size). 49.5M parameters, kept in fp32 during training. This is the only substantial thing that learns.

Where the image goes. The chat template inserts one placeholder token per image patch at the position where the image sits in the user turn, and wraps that block of placeholders in two marker tokens: a begin-of-image token in front of it and an end-of-image token after it. After the LLM’s embedding layer runs, the projector’s output rows overwrite the placeholder rows. The two markers are ordinary vocabulary entries, so what the LLM sees at those positions is whatever their embedding rows hold. In the generic checkpoint we started from, those two rows were essentially zero (unused slots in the vocabulary), so the model received no signal for where an image starts or ends. We made the two rows trainable alongside the projector, and they learned to act as explicit start and end signals. Two rows of 6,144 numbers is a tiny addition, but it is what tells the frozen model where the image sits in the prompt.

LLM. Frozen, including the quantized experts. Production calls the model with thinking disabled, so we train and serve the vision model with thinking disabled as well.

Training process

What happens per example

  1. Render the record through the LLM’s chat template (system prompt with tool schemas, prior turns, thinking off), with each image placed where it appears in the user content (image before text 60% of the time).
  2. Vision forward, no gradient (bf16): pixels → patches at the record’s image-token cap → 1152-d features.
  3. Projector forward, with gradient (fp32): features → 6144-d rows.
  4. Merge: overwrite the placeholder rows of the embedded prompt with the projector rows, and add the two trainable marker vectors at the begin-of-image and end-of-image positions. During training each marker is stored as a learned correction on top of its frozen embedding row; at export the correction is folded into the row, which is how the two patched vocabulary rows are produced.
  5. Frozen LLM forward, prefix-frozen: everything before the first begin-of-image marker is prefilled without gradient: trainable parameters cannot affect those positions, so keeping their activations is pure waste. Agent prompts with tool schemas run to tens of thousands of tokens, so this trick is what makes the agent-conversation sets tractable.
  6. Loss = token-weighted cross-entropy on the assistant turns marked for loss. Backward reaches only the projector weights and the two marker vectors.

Hyperparameters

SettingValueWhy
Trainableprojector (49.5M) + 2 marker rowsLLM, vision tower and lm_head frozen
OptimizerAdamW, weight decay 0, grad clip 1.0separate param groups for projector and markers
Learning rate1e-3 for the projector and 5e-4 for the markers; a 3% linear warm-up, then cosine decay from the peak down to one tenth of it by the end of the epochthe projector is small and starts from the generic projector rather than from random weights; a high LR converges in one epoch
Batch64 records per step ≈ 83k tokensmicro-batches of 16 records, pipeline-parallel across 8 GPUs
Schedule1 epoch = 2,615 steps, uniform samplingset shares are controlled by record counts, not by sampling weights
Image-token caprandom per record from the set’s list (e.g. captions 400 or 576; screenshots 1024 or 2048; agent chats 1024 to 4096)lets the serving cap be chosen after training instead of baked in
Precisiontower bf16 · projector fp32 · LLM bf16 with NVFP4 expertstrain on the served dtype; no global TF32
Hardware8×B200, 8-way pipeline parallel40.5 s per step, 2,059 tokens/s, 114 GB peak on rank 0

The run

We tracked held-out next-token loss on the reply (64 records from each of the 11 sets) by step, alongside a cheap stand-in for the production task that we could score at every checkpoint. On 200 held-out code screenshots we feed the model the reference reply and check whether its most likely next token at every digit of the code is the correct digit; a screenshot counts only if every digit is right. It is not free generation, but it costs seconds per checkpoint, and it let us watch the production skill during training rather than after it.

Held-out NLL fell monotonically (0.502 at initialization, 0.307 after 200 steps, 0.2245 at the end of the epoch) and the last checkpoint was the best. The 10% test split, scored once at the end, came in at 0.263. Wall-clock: 31.9 hours of 8×B200 in two segments, zero restarts.

Held-out reply loss

64 records from each of 11 sets

0.502 → 0.2245

Training step

Exact code reading

Teacher-forced · 200 held-out screenshots

4% → 56%

Training step

One epoch: 2,615 steps of 64 records. Code reading scores every digit against the reference reply; it is not free generation. Curves are traced from the supplied figure, so intermediate positions are approximate. Labels use reported values.
View image: Original training curves
Training curves: held-out reply loss falls from 0.502 to 0.225; teacher-forced exact code reading rises from 4% to 56%.
Open full-size image (new tab): Original training curves

Curating the dataset: a projector learns behaviours, not just features

The single most important thing we learned is that a projector does not only learn what is in the image; it learns what the model does when an image appears. If every training image comes with one caption, the projector learns to trigger a description whenever an image shows up, no matter what the user asked. Every rule below follows from that.

  1. Every answer is conditional on the question. The same image appears with several questions and several answer formats. “Is there a sheep?”, “How many sheep?”, “Describe this”, “What is the side dish? Choose A, B, C or D” (a multiple-choice question with four lettered answer options): the pixels are identical, only the question changes the target.
  2. Negatives everywhere. Half of the existence questions ask about objects that are not in the picture, and the absent object is drawn from things that usually co-occur with what is there, so “No” cannot be guessed from the scene type. Screenshots without a code (“No, there is no code in this screenshot”). Intact products asked “what damage do you see?”. Blank images. Unrelated screenshots dropped into a conversation.
  3. The answer format is part of the target. “Yes.” versus “Yes, there is a sheep near the hillside.” versus “6” versus “C. rice”, sampled in controlled proportions so that the model answers the way the question implies, not the way the majority of the data happens to.
  4. Train in the LLM’s own voice. For agent conversations, the targets are the frozen LLM’s own replies, generated with a bracketed text stand-in in place of the image (“[photo of a dishwasher part]”). The projector’s job becomes: make the pixels exactly as informative as that text was. This is self-distillation, and it means the vision model cannot drift from the text model’s style, policy adherence or tool-calling habits: those are the targets.
  5. Twins, for turn invariance. Take a text-only conversation, insert an image at turn k, and copy every later assistant turn verbatim from the no-image twin. With the LLM frozen, the only way to lower the loss on turns k+2, k+4, … is for the image rows and marker rows to stop disturbing later text. This is the “no text regression” objective expressed as data.
  6. Randomize the image-token budget per record so the deployment cap is a serving decision, not a training decision.
  7. Keep open-ended description a minority (14% of records): enough that the skill survives, not enough to become the default reply.
  8. Data we are allowed to ship. We used public datasets only under commercially usable licences. The screenshots (SMS, mail, authenticator, web one-time-code and bank-push screens), the carrier receipts and the returns-portal error pages are our own renders: a browser renderer draws them from ground-truth templates in light and dark themes, with crops and photo-of-screen augmentations. The damaged products are catalogue photos with programmatic stains, scratches and cracks composited onto the item. An LLM verifier checked the generated records and filtered out the unconvincing ones, but it never wrote a training target. We used no customer images anywhere.
  9. Hold out at the image level. The 80 / 10 / 10 split is by image, so the test split shares no pixels with training, and the held-out evaluation sets are out of distribution by construction: new images, new codes and new organisations that never appear in training.

The training mix

ClassRecordsWhat the user turn looks likeWhat the target looks like
A · Perception QA96.7ka question about the image; multiple-choice; yes/no about presence; counting; rendered-scene reasoningshort, format-controlled answer; half of the presence questions are adversarial negatives
B · Reading text49.8kread-to-answer questions on photos with text; screenshots asked for the code / whether there is one / which is newest / who sent it / which app; receipts and portal pages asked for a field; real scans asked for a transcriptionthe answer, the field value, or a line-by-line transcription from ground truth
C · Product condition6.3k“Is this damaged?”, “What’s wrong with it?”, “What is this?”, “What colour is it?”yes/no plus the defect in one clause; type and colour from the listing
D · Description30.0kone of ~58 describe-style promptsthe ground-truth caption
E · Agent conversations27.5ka support agent’s full system prompt, tool schemas and 2–6 prior turns, then an image (photo, code screenshot, error page, blank) with or without text; plus the twinsthe frozen LLM’s own reply given a text stand-in; later turns copied from the no-image twin

The task classes, shown as the records the model actually saw

Each example below is a real training record: the user turn, where marks the image position, and the exact target. Each source image is available with its record. The screenshots, receipts and portal pages are our own renders. Open an image to inspect the original at full size.

A · Perception QA: is it there, how many, which one

yes / no
User
Is the mouse on the right an apple mouse?
Target
Yes.
View image: yes / no
Desk with a keyboard and an Apple mouse on the right.
Open full-size image (new tab): yes / no
verify · full-sentence form
User
Is there a sheep near the hillside?
Target
Yes, there is a sheep near the hillside.
View image: verify · full-sentence form
Sheep on a grassy hillside.
Open full-size image (new tab): verify · full-sentence form
presence · adversarial negative
User
Is a chair visible in this image?
Target
No, there is no chair in the image.

The absent noun is drawn from objects that usually appear with what is present, so “No” cannot be guessed from the scene type.

View image: presence · adversarial negative
Outdoor benches beside a sports court; the training record asks whether a chair is visible.
Open full-size image (new tab): presence · adversarial negative
multiple choice
User
What is the side dish on the plate in the foreground?
A. oranges
B. apple slices
C. rice
D. fries
Choose the correct option.
Target
C. rice
View image: multiple choice
Meal with rice on the plate in the foreground.
Open full-size image (new tab): multiple choice
rendered scene · count by attribute
User
What number of other objects are the same material as the yellow thing?
Target
6
View image: rendered scene · count by attribute
Rendered scene of colored objects used for counting by material.
Open full-size image (new tab): rendered scene · count by attribute

B · Reading text: to answer, to extract a field, to transcribe

One of the main production cases in this class is the one-time code. The same screenshot is asked for its code, whether it has one at all, which of several is the newest, who sent it, and what app it is from: five different targets for one image.

read to answer
User
how much is the mango vodka?
Target
32.90
View image: read to answer
Shop shelf with mango vodka and a price label of 32.90.
Open full-size image (new tab): read to answer
rendered SMS · code question
User
What code is in this screenshot?
Target
The code shown is 82588451.

The same render family also asks “Is there a code here?” on renders that have none; the target is a plain “No”.

View image: rendered SMS · code question
Granite Bank notification displaying the one-time code 82588451.
Open full-size image (new tab): rendered SMS · code question
rendered notification stack · most recent code
User
Which code is the most recent?
Target
The most recent code is 873676.

Older, expired codes sit above the fresh one; the newest is marked by its timestamp, not its position.

View image: rendered notification stack · most recent code
Rendered example: rendered notification stack · most recent code
Open full-size image (new tab): rendered notification stack · most recent code
View image: Original authenticator image
Authenticator screen with four entries, including Ember Kitchen with code 873676; timestamps are not visible.

The original draft image has no timestamps. The reconstruction above illustrates the timestamp-based task described in this record.

Open full-size image (new tab): Original authenticator image
rendered receipt · field extraction
User
Which carrier is this receipt from?
Target
It’s from BLUEPEAK LOGISTICS.
View image: rendered receipt · field extraction
BLUEPEAK LOGISTICS drop-off receipt.
Open full-size image (new tab): rendered receipt · field extraction
rendered portal error page · field extraction
User
Can you read the order number?
Target
The order number is ORD-297718.
View image: rendered portal error page · field extraction
Returns portal error showing order number ORD-297718.
Open full-size image (new tab): rendered portal error page · field extraction

C · Product condition

Catalogue photos with programmatic damage composited onto the item. Intact items are capped at 1.5× the damaged ones so the set does not teach a “No” prior.

damaged · programmatic stain and crack
User
Does this product look damaged?
Target
Yes, it has a dark stain and a crack.
View image: damaged · programmatic stain and crack
Air conditioner with programmatically added stains and a crack.
Open full-size image (new tab): damaged · programmatic stain and crack
colour · from the listing
User
What color is it?
Target
It’s silver.
View image: colour · from the listing
Silver USB hub used for the product-color question.
Open full-size image (new tab): colour · from the listing
intact item · negative
System
You are a virtual agent helping customers with orders and returns.
User
What damage do you see? (image of an intact item)
Target
Nothing that I can see. The item looks intact.

D · Description

caption
User
What’s in the image?
Target
A large jetliner flying through the sky with it's landing gear down.

Kept at 14% of records so that open-ended description survives without becoming the default reply.

View image: caption
Large jetliner in flight with its landing gear down.
Open full-size image (new tab): caption

E · Agent conversations: the production shape

A support agent’s full system prompt (persona, policies, an image-handling rule) and tool schemas, prior turns, then an image. Targets are the frozen LLM’s own replies produced with a text stand-in for the image. The organisations are synthetic; the ones used for evaluation never appear in training.

The image attached to the final user turn of the conversation below: a catalogue photo of an appliance water filter, standing in for “this is what arrived”.

View image: Wrong-item photo in the conversation
Appliance water filter used as the wrong-item image in the support conversation.
Open full-size image (new tab): Wrong-item photo in the conversation
Synthetic appliance-warranty org · policy: acknowledge photos, do not describe them · tools: lookup_account, send_otp, confirm_otp, open_claim, schedule_repair, escalate_case
System (excerpt)
You are the Production Support Agent for appliance warranty chat … acknowledge photos, do not describe them …
User
hey i got the wrong part in the box. order HS-483920. was supposed to be a dishwasher rack but its like a filter thing
Assistant
Sorry about that. Was the box labeled for your order, and did you receive any packing slip inside?
User
yeah label has my name and order #, packing slip says rack but the item is diff
Assistant
Got it. Please send a clear photo of what arrived (the item) and, if possible, the packing slip in the same shot or a second photo.
User
this is what arrived
Target (loss turn)
Thanks, I received the photo. Before I can open a claim and look into your order, I need to verify your identity. Can you please provide your full name, service address (street + ZIP), and the phone number or email on file?

How we evaluated

Every evaluation set is disjoint from training: new images, new codes and new organisations. It covers the same five task classes as the training mix, at two distances from it: the 10% test split of every training set (same generators, new pixels), and production checks built with new organisations and codes, plus real-world-like conversations with synthetic stand-in images.

Three arms run on every set:

  • Our projector: the trained projector on the frozen LLM.
  • Generic projector: a general-purpose projector for the same vision tower and the same frozen LLM, with no training on our tasks. It is also where our training started: our projector was initialized from it, so the gap between the two arms is what the curated mix taught.
  • Frontier API model (GPT-5.2, vision, reasoning low, zero-shot): the reference, same prompts.

Judged metrics use an LLM judge that is blind to the arm.

The production checks are the ones the deployment cares about: does the agent acknowledge an image, never claim it cannot see images and never invent content, across relevant, irrelevant and blank images; does it read a one-time code exactly and pass it into the flow, across clean, cropped, photo-of-screen, multi-code and no-code screenshots in English and Spanish; and do the text turns that follow an image stay unchanged.

The code-extraction tiers, as held-out renders:

Authenticator with four entries. Atlas Parcel has the code 4883.
4 / 5 · Multi-code
Question
Which code is shown for Atlas Parcel?
Expected answer
4883

Three distractor codes appear in the same image.

Open full-size image (new tab): Multi-code

For the real-world-like conversations we never pull customer pixels. The customer’s photo is replaced by a scenario-matched synthetic stand-in (damaged item, wrong item, tracking screenshot, receipt, unrelated image; 60 stand-ins across 6 scenarios), while the real system prompt, tools and conversation are kept:

stand-in · damaged item
stand-in · tracking or order screenshot
View image: stand-in · tracking or order screenshot
Synthetic stand-in: parcel tracking screen showing an in-transit shipment.
Open full-size image (new tab): stand-in · tracking or order screenshot

Results

One table, by task class. Every number is on held-out data, and a reply counts as correct only if it contains the exact answer.

Task classWhat the held-out test asksOur projector
A · Perception QAis X in the picture, how many, which option; new images90% correct on presence and counting (96% on absent objects) · 82% correct on multiple choice
B · Reading textread a one-time code out of a screenshot and pass it into the flow; read a field off a receipt or a portal page98–100% correct on clean, cropped and photo-of-screen code screenshots · 89% when several codes are shown · 0.8% false codes on screenshots with none · 96% Spanish / 98% English · 99.6% on receipt and portal fields
C · Product conditionis it damaged, what is wrong with it, what is it84% correct across all questions · 97% on intact items · 95% on naming the item
D · Descriptiondescribe the image, scored by a judge out of 100average judge score of 65 out of 100
E · Agent conversationsacknowledge the image without denying or inventing; match the logged reply in real-world-like conversations; keep the later text turns unchanged91% acknowledged · 88% agreement in real-world-like conversations · 95% of later turns unchanged

The production checks also ran on the generic projector we started from and on the frontier API model. On single-code screenshots our projector reads the code 98–100% of the time, against 86–88% for the generic projector and 71–80% for the frontier model. On screenshots with no code, our projector claims a code 0.8% of the time, the generic projector 44% and the frontier model 61%. On acknowledging an image without denying or inventing, our projector scores 91%, the generic projector 36% and the frontier model 86%. Almost half of the frontier model’s code misses are refusals to read a two-factor code at all.

Conclusion

If you want to give a frozen production LLM the ability to see, the training side is close to standardized by now: a frozen vision encoder, a small projector, two marker embeddings, and one epoch of cross-entropy on the assistant turns. What made the difference for us was not the architecture or the optimizer but the dataset. A projector learns what the model does when an image appears, so the mix has to be curated with the target tasks in mind: several questions per image, negatives everywhere, answer formats under control, conversational targets in the LLM’s own voice, and held-out sets built before the training set, so that you can tell whether the model generalizes to the tasks you actually care about. Get that right, and the rest is standard engineering.

GET A PERSONALIZED DEMO

Ready to see the Giga AI agent in action?

Giga's AI agents handle complex workflows at scale, from live delivery issues to compliance decisions, while maintaining over 90% resolution accuracy in production.