GPT-6 Sol and Luna vision fix: what to retest after September 25

OpenAI fixed an image-encoding bug in GPT-6 Sol and Luna. Use a controlled screenshot, OCR and chart checklist before trusting older vision results.

Binoculars drawn in black ink on a pale card, a blue-gray background and the title GPT-6 Sol + Luna.

OpenAI’s September 25, 2026 update fixes an image-encoding bug that degraded image understanding in GPT-6 Sol and GPT-6 Luna. If a screenshot, document image or visual automation task failed earlier, rerun the same task before concluding that the model cannot do it. The official API changelog says the fix affects visual tasks in both the API and Codex, including computer use.

That is a reason to revisit image-based evaluations, not evidence that every coding complaint has been resolved. This guide provides a reproducible retest procedure. It does not report an Ofox benchmark or a measured before-and-after improvement.

What changed, and what the announcement does not establish

QuestionWhat can be concluded
Which models were named?GPT-6 Sol and GPT-6 Luna
What was fixed?A bug in image encoding that degraded image understanding
Which surfaces were named?API and Codex visual tasks, including computer use
Is there a published improvement percentage?The cited announcement does not provide one
Did every third-party route receive the change simultaneously?The announcement does not establish that

A model name is only one part of an evaluation record. Provider route, execution time, image preprocessing, prompt, reasoning settings and tools also matter. Preserve those details alongside the answer. If your gateway changes the image before forwarding it, a remaining failure may come from that transformation rather than the model itself.

For a separate coding configuration decision, use our Sol High versus XHigh reasoning guide. Keep that choice separate from this particular visual regression.

Build a small test set with answers you can check

Choose images from the workflow you actually need. Use material you are allowed to send to the provider, and remove private information before submission. Keep an original copy rather than repeatedly exporting a screenshot through messaging apps.

TestExample questionA checkable result
Screenshot readingWhich control is disabled?The exact label and its visible state
Document OCRWhat is the invoice identifier?Exact characters, including leading zeroes
Chart readingWhich series is highest at the final point?Correct series and location, with uncertainty if unreadable
Visual navigationWhere is the requested button?A location that can be verified against the same screenshot

These are suggested fixtures, not completed tests. Include one easy image and one difficult image for each relevant task. A blurry or cropped source should be allowed an “unreadable” outcome: inventing an answer is not a successful extraction.

Before running anything, write the expected answer and acceptable tolerances. For OCR, decide whether spacing and punctuation matter. For charts, distinguish exact values from estimates based on an axis. For navigation, decide whether you are judging the interpretation alone or also the subsequent tool action.

Rerun without changing the entire experiment

  1. Record the timestamp, exact model ID and provider endpoint. For Codex, also record the client version, selected model, effort and relevant tool settings.
  2. Keep image bytes, image order, question and requested answer format unchanged. Save a SHA-256 hash of each input file.
  3. Use the same configuration across the models you are comparing. If a parameter is unsupported by one model, document that difference instead of silently dropping it.
  4. Run each case more than once when variability affects the decision. Keep every answer, including failures and refusals.
  5. Score the answers against the criteria written before the run. Record latency and reported usage separately from correctness.

Do not claim a before-and-after improvement unless you have actual saved results from before the fix. If the older run is missing, label the new table “post-fix evaluation” and compare only the results you really obtained. Reconstructing a bad earlier answer from memory is not a baseline.

A compact record can be a CSV with these columns:

run_time_utc,provider,model,client_version,effort,image_sha256,case_id,expected,actual,correct,latency_ms,request_id

The schema is a suggested logging format, not a provider response schema. Avoid placing API keys or private image URLs in the CSV. Store the raw responses separately with access limited to the people running the evaluation.

If image understanding is still wrong

Check the input before changing providers. Open the exact file sent to the API, inspect its dimensions and make sure a small label has not become illegible. Confirm that the request includes an image part rather than a text-only reference to a local filename. If the image is supplied by URL, verify that the provider can retrieve it without your browser session.

Next, separate perception from action. Ask the model to describe the target and its visible state before asking a tool to click it. A correct description followed by a failed click points to a different problem from an incorrect description. In an extraction workflow, check that the application parsed the full answer rather than discarding a field.

Finally, compare the provider’s supported image-input format and the actual outgoing request. Do not assume a successful text-only request proves that the image path works. The OpenAI vision guide is the reference for direct OpenAI requests; a gateway’s own documentation defines its route.

Decide whether to keep the model for this task

Use your post-fix results to answer a narrow question: does this configuration meet the acceptance criteria for your screenshots or documents? A model that reads one clean chart correctly has not demonstrated reliable navigation across an entire application.

If you also compare API costs, preserve usage and actual route information rather than assigning a token estimate from the image file size. Our Sol API cost guide explains why input, output and caching should be recorded separately. No new price or discount is assumed in this retest procedure.

Frequently Asked Questions

Does this fix prove Sol or Luna is better than Astra at vision?
No. The announcement identifies a bug and its correction; it is not a controlled comparison against Astra. Use the same images, scoring criteria and recorded settings if you need that comparison.
Should I rerun a text-only coding benchmark?
The cited change concerns image understanding. A purely text-based benchmark does not become invalid merely because this visual bug was fixed. Rerun it only when another relevant change or an evaluation problem justifies doing so.
Can I call a new result a percentage improvement?
Only with a valid earlier measurement and a defined metric. Without a saved pre-fix baseline, report the post-fix results on their own.