openai
Detection climbs from 13.8 to 46.2 mAP at Roboflow
Promtime
openaiRoboflow ran OpenAI's GPT-5.6 lineup through its upcoming VLM benchmark and recorded 46.2 mAP@50 on object detection for Sol, against 13.8 for GPT-5.5. The benchmark covers detection, counting, OCR and data extraction, and Roboflow calls Sol the best vision model OpenAI has released so far.
At a glance
- Detection results depend on the prompt format: the GPT-5.6 models score best when asked for absolute XYXY coordinates in image pixels, and the wrong format costs roughly 15 mAP points.
- Counting rose to 73.0% for Sol from 64.9% for GPT-5.5, with Terra at 67.6% and Luna at 66.2%, meaning even the cheapest model in the lineup beat the previous baseline.
- Gemini 3.5 Flash still leads Roboflow's detection and counting results at 0.8 cents per image, while Sol averaged roughly 2.5 cents and close to 10 seconds per image.
Detection and counting are the tasks that sit underneath the agent demos OpenAI put on stage, so a jump from 13.8 to 46.2 mAP@50 reads as the missing piece rather than a side upgrade. The gap that remains is economic: Gemini 3.5 Flash holds the top detection and counting results in the same benchmark at roughly a third of Sol's per-image cost, which likely keeps high-volume pipelines on Google's model for now.
Detection rose to 46.2 mAP@50 for Sol, with Terra at 44.7 and Luna at 43.3
OpenAI presented the GPT-5.6 lineup with a heavy emphasis on computer use, demonstrating models that navigate and operate desktop applications along with detailed 3D visualizations. In Roboflow's benchmark, Sol reached 46.2 mAP@50 on object detection, Terra 44.7 and Luna 43.3, all far above the 13.8 recorded for GPT-5.5.
Document layout detection was among the clearest strengths, with Sol locating titles, paragraphs, tables, images and signatures. Sol also held up on dense scenes of pills and eggs, where VLMs commonly fail because each label and coordinate set is generated as text, so longer responses raise the chance of missed objects, duplicates or coordinate errors.
For best detection results, Roboflow prompts the GPT-5.6 models to return absolute XYXY coordinates in image pixels, unlike Gemini 3.5 Flash, which performed best with YXYX coordinates normalized to a 0–1000 range. The wrong format cost about 15 mAP points in the benchmark.
OpenAI confirmed Sol loses stability on images around 2,000 by 2,000 pixels or larger
Sol sometimes returned boxes in seemingly random parts of an image, with little or no overlap with the ground truth and arranged in unnatural layouts such as straight rows or evenly spaced groups. Roboflow shared the cases with OpenAI, whose team confirmed that Sol becomes less stable on images around 2,000 by 2,000 pixels or larger, especially at lower reasoning effort.
Raising reasoning effort improves stability but increases token use, latency and cost; resizing or cropping large images before sending them to the OpenAI API is the practical workaround. The vision gains also come with higher token usage across the whole GPT-5.6 lineup, which matters mainly at scale, where token volume drives processing costs.
Counting reached 73.0% for Sol while text extraction slipped to 82.5%
Counting improved across the lineup: Sol scored 73.0% against 64.9% for GPT-5.5, with Terra at 67.6% and Luna at 66.2%. Sol counted heavily overlapping metal brackets, a case Roboflow calls difficult for traditional object detectors and VLMs, and counted bullet holes only inside selected scoring zones.
OCR stayed close to the previous generation, with Sol at 90.7% mean similarity against 91.2% for GPT-5.5, Terra at 88.8% and Luna at 88.4%. The gap was wider on targeted text extraction, where Sol reached 82.5% compared with 87.6% for GPT-5.5, followed by Luna at 81.4% and Terra at 79.4%.
Sol transcribed handwritten notes and pulled a requested date from another note, read a tire size printed on a worn tire, and returned a hockey broadcast score in the requested format. It failed to read the expiration date on a blister pack, struggled to count empty and sealed slots in blister packs, and returned a wrong total on an abnormal candy example.
What Terra and Luna cost
Terra averaged about 1 cent per image and around 6 seconds, while Luna cost less than 0.5 cents at slightly over 5 seconds, which Roboflow describes as the lineup's best latency-quality balance, close to Gemini 3.5 Flash in speed. Sol was the second most expensive model in the benchmark after Claude Fable 5. Roboflow plans to publish the full VLM benchmark within the next few weeks.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
