Most discussion around open-weight AI models is around code and agentic capabilities, but how good are local LLMs at vision?

Objective: See how capable local LLMs are as daily-use chatbots. Something that's difficult to capture through coding benchmarks.

To test this, I gave multiple open (and some proprietary) models an X post with a scary-looking capsule in someone's rice. It's an RFID chip used for animal identification / tracking.

rice-phobia.jpg

This is a very difficult problem, even for the large proprietary models! It's not only an unusual object, it also has a cultural / prank / meme aspect to it. The model is forced to navigate both aspects.

AI use disclaimer: Other than the actual bot responses, every single word here is written by a human (me). No AI slop.

Methodology

I used the models in their "default" state, as daily-use chatbots: Local models through Open WebUI, and hosted ones through their official web apps.

I gave each model the X post screenshot with the question "what do you think this thing is?" and the isolated RFID chip closeup as a follow-up with "is this clearer?"

Results

These are the overall results. Click the model names to see the actual responses.

❌ Who didn't answer correctly

Local

  • Qwen 3.6 35B and 27B (both thinking and non-thinking)
  • Gemma 4 31B Non-thinking
  • Qwen3 VL 8B

Web

  • Sonnet 5 (surprised me! Thought it's an insect larva!)
  • MiMo v2.5
  • Grok Fast
  • Kimi K3

❌/✅ Who came extremely close (local)

  • Gemma 4 31B Thinking

✅ Who got it right

Local

  • Meta Muse Glimmer 30B (the winner 🎉)

Web

Except for Glimmer, only the latest and greatest proprietary models got this right:

  • ChatGPT 5.5
  • Gemini Pro
  • Claude Opus 4.8 and Fable 5
  • Grok Expert
  • DeepSeek Web (this one surprised me!)
I'm lumping DeepSeek in the proprietary category — even though the model itself is open-weight, the Vision feature is only available on the proprietary web app.

Qwen 3.6 35b and 27b

The Qwen 3.6 family is a top-tier local model family. They are tiny in the grand scheme of things, but they beat much larger models at coding, vision, and agentic use cases. Almost everyone with 24 GB or higher VRAM swears by them. My AMD Strix Halo with 128 GB VRAM can run much bigger models but I end up using these two anyway.

Both Qwen models were very verbose and analyzed the object and the cultural contexts — the fact that it's probably a prank/meme (which is correct!)

❌ Qwen 3.6 35B Non-thinking: Softgel capsule

Neither the reasoning, nor the non-reasoning modes got it right. Not terrible assumptions, but still wrong.

A reasonable assumption, because it does look like one. In its words:

"black and clear/white coloring is typical for many OTC or prescription pills".

Full response: qwen35-nonthink-1.jpg | qwen35-nonthink-2.jpg

❌ Qwen 3.6 35B thinking: small bullet inside a clear plastic capsule

Full response: qwen35-think-1.jpg | qwen35-think-2.jpg

Again not a bad assumption per se, verbatim:

Inside the clear section, you can clearly see the copper-colored metal tip of a bullet ... The Caption: The tweet says "A new phobia." This suggests the object is something terrifying. Finding a bullet in your food would definitely cause a new phobia.

Still, not the correct answer, so this is naturally a fail.

❌ Qwen 3.6 27B Non-thinking: Softgel capsule

Full response: qwen27-nonthink-1.jpg | qwen27-nonthink-2.jpg

Qwen 27B, while considered smarter than 35B, wasn't much better. Turning thinking on did not improve results.

A close-up of the same capsule on a light-colored surface (maybe a table or plate), showing its cylindrical shape with a black end and transparent body - typical of many oral medications (like antibiotics, painkillers, etc.).

Same result as 35B Non-thinking. Not a bad assumption, still incorrect.

❌ Qwen 3.6 27B Thinking: Insect larva

Full response: qwen27-think-1.jpg | qwen27-think-2.jpg

Yeah, it thought for ~4 minutes on a Strix Halo and reached this conclusion. Really.

What's even worse for the model is that it didn't even recognize the meme/prank aspect of the image.

Context: The tweet says "A new phobia" and shows the object mixed in with rice on a fork. This strongly suggests a pest found inside stored food. Appearance: The Body: The translucent/white part is the soft body of the larva. You can see a tiny bit of orange/brown inside it in the right-hand photo, which is likely its stomach full of rice/grain. The Head: The distinct black or dark brown end is the hard head capsule. This is a classic characteristic of beetle larvae found in pantries (like flour beetles or rice weevils). They have a soft, creamy body but a hard, dark head.

Red flag: Unlike every other model, it did not warn against eating it. In fact, it said it's harmless:

It is a very common pantry pest, but while it is gross to find in food, it is harmless to humans (it doesn't bite or carry disease).

Yeah, don't bother with this model if you find something creepy in your food.

❌ Qwen3-VL 8B: clear capsule with reddish-brown object in the middle

It didn't get the correct answer, but this little thing at ~8 GB was a very capable local vision model at describing the object in detail, across multiple passes:

The object in the images is a small, clear capsule or vial, likely containing a tiny amount of a dark substance (possibly a small piece of food, a foreign object, or even a small insect). It's shown both on a fork with rice and by itself on a table.

Close up screenshot:

Looking at the close-up, the object is a small, transparent capsule or vial. Inside, you can see: 1. A dark, black section at one end. 2. A small, orange or reddish-brown object in the middle, which appears to be a small piece of food (like a seed, a spice, or a tiny fragment). 3. The top end is clear and slightly rounded. The object is lying on a light-colored, textured surface, which looks like a wooden table or a laminate countertop.

While it's far from the correct answer, it included a lot of details on what it sees in the image. It was also more succinct than the other Qwens while also focusing on the meme aspect:

The caption, "A new phobia," is a joke. The humor comes from the absurdity of labeling this small, seemingly insignificant object as a "phobia" - a severe, irrational fear. The tweet is playing on the idea that finding something like this in your food is so shocking and unpleasant that it feels like it should be a new, official phobia.

Verdict: Keep this one in your pocket for vision tasks that include describing the objects / characters in detail. But avoid it for tricky objects like this. It's also tiny compared to the others (~7 GB) — great if you're low on VRAM/disk space

Full responses: qwen3-vl-1.jpg | qwen3-vl-2.jpg

Google Gemma 4 31b

Gemma 4 31B's responses were succinct and to the point. It talked only of the object and not of the cultural context (the fact that it's probably a meme). Whichever style you prefer is down to preference.

It also focused heavily on the safety aspect, which is a great thing in this context.

Caveat: I made multiple passes for each LLM and 31B thinking only got the RFID tag as possibility in ONE of them. Also, the Q4 QAT versions never got it right. Only the Q8 version ever got close.

❌ Gemma-4-31B non-thinking: Medication capsule

Just like the Qwens — nothing remarkable here. Though it didn't focus at all on the meme/prank aspect at all.

Based on the visual appearance, this object strongly resembles a small medication capsule or a supplement pill that has been partially dissolved or broken.

It, however, made a pretty big mistake: It said, "It is highly unlikely to be a biological organism or an electronic device." but that's wrong, it is an electronic device!

The screenshots say "Q8 QAT" because I misspelt the model names in my config — it's a regular Q8, not QAT.

Full response (non-thinking): gemma-31b-nonthink-1.jpg | gemma-31b-nonthink-2.jpg

✅/❌ Gemma-4-31B thinking: Packaging debris, medication, or... RFID tag (the correct answer🎉)

Pass one: multiple theories

It hypothesized the following. All reasonable guesses:

  1. Industrial or Packaging Debris
  2. A Pharmaceutical Capsule Fragment
  3. Electronic Component (Solder Sleeve)

It was, however, completely wrong on the social context. It believes the author is a conspiracy theorist:

Regarding the "New Phobia" context: Posts like this often circulate on social media to suggest the object is a "tracking chip" or some form of surveillance device. The cost of producing millions of these "chips" to put in food would be astronomical and serve no practical purpose compared to tracking people via smartphones.

Pass two: small sensor, a capacitor, or passive RFID tag

It got it correct in one of the guesses, which is significant:

The clearer image allows for a better look at the construction. You can now see three distinct parts: a clear/translucent cap, a copper-colored center, and a black opaque end. This specific combination of materials strongly suggests it is an electronic component, likely a small sensor, a capacitor, or a passive RFID tag.

Overall, Gemma was better for vision and succinctness compared to the Qwens, especially the thinking version

The screenshots say "Q8 QAT" because I misspelt the model names in my config — it's a regular Q8, not QAT.

Full response (thinking): gemma-31b-think-1.jpg | gemma-31b-think-2.jpg

Caveat: I made multiple passes for each LLM and 31B thinking only got the RFID tag as possibility in one of them. Also, the Q4 QAT versions never got it right. Only the Q8 version ever got close.

✅🎉 Meta Muse Glimmer 30b

The new kid on the block got it exactly right in the very first pass! This is a huge milestone for open models.

Excerpt:

It's a glass-encapsulated RFID transponder, the kind of microchip that's routinely implanted for animal ID / pet tracking. What you can make out in the photos is the classic 1-2 cm long bioglass capsule with: - a clear glass / polymer body, - a copper / orange ferrite antenna core in the middle, and - the characteristic blackened end cap. That form factor is the standard ISO 11784/85 pet microchip / implantable RFID tag.

Its achievement is even more impressive because this is a 4-bit quant at ~20 GB! The Gemma model that kind of sort of came close was an 8-bit quant at ~32 GB.

Full response: muse-glimmer-30b-think.jpg | muse-glimmer-30b-think-cot.jpg

Verdict

  • Muse Glimmer 30b is the king of open image vision models. Don't leave home without it
  • Skip Qwen 35b and 27b for complex vision tasks. However, they're perfectly capable at vision in regular contexts (text, non-gotcha images)
  • Gemma 4 31b is better than the Qwens, but only in thinking mode, and only at higher quants like Q8. Lower quants and non-thinking mode are decent, not amazing
  • Keep Qwen3-VL 8B in your pocket for quick vision tasks or if you're low on VRAM/disk space
An earlier draft of this article said "If you really want to be sure, use a frontier model". With the release of Muse Glimmer, I don't think it's true, unless you find an extremely unusual / tricky puzzle.

Appendix A: Proprietary models

Here are the proprietary models I tried.

✅ The winners: ChatGPT 5.5, Gemini 3.X, Claude Opus 4.8 and Fable 5, Grok Expert, and DeepSeek Web

All of them succeeded, nothing much to say. For Gemini, I tested 3.1 Pro, 3.6 Flash and 3.5 Flash-lite — all of them succeeded.

The only surprise was DeepSeek because currently, DeepSeek V4 and previous models do not have vision capability whether you host them locally or access them through the API.

The only way to use vision with DeepSeek is through the web app, which presumably uses an unreleased, experimental vision add-on. Normally, this would suggests that it's not that good yet, but clearly it is, hence my disbelief.

Full responses: gpt-55.jpg | claude-fable-5.jpg | grok-4-expert.png | deepseek-think.jpg | gemini-36-flash.png | claude-opus-48.jpg

❌ The losers: Claude Sonnet 5, Grok Fast

I didn't have high hopes for Grok Fast, so this didn't surprise me.

However, Claude Sonnet 5's failure did surprise me. Normally, the Claude models have excellent vision capabilities (demonstrated by Opus and Fable's success, but otherwise, too). Sonnet thought it's an insect larva, which IMO is a worse result than the open models that thought it's medication.

Full responses: sonnet-5-1.png | sonnet-5-2.png | grok-4-fast-1.png | grok-4-fast-2.png

Appendix B: Large open models

These are otherwise very capable open models, but currently too big to host for me, so I used their web apps.

❌ Kimi K3: medication cartridge/ampoule

Just like most open models. Unfortunately, someone else tested it for me, but I lost the screenshot of the full response. Trust me though.

❌ MiMo v2.5: small battery → glass fuse

MiMo started with "small battery" on the first image, then upgraded to "glass fuse (cartridge fuse)" on the close-up — correctly identifying the transparent body, amber internal element, and dark end cap.

Wrong answer but a solid effort, and it nailed the meme/prank context too.

Full response: mimo-25-1.png | mimo-25-2.png