LensVLM-9B by Apple

- Apple: 92 events in the last 90 days
- Previous: 1 days earlier · I said no and Apple said yes
What happened
Vision-Language Models can process text as rendered images, but accuracy degrades with compression; LensVLM addresses this by scanning compressed images and selectively expanding relevant parts through learned tools, maintaining high accuracy even at high compression ratios.
Summary assembled by rule from the sources below