← Back to events
ActiveTech

LensVLM-9B by Apple

Photo: Hacker News

What happened

Vision-Language Models can process text as rendered images, but accuracy degrades with compression; LensVLM addresses this by scanning compressed images and selectively expanding relevant parts through learned tools, maintaining high accuracy even at high compression ratios.

Summary assembled by rule from the sources below

Why it's spreading

Sources

Community