← 返回事件
持续讨论科技

LensVLM-9B by Apple

图:Hacker News

发生了什么

Vision-Language Models can process text as rendered images, but accuracy degrades with compression; LensVLM addresses this by scanning compressed images and selectively expanding relevant parts through learned tools, maintaining high accuracy even at high compression ratios.

摘要按规则整理自下方来源原文

为什么在扩散

来源

社区讨论