AI benchmarks have a trust problem and Google wants to fix it

- Google 近 90 天出现 12 次
- 上一次:同一天稍早 · Marvell Technology shares slump as investors question Google deal. Here’s what Wall Street is saying.
发生了什么
Google Deepmind is testing a double-blind evaluation of a frontier AI model for the first time. Cryptographic protection through Confidential Space is meant to keep Google from seeing the test questions and keep evaluators from seeing the model weights. The pilot project with the Singapore AI Safety Institute uses a Gemini Flash Lite and could set a new standard for tamper-proof AI benchmarks.
摘要按规则整理自下方来源原文