启用自动模式后,独立评估中没有任何攻击成功攻击 Anthropic 的任何模型。而在 Codex v0.144.5 自动审核权限模式下运行的 GPT-5.6 Sol 的攻击成功率为 5.83%。GPT-5.6 Sol 在 Max ...
现有代码智能体评测,真的靠谱吗? 当前主流代码/AI智能体基准主要分两条路,一条是SWE-bench这类公开静态基准,任务直接来自公开GitHub ...
A critical vulnerability in IBM-owned, low-code AI builder Langflow lets unauthenticated attackers execute code remotely on ...
Kindle firmware 5.19.6 is rolling out now and quietly fixes two rendering bugs Amazon never acknowledged — restoring AZW3 ...
DOUBLECUP hides malware stages in cached PNG files, then uses ClickFix commands to deliver CountLoader variants and the ...