Kimi K2.7 Code delivers a 21.8% improvement in real-world coding benchmarks, costing 13¢–78¢ per prompt with mixed speed and ...
Florida is employing AI-powered robot rabbits to help locate invasive Burmese pythons. These solar-powered decoys mimic prey ...
ARC-AGI-3 benchmark gains its first fully open-source agent: NIMI's Tycho writes Python code as falsifiable hypotheses about ...
Researchers found AI coding agents build less reliable pipelines when forced into structured formats — DataFlow-Harness ...
The performance of many next-generation devices depends on controlling how energy flows at extremely small scales. In the ...
Anthropic reviewed 141,006 of its own test runs after OpenAI's Hugging Face hack, and found three Claude models had broken ...
Anthropic found three cybersecurity evaluation incidents in which Claude models gained unauthorized access to real organizations.
Wrote and published malware during tests, which is apparently OK because leaky test environments were the real problem ...
On TerminalBench, Sarvam Code solved nearly as many tasks as leading closed-model coding systems. It also scored 82% on Data ...
Anthropic disclosed Thursday that three of its Claude models gained unauthorized access to the production systems of three organizations during cybersecurity testing.
Frontier AI systems are increasingly capable of translating narrowly defined objectives into complex, real-world cyber ...
According to Anthropic, the third cybersecurity incident involved an unnamed “internal research test model.” It compromised ...