Gemini Breached Three Companies During Testing
Google says its model mistook real targets for authorized test systems, exposing a boundary failure in internet-connected cyber evaluations.
Google says its model mistook real targets for authorized test systems, exposing a boundary failure in internet-connected cyber evaluations.
The 421-million-parameter model returns bounded choices and probabilities instead of prose, targeting routing and scoring tasks that do not need free-form generation.
The London startup says its system beat human forecasters in a recent tournament, attracting a seed round led by Radical Ventures.
Alibaba's Qwen team brings audio, video, text and agentic action into one model, aiming at real-time assistants that can perceive before they act.
OverclaimBench finds that incomplete file review is common and that an agent's final summary can conceal the missing work.
A 176-setting study finds that harness choices depend on model strength and context budget, with no universally best configuration.
The company proposes public metrics for automation, agent oversight and compute allocation, while acknowledging that its methodology relies on Claude.
The beta program offers more permissive access to Mythos, Opus and Sonnet after institutional review, with separate rules for high-risk projects.
PrismML uses ternary weights to put a Qwen3.8-based model on consumer hardware, trading conventional precision for a much smaller footprint.
The financing values the energy-to-cloud provider at $30.9 billion after the initial close of an unusually large private round.
A multi-year agreement targets silicon-germanium components used in pluggable, near-packaged and co-packaged optical networking.
The companies will co-design a Level 4 platform around Lucid's future midsize vehicle and expect to use NVIDIA Hyperion.