Reflection previews Beam before releasing its weights
The 501-billion-parameter mixture-of-experts model is available only through early access. Reflection says weights, a technical report and developer tools will follow later in October.
The 501-billion-parameter mixture-of-experts model is available only through early access. Reflection says weights, a technical report and developer tools will follow later in October.
A preprint finds a shared internal direction that points language models to the first, second or later fact in a passage. Moving a question along that direction can change which fact the model retrieves.
A preprint finds that much of the benefit from an agent's history survives after its past actions are shuffled. Explicitly pairing each action with its result improves task completion.
The open-source command-line tool embeds llama.cpp and ships two small fine-tuned models. Its author also released the 401,975-pair training set and benchmark code.
Context Language Models replace an append-only transcript with a file the model can rewrite, delete and reorder. The authors report better long-task accuracy with less repeated computation after training the editing policy.
A practical guide measures the trade-off in separating prompt processing from token generation. Tail latency improves under load, but moving the model's working memory delays the first token.
Today's remaining AI news covers an unusual hybrid model, spatial-memory and honesty studies, community measurements, agent security incidents, EU watermarking and practical talks.
A preprint adapts a clinical spatial-memory test for 16 vision-language models. All of them can recognise a landscape from the angle they studied, but most fall to guessing once the camera moves.
A preprint tests eight published probes on language models playing characters who reject basic facts. Many fail once true and false answers share the same prompt, and a probe trained to separate truth from obedience holds up.
The open translation family now has GGUF, FP8 and NVFP4 packages, a free compatible API and four public benchmarks. Its performance numbers remain the developer's own.
A controlled study finds that one sentence of confidence or doubt can sharply change whether a reasoning model calls a tool, but the changes rarely target the problems where help is needed.
An interpretability study finds the same abstraction, induction and retrieval pathway before few-shot accuracy rises, suggesting demonstrations energise existing machinery instead of building a new algorithm.