Today’s wider coverage ranges from open audio-video generation and agent sandboxes to research on memory, safety and recurrent model depth. Preprint results, vendor measurements and community experiments remain attributed to their publishers.

Releases

KandinskyLab releases audio-video generation weights. Kandinsky 6.0 Lite and Pro generate five-second clips with synchronised audio from text or a first frame; the 3B Lite weights and code are MIT-licensed, while the 29B Pro repository has restricted access. The inspectable Lite artefacts make the release useful despite author-run evaluations. Direct source: huggingface.co/kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers

GPT-6 reaches all ChatGPT users. OpenAI says its new model is rolling out globally with “Intelligent UI,” which can produce interactive answer components, and the API’s chat-latest alias has changed alongside it. No technical report or model card accompanied the announcement, so this is a product rollout rather than an assessable model release. Direct source: openai.com

Research

Queen links chess positions to language explanations. Princeton researchers connect a four-billion-parameter chess encoder to a language model through cross-attention, aiming to preserve strong play while explaining moves. The authors report roughly grandmaster-level chess, giving researchers an inspectable test of whether strategic representations can support language. Direct source: arxiv.org

SafeActBench measures the step from evidence to safe action. The benchmark tests whether a model that can recognise a risk also changes its behaviour appropriately. That separation matters for systems whose safety evaluations currently stop at verbal awareness. Direct source: arxiv.org

Memadapter studies memory-induced agreement. The preprint examines how retrieved memories can make an assistant follow prior claims even when they conflict with current evidence. It is worth attention because persistent agent memory can amplify sycophancy across sessions. Direct source: arxiv.org

Looped models approach stable internal states. Researchers repeatedly apply shared model blocks and analyse when the hidden state converges toward a fixed point. Their reported speed and training gains remain author results, but the mechanism offers a concrete account of computation through recurrent depth. Direct source: arxiv.org

Multilingual GSM-Symbolic tests reasoning beyond English. The dataset varies grade-school mathematics problems across languages and symbolic forms to separate memorised phrasing from calculation. It gives multilingual evaluations a more controlled stress test than translated fixed questions. Direct source: arxiv.org

LoopCD treats model depth as a continuous process. The method trains a repeated block to refine representations over a variable number of steps rather than assigning every layer separate parameters. The paper is useful for comparing recurrent computation with conventional depth under a shared parameter budget. Direct source: arxiv.org

Prompting techniques

Terse makes concise Claude Code style persistent. The plugin packages reply rules, document conventions, subagent instructions and a context meter; its author reports 46–54% fewer words across a 20-prompt test. Those cost and style results are self-measured, but the repository is a practical example of moving durable behaviour into a plugin. Direct source: github.com/lowenbjer/claude-terse

An AWS talk separates four context operations. Elizabeth Fuentes Leone groups agent context work into externalising, selecting, compressing and isolating information, demonstrated with Strands Agents. No transcript was checked, so the English video is a technique pointer rather than verified benchmark evidence. Direct source: youtube.com

What people are building

nanoMuse links a self-hosted phone and desktop agent. The GPL project combines Android, desktop and web clients with editable Markdown memory and approvals for consequential actions. Version 0.1.41 is the new development; security and capability claims remain the team’s own. Direct source: github.com/nano-muse/nanoMuse

A browser walks through every number in a tiny GPT. Manogya Singh’s visualisation computes a miniature transformer’s forward pass, backpropagation and one reinforcement-learning step with inspectable arithmetic. It is teaching material rather than a new method, and the repository currently lacks a licence file. Direct source: github.com/manogyasingh/transformer-visualisation

An agent-built SimCity clone raises verification questions. Its creator says Claude Opus 5.5 used 27 subagents to translate original PowerPC code into a WebGL version in about two hours, then compared game state through an emulator harness. The playable demonstration exists, but the process and fidelity claims are unverified and no repository was found. Direct source: dator.dev

Worth reading

Cloudflare keeps evidence collection outside the model. Its security-operations design performs a fixed reconnaissance pass in code, then gives narrow agents evidence with explicit states for absent, present and unchecked data. The account lacks product accuracy figures, but the separation between collection and judgement is portable. Direct source: blog.cloudflare.com

Conflicting training values can override stated reasoning. A research note reports models producing a final answer that contradicts their visible chain of thought after training on clashing traits. Rates outside the constructed setup were low, making the work a measurement of monitor limits rather than evidence of a general hidden plan. Direct source: lesswrong.com

NVIDIA details its olympiad-specialist training. The company describes curated programming data, generate-evaluate-refine loops and separate supervised and reinforcement-learning checkpoints behind reported gold-level IOI and IMO results. Checkpoints and datasets are published, while the competition comparisons are NVIDIA’s own. Direct source: huggingface.co/blog/nvidia

Epoch estimates rapid growth in internal coding-agent use. Values extracted from OpenAI charts suggest the median researcher’s daily usage, priced at API list rates, roughly doubled each month through mid-August. The figures measure implied usage rather than OpenAI’s actual cost and come from one company. Direct source: epoch.ai

An essay separates alignment engineering from misalignment science. Edward James Young argues that optimising current safety metrics can improve capabilities without building an understanding of failure, and calls for more diagnostic science. It is an opinionated map of the field, useful for its categories rather than as empirical proof. Direct source: lesswrong.com

A short essay asks what mathematical creativity means. Raemon proposes examining whether OpenAI’s new proofs contain concepts that do real work or mainly search combinations of known ideas. The post offers a testable question and requests detailed mathematical analysis rather than claiming an answer. Direct source: lesswrong.com

Hacker News

Readers dispute what a formalised fluid proof represents. A rebuttal argues that a Lean proof tied to OpenAI’s Navier–Stokes work does not correspond to the stated natural-language argument. Commenters distinguish validity inside Lean from the separate claim that the formal statement captures the intended proof. Direct source: news.ycombinator.com

GPT-6’s generated interfaces divide practitioners. Some commenters praise interactive explanations, while others criticise the layouts and speculate that generated interfaces could displace browser conventions. The reactions are product impressions, not a controlled usability study. Direct source: news.ycombinator.com

A formal repository tackles packing 11 squares. The thread examines an AI-assisted Lean proof for the optimal square-packing problem and links visual explanations. It adds an inspectable example to the discussion of AI-assisted formal mathematics without independently validating the whole development process. Direct source: news.ycombinator.com

Reddit

llama.cpp experiments with a GPU cache for MoE experts. A pull request keeps mixture-of-experts weights in host memory while caching active experts on a limited GPU. Users expect gains but had not published stable benchmarks, so the thread is an early tuning report. Direct source: old.reddit.com

A retrained draft model speeds Ternary Bonsai 2. The author reports 1.2–3.2 times faster decoding across browser, Mac, L4 and code-edit tests after training DFlash 2 on the target model’s output. Weights and setup are public, but the heterogeneous self-run measurements are not a general comparison. Direct source: old.reddit.com

MacroStories compresses a narrow language model to 20K parameters. The 81 KB model reportedly writes short stories inside a small training distribution using a shared recurrent decoder block. Its interest is the extreme size and microcontroller discussion, not broad language competence. Direct source: old.reddit.com

Repurposed mining cards run a long-context model. A user reports about 90 tokens per second for GLM-5.3-Flash at 384K context on two modified CMP 170HX cards. Different engines, quantisation and speculative-decoding settings make the comparison anecdotal, but the build details help local inference work. Direct source: old.reddit.com

YouTube

Periodic Labs argues for physical experiments. In an English Latent Space interview, Liam Fedus and Ekin Dogus Cubuk discuss reinforcement learning from laboratory work, noisy measurements and materials discovery. The summary rests on chapters and description rather than a transcript, so the scientific claims remain the speakers’ views. Direct source: youtube.com

Runlayer proposes agents that write reusable skills. Rafal Wilinski describes serving governed skills over MCP and distilling new ones from successful and failed runs. A reported improvement from 36% to 100% is a single vendor test; the English talk is most useful for its architecture. Direct source: youtube.com

Langfuse lets coding agents inspect the improvement loop. Marc Klingen demonstrates an agent updating evaluation data after finding jargon in generated changelogs. The English vendor talk shows how tracing, datasets and back-tests connect, while its outcomes remain product demonstrations. Direct source: youtube.com

Tessl divides software factories into three loops. Dru Knox describes inner, outer and meta loops and proposes measuring takeovers, human review comments and autonomously started pull requests. The English presentation is partly a vendor pitch, but its operational metrics are concrete. Direct source: youtube.com

Google presents sandboxed Gemini agents. Philipp Schmid demonstrates the Interactions API, persistent timelines and an agent that reads skills and AGENTS.md files inside a managed sandbox. The English talk is a product demonstration rather than independent evaluation. Direct source: youtube.com

A former OpenAI safety writer explains his resignation. David Robinson tells Ezra Klein that release pressure and financial incentives outpaced the safety culture he considered necessary. The English interview is one insider’s account and should be read as testimony and opinion. Direct source: youtube.com

Apollo reconstructs damaged Greek papyri. A German-language DER STANDARD episode interviews papyrologist Bernhard Palme and project lead Anna Dolganov about an Austrian model for filling gaps in historical Greek text. Claims that it sometimes suggests solutions missed by researchers come from the programme description. Direct source: youtube.com

AI v kostce compares the latest assistants. The Czech-language episode tests OpenAI and Claude products on documents, development, images and security work, then discusses lab leaders’ calls to slow deployment. Its value is hands-on regional commentary rather than controlled benchmarking. Direct source: youtube.com

In brief

ARTEX appears in South Korean bank breaches. CrowdStrike attributes use of the open-source penetration-testing agent to a probably Chinese-speaking, financially motivated actor and says sessions used DeepSeek, GLM and Grok models through Claude Code. The attribution carries moderate confidence and the extent of stolen data remains unclear. Direct source: crowdstrike.com

LMCache has a critical unauthenticated code-execution flaw. JFrog says versions 0.3.9 through 0.5.5 and 0.5.6 release candidates can unpickle a crafted ZeroMQ message before checking its type, with no fixed release yet. Operators should keep the service on trusted networks while a patch is pending. Direct source: thehackernews.com

PoeLLM hides command infrastructure in a poem. Lumen Black Lotus Labs reports malware deriving changing command addresses from text hosted on GitHub and infecting more than 3,000 exposed LLM-related servers. The technique deserves attention because ordinary link scanning sees neither a URL nor executable payload in the poem. Direct source: theregister.com

AWS releases a history-aware agent sandbox. Strands Box can apply rules based on the current tool call and the agent’s earlier actions, including rate limits and conditions on pushes or expensive APIs. The open design is inspectable, while enforcement claims are AWS’s own. Direct source: theregister.com

Microsoft adds local-cloud routing for Windows agents. HydraFusion chooses between on-device and hosted models, while Execution Containers apply permissions to local agents accessing PC files. Rollout begins with GitHub Copilot, making containment behaviour the part to watch as access expands. Direct source: theregister.com

Teen-safety assessments conflict. OpenAI says typical teen use is brief and focused on learning, while Common Sense Media says parental alerts and other safeguards often fail. The BBC account is a pointer to competing measurements rather than a resolution. Direct source: bbc.co.uk

Google opens a public SynthID detector. The site checks uploaded media for watermarks from Google and several partners, returning a yes-or-no result after login. Wider access makes provenance checking easier, though absence of a detected watermark does not establish human authorship. Direct source: arstechnica.com

Singapore issues financial-sector AI risk guidelines. The Monetary Authority of Singapore expects independent review of every AI use case and places responsibility with boards and senior management, including for third-party systems. Only secondary reporting was available for this digest, so the exact regulatory language needs the MAS document. Direct source: theregister.com

RTX Spark puts unified-memory AI systems in laptops. New machines pair Arm CPUs with NVIDIA Blackwell graphics and offer up to 128 GB of unified memory, with listed prices from $2,599 to $7,000. Availability and independent sustained-performance tests will decide whether they serve local models better than desktop alternatives. Direct source: theverge.com

Business, briefly

McDonald’s faces a proposed US class action alleging that its franchise pricing platform enables unlawful coordination; the company says the tool does not automate or fix prices. Direct source: theguardian.com

Nous Research confirmed a $1.5 billion valuation and launched agents for business users, a funding and product-positioning event without a separately assessable technical release. Direct source: techcrunch.com

What this suggests: Today’s secondary stories repeatedly move the boundary around an AI system: which context it sees, which evidence it can trust, which hardware it reaches and which actions its sandbox permits.

What’s next: LMCache still needs a fixed release, Microsoft is beginning its Copilot file-access rollout, and several research preprints now await code review or independent replication.

Verification

Claim groupLabelPrimary sourceIndependent check
Releases and project existenceVERIFIEDDirect links in each itemnone unless named
Model, benchmark and latency performance claimsVENDOR-REPORTEDDirect links in each itemnone
Research findingsVENDOR-REPORTEDLinked papers and research notesnone
Community measurements and demonstrationsUNVERIFIEDLinked HN and Reddit threadsnone
Essay arguments and interview positionsOPINIONLinked essays and videosnone
Security and policy eventsPARTIALLY VERIFIEDLinked vendor reports and independent coverageSources named in each item