Strata fork revives an IBM AI server for local LLMs

A hardware-specific fork runs Qwen3.8-Flash-Next across two POWER9 processors and four V100 GPUs, showing what model-aware optimisation can recover from a 2018 system.

5 October 2026 · 3 min · Martin Seckar

Wagtail's one-model month spent half its tokens elsewhere

A plan to run September engineering on GLM 5.3 Flash consumed two billion tokens, but only half reached the target model. Prototypes, provider capacity and evaluation explain the gap.

5 October 2026 · 3 min · Martin Seckar

Matthew Green says agent sandboxes need a warden

The cryptographer argues that containment remains necessary but cannot solve the hardest part of agent security: deciding which information and instructions are authorised.

5 October 2026 · 3 min · Martin Seckar

AI Daily Digest for 5 October 2026

Speaker diarisation, model-cognition papers, local projects, prompting practices, community tests and the day's safety and policy developments.

5 October 2026 · 16 min · Martin Seckar

Aleph Alpha releases open-weight Kolibri

The German-English reasoning model ships with Apache-2.0 weights and serving instructions. Its small active parameter count still leaves a large memory requirement.

4 October 2026 · 2 min · Martin Seckar

Zou team traces how models report internal changes

A controlled preprint separates detecting an activation change from reporting its location. The experiment uses fixed input text and scores answers without an AI judge.

4 October 2026 · 4 min · Martin Seckar

Georgia Tech team extends model reference tracking

A small trained intervention sharply improves Qwen3-8B on synthetic reference chains. The preprint traces the change to how middle layers pass information between lines.

4 October 2026 · 4 min · Martin Seckar

Mingbird adds skills to its local agent harness

The Ollama-based project adds voice input and everyday document tools. Its published benchmark argues that small-model results depend heavily on the surrounding software.

4 October 2026 · 3 min · Martin Seckar

Kevin Liao proposes document-based agent memory

The developer's workflow keeps specifications and decisions in editable Markdown. His argument comes with an open-source implementation, but no comparative benchmark.

4 October 2026 · 3 min · Martin Seckar

Simon Willison calls for default agent spending caps

His proposal favours automatic shutdown at a budget limit. Removing the limit would require an explicit choice.

4 October 2026 · 1 min · Martin Seckar

AI Daily Digest for 4 October 2026

Research, projects and discussions beyond today's six articles, with source links and evidence labels.

4 October 2026 · 23 min · Martin Seckar

Percepta tests growing memory for long-context recall

Spotlight Memory addresses a small region of expandable storage at each step. Percepta reports strong retrieval beyond the models’ training length.

3 October 2026 · 4 min · Martin Seckar