Radar: AI & dev — 2026-09-02
Part of: Radar
First edition of a weekly digest. Candidates are gathered automatically from the
feeds in scripts/news/sources.json and summarized by a local model; every line
below was then checked against the linked source by hand, which is where the
week’s actual theme showed up — the gap between finding a bug and exploiting it
has collapsed, and so has our ability to watch a model think.
Just a rumour of a bug is enough to find a security exploit these days
Anil Madhavapeddy, an OCaml compiler maintainer, saw probes for percent-encoded traversal sequences within about ten minutes of a patch being posted for public discussion — automated watchers reading public repos, with agents capable of turning a hint into a working exploit. rclone’s maintainer confirms the same shape: roughly 20 security disclosures in the project’s first ten years, and over 40 in the last month alone. The uncomfortable conclusion is not about any single bug: open source embargo practice assumes days, and now has minutes. If you maintain anything public, your disclosure process is already out of date.
Simon Willison · 2026-08-28Researchers fear safety disaster ahead of OpenAI’s Astra release
Astra reportedly uses a recurrent-depth (looped) transformer, cycling information through internal layers instead of emitting a readable chain of thought. That can improve performance, but it moves the reasoning somewhere nobody — researcher or automated safety system — can inspect. The practical angle for anyone building on top of models: chain-of-thought monitoring is a safety technique with an architectural dependency, and it stops working when the architecture changes underneath you.
The Verge · 2026-09-02Claude’s new system prompt really doesn’t want to reproduce song lyrics
Anthropic publishes the system prompts for its consumer apps — Claude.ai and the
mobile apps, though not Claude Code or Cowork — including the historic versions.
They are now split into a page per model, and appending .md to any page on
platform.claude.com/docs returns Markdown. That last detail is the useful one:
it makes the prompts diffable, so a change in behaviour can be traced to a
change in text rather than guessed at.
Claude Fable 5.1 made me a really nice animated pelican
Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1, well clear of the 20-30% band around it. Willison’s own note is the more interesting part: he had already written about losing faith in his pelican benchmark, so the post is as much about a benchmark aging out of usefulness as about the model that beat it. Worth reading if you maintain any eval of your own.
Simon Willison · 2026-09-01Claude Fable 5.1 is generally available in GitHub Copilot
Aimed at long-horizon, autonomous coding: deep codebase research, feature work, complex agentic flows. The line to read before enabling it is the one that is easy to skim past — unlike the other Claude models in Copilot, this one requires data retention. That is a policy decision for your org, not a model preference.
GitHub Changelog · 2026-09-01Fragments: September 1
Fowler points at Simon Willison’s LLM cliché highlighter: paste text or a URL and it flags the patterns common to LLM prose, referencing the Wikipedia page on signs of AI writing. Useful, with the caveat that page itself makes — humans write clichés too, so a hit is a prompt to look, not a verdict.
Martin Fowler · 2026-09-01Enterprise-grade precision for long-context multimodal embedding inference on Cloud TPU
vLLM on TPU with GKE for embedding workloads, with a runnable Qwen3-Embedding-8B example and the pooling-state fixes needed for requests to survive preemption. Relevant if you are scaling an embedding pipeline past prototype size; skippable otherwise.
Google Developers Blog · 2026-08-26On this digest’s own accuracy. Reviewing eight machine-written summaries against their sources changed five of them and removed one entirely. The removed item was a month old and had slipped in because its feed publishes no dates at all, so nothing could age it out — a gap now closed in the generator. The local model’s failures were not invented facts but confident vagueness: “provides valuable insight into how language models learn” in place of a post about diffing published prompts. The numbers it quoted did check out.
Found something that belongs here? Send it through the recommendations page.