AI Pulse — Week 35, 2026
Two security results landed the same week: a credible researcher broke Claude Code’s auto mode roughly 80% of the time, and an OCaml maintainer clocked exploit probes ten minutes after posting a patch for discussion. Debian, meanwhile, voted to allow AI-assisted contributions and declined to require disclosure.
The default agent permission mode did not hold
Johann Rehberger found an attack against Claude Code’s auto mode that he claims works about 80% of the time. Anthropic recently made auto mode the default and has made strong claims about it stopping prompt injection. The attack does not fight the model at all. It gets Claude Code to download and extract a zip archive, then run code that imports base64 — and the archive contains a malicious local struct.py, which Python loads first because of import path order. Extracting an archive into the working directory is the code execution.
One detail is worse than the bypass. Rehberger reports that Claude sometimes detected the compromise and auto mode then blocked its own cleanup command.
Why it matters: the permission model here guards the wrong verb. It asks whether a command looks dangerous, and python -c "import base64" does not. The dangerous step already happened when untrusted bytes landed on a path the interpreter searches. I have been treating “extract this archive” as a read-only operation in my own head, and it is not — on any language with a path-relative import order, unpacking a file is arming it. The blocked-cleanup part is the design lesson I would generalize: a guard that can veto remediation but not the initial compromise has inverted its own purpose. If you run agents with a permission mode, check that the allowlist covers recovery paths, not just entry points.
The disclosure grace period is now about ten minutes
Anil Madhavapeddy, a Cambridge CS professor and an OCaml compiler maintainer, reported that security issues in OCaml projects now see attempted exploits within minutes of a patch being shared for discussion — not after release, after discussion. His line: “Within about ten minutes (!) this website was fielding probes for percent-encoded traversal sequences, indicating that automated watchers are keeping an eye on public repositories.” He attributes it to automated watchers plus coding agents that are now good at deriving a flaw from very little information.
In the comments, rclone maintainer Nick Craig-Wood added a second datapoint: rclone received more than 40 security disclosures in one month, against roughly 20 in the project’s first decade.
Why it matters: coordinated disclosure has always rested on an unstated assumption — that turning a patch into a working exploit costs the attacker real effort, which buys defenders a window. That cost just fell, and the window closed with it. This pairs directly with the auto-mode result above: the same capability that makes an agent good at finding your bug from a diff makes it good at finding the gap in your agent’s permission list. The practical change I would make is small and unglamorous: stop discussing security fixes in public before the release is ready to ship. The habit was safe because exploitation was slow, not because discussion was private.
Craig-Wood’s numbers are worth holding loosely — a jump in disclosures measures reporting volume, which now includes a lot of agent-generated noise, not only real severity.
Debian wrote its AI policy down, and chose obligations over provenance
Debian resolved a general resolution on generative AI on 29 August. Choice 5, “Responsible Use of Generative AI”, won; both of the strictly anti-AI options failed to clear the “None of the Above” threshold. The resolution neither endorses nor prohibits AI tools. It requires that contributors “understand, review, test, and, where appropriate, modify AI-assisted output before incorporating it into Debian,” and holds submissions to the existing bar for quality, correctness, maintainability and legal compliance. It does not mandate disclosure of AI use, and it deliberately takes no position on the copyrightability or licensing of AI-generated material.
The scale question sits next to it. Simon Willison quoted Paul Dix on a port where AI wrote a million lines that were then refined over a couple of months into software now running on millions of developer machines. Dix pre-empts the obvious objection — that a language-to-language port has an oracle to check against, so it is the easy case — and thinks that undersells it anyway. He does not name the project in the quoted passage, so I would treat it as a striking anecdote rather than a measured result.
Why it matters: the no-disclosure choice is the interesting one, and I think it is correct. A disclosure tag tells a reviewer where a patch came from, which is exactly the fact that does not determine whether the patch is right. Debian instead put the obligation on the human submitting it — understand it, test it, be able to maintain it — which is what they already demanded of hand-written contributions. That is a policy that stays true as tools change, and it does not require anyone to police an unverifiable claim about provenance. The part they punted on is the part nobody can settle alone: if AI-generated material turns out not to be copyrightable, a distribution built on licence grants has a real problem, and a project vote cannot fix it.
Open weights got much bigger and much sparser in the same week
Two Chinese labs shipped, and the shape of both models is the story. Tencent’s Hy4 Preview is 770B total parameters, 49B active, with a 1M token context window, and it is 1.56TB on Hugging Face. Their Hy3 in July was 295B total, 21B active, 256,000 context, 598GB — so total parameters roughly 2.6× in about six weeks. Alibaba’s Qwen3.8-Flash-Next is a multimodal MoE at 125B total with only 6B active, described as an early preview of the architecture Qwen4 will use. Willison ran it locally on a DGX Spark using Unsloth quantizations, preferring the 78.9GB UD-Q2_K_XL build over the 72.5GB UD-IQ1_S.
The rest of the stack moved in the same days:
- Ollama v0.33.1 added MLX support for Qwen3.8 Flash Next, and structured output in the MLX runner — days after the model landed.
- vLLM v0.28.0 shipped 584 commits from 270 contributors, most of it a Kimi-K3 performance push: decode context parallelism, fused decode and prefill kernels, and combined all-gathers.
- IBM published how Granite 4.2 was built.
Why it matters: active-parameter ratio is the number I would actually track. Hy4 activates about 6% of its weights per token and Qwen3.8-Flash-Next under 5%, which means serving cost tracks a small model while the thing you must physically hold tracks a huge one. That splits the audience cleanly. If you rent inference, sparsity is a discount. If you run locally, 1.56TB is the whole story and the active count is irrelevant — which is why the interesting local news was a 78.9GB quantization, not a 770B release. It also explains why Ollama shipping MLX support within days matters more than it sounds: for open weights, the lag between release and something runnable on hardware you own is the real distribution channel.
Evaluation got cryptographic plumbing, and a new axis to fail on
Google DeepMind is piloting what it calls the first double-blind evaluation of a proprietary frontier model, with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. A Gemini Flash Lite model is tested against confidential benchmarks inside Google Cloud’s Confidential Space: evaluators cannot see the weights, and Google cannot see the test prompts. The point is contamination — the previous arrangement relied on contracts and zero-logging promises, and this replaces those with cryptographic attestation.
The Open ASR Leaderboard added Hindi and Indian English, its first Global South languages, using a Monsoon dataset built by Voice Arena and Hugging Face from 4,888 speakers across hundreds of districts, with twelve demographic attributes per clip and nine axes of variation. The headline result is a gap, not a ranking: eight leaderboard models sat within 0.18 WER points of each other on the aggregate score, yet one varied 1.68 points across Indian zones while another varied 0.46. Two systems that look identical on the leaderboard differ almost fourfold in how much their accuracy depends on where the speaker is from.
Why it matters: these are two halves of the same repair, and they follow directly from last week’s finding that ASR leaderboard leaders are partly fitted to their benchmarks. Double-blind evaluation fixes who can see the test. The Hindi datasets fix what the test measures — an aggregate WER averaged over a population hides exactly the variance you would feel in production, because your users are not distributed like the benchmark. The regional-variance number is the one I would put in a model card. “Indistinguishable on aggregate, fourfold difference in regional sensitivity” is a procurement-relevant fact that no leaderboard position carries. I would also note DeepMind is both the subject and a co-designer of its own double-blind pilot, which is a reasonable place to start and not yet independent evaluation.
Agent memory as derived state instead of retrieval
Jordy Zomer built Lemmalog, a Datalog engine that holds an agent’s knowledge as logical facts and rules rather than as conversation history to search. The motivation is a failure mode worth naming: during long vulnerability-research sessions the model “slowly lost track of what we had actually established,” resurrecting disproven hypotheses. His diagnosis is that this is a state-management problem, not a model problem. So the architecture mirrors static analysis — the LLM is the front end that turns messy observations into facts, and the engine derives conclusions, tracks provenance, and invalidates dependents when a fact is retracted.
The numbers are mixed and he reports them honestly. On LongMemEval, Lemmalog reaches 0.463 F1 against PropMem’s 0.550, while passing roughly 38× less context than the full-context baseline (about 2,700 versus about 104,000 tokens per question), and it leads on knowledge updates. On LoCoMo it trails again on F1 (0.533 versus 0.605) but is much stronger on adversarial questions (0.707 versus 0.509 full-context) and much weaker on inference (0.164 versus 0.289). His own takeaway is the deflating one: most of the gain came from better extraction and retrieval — entity reconciliation, alias matching, hybrid BM25 plus embeddings — not from the Datalog engine.
Why it matters: retraction is the capability I want and do not have. Every agent memory I run is append-only in practice, so a hypothesis I disproved an hour ago is still sitting in the context with the same weight as a fact I verified, and the model will happily pick it back up. Provenance and invalidation address that directly. But read the benchmark table before adopting: this loses on aggregate F1 to a simpler propositional memory on both suites, and the inference score is bad. The honest summary is a promising shape at 38× less context, not a win — and the author’s note that the wins came from retrieval plumbing rather than the logic engine is the part most likely to survive.
What I’m watching
- Whether the auto-mode bypass gets a structural fix or a pattern blocklist. Import-path shadowing is one instance of a general class — untrusted files landing where a runtime looks — so a fix that only knows about
struct.pyhas not fixed anything. - OpenAI wound down its contract supplying models to Cursor after Cursor’s acquisition by SpaceX. I could not load the announcement directly, so I am relying on OpenAI’s own feed summary for the fact and nothing more. Either way, model supply for a coding agent is now revocable on ownership grounds, which belongs in your dependency risk list.
- Whether any other distribution or foundation copies Debian’s no-disclosure line. It is the first vote I have seen that resolves the provenance question by declining to ask it, and if it holds up it is a much cheaper policy to run than an audit trail.
Sources
- Simon Willison — Breaking Claude Code Opus 5 Auto Mode
- Simon Willison — Just a rumour of a bug is enough to find a security exploit these days
- Hacker News — Debian votes to allow "responsible use of generative AI" (LWN)
- Simon Willison — Quoting Paul Dix
- Simon Willison — Introducing Hy4 Preview
- Simon Willison — Qwen3.8-Flash-Next
- Ollama releases — v0.33.1 (MLX Qwen3.8 Flash Next support)
- vLLM releases — v0.28.0
- Hugging Face — Granite 4.2 LLMs: How They're Built
- Google DeepMind — Piloting the world's first double-blind AI evaluations
- Hugging Face — The Open ASR Leaderboard Adds Its First Global South Language
- Hacker News — I accidentally turned LLM memory into program analysis
- OpenAI — Our decision on Cursor following its acquisition by SpaceX