September 2026: The Factory Floor
July said verification was the limit on how much you can delegate. August built the machinery underneath it, and sent the invoice. 83 of 240 matrix items were rewritten, and three surplus rows were removed.
At a glance
Why: Claude Code made auto mode the default on August 14, six weeks after making Manual the default. The numbers behind the reversal: developers approve 97% of permission prompts, and in a 1,053-person study humans caught 13.6% of dangerous commands where a classifier caught 89%.
Why: Vercel published a working software factory in which every agent run terminates as success, flawed, blocked or manual, and each failing class routes to a different repair. Only success ships.
Why: A study of 3.52M changes in one enterprise codebase found AI-generated code carries higher coupling and copy overhead, showing up as 5-8% more compute consumed in production. Meta's telemetry pairs a 220% rise in changes with a 40% rise in incidents.
Why: The ChainDrop compromise planted persistence in .vscode/tasks.json and .claude/settings.json, so the payload ran on folder open with no install step. The package signatures were valid; they proved the source, not its safety.
Why: Across 8.1M pull requests, agent-authored ones merge at 79% in the strongest organisations and 37% in the rest. The variable is not tooling; it is whether a person's name is on the change.
Why: A controlled comparison found agents told to write tests first scored consistently worse at three to eight times the token cost, because they harden a bad design around the first test. The sample is small and the author says so, but mutation score is the better instrument either way.
What changed in the model
Three changes to how the model is written. Each one changes how your own score should be read.
Some practices were claimed by two areas at once, so the same strength could lift two scores. Each capability now belongs to exactly one area. If your score moves slightly against last month, this is the most likely reason, and the new number is the honest one.
Where a named product or framework had become the bar, teams that solved the same problem another way scored lower for no good reason. A dependency-graph build with a shared cache is the criterion; Bazel, Nx and Turborepo are all ways to meet it. The same applies to delivery metrics, where DORA is now one option among equivalents.
Bars like "50+ pull requests a day" or "400+ tools" measured headcount rather than maturity, and a small team could not reach them however well it worked. Those now read against your own baseline before agents. Percentages of your own organisation were left as they were, because they already scale.
Every change, by area and level
2026-07 2026-08 - July taxonomy preserved unchanged.
Development
Coding Agent Usage
Context Engineering
Code Review & Quality
Testing Strategy
Delivery Management
CI/CD Pipeline
Merge & Deploy
Metrics
Governance & Compliance
Organization
AI Adoption Model
Knowledge Management
Team Structure & Roles
Tech Debt & Modernization
Infrastructure
Agent Runtime & Sandboxing
MCP & Tool Integration
Build System
Observability & Feedback Loop
The month in numbers
Context behind the edition. The figures that drove specific changes are cited with those changes above.
report any / significant EBIT impact from AI - both flat year on year
McKinseysecurity pass rate for generated code, unchanged in a year - model upgrades do not buy it
Veracodeof developers use coding agents weekly / daily, across 15,000 surveyed
JetBrainsof accepted AI completions are deleted outright, most within 15 minutes
arXivof AI-written vulnerability patches fully fix the flaw without changing behaviour
1Passwordin Linux 7.2 tagged Assisted-by, with some maintainers stripping the tag
LWNUpdated Guides
83 of 240 matrix items were rewritten this edition and three surplus rows were removed. 72 guides changed; the 44 below are the ones rewritten around new August evidence.
How the matrix is built
4 perspectives x 4 areas x 5 levels x 3 items = 240, and 240 is also the guide count. Every item owns exactly one guide and no guide is claimed twice. Levels run Assisted, Delegated, Systematic, Governed, Self-improving.
The matrix is what we show: one sentence per item, in a practitioner's language. The maturity gates are what we assess: Must, Should, Prerequisites and Evidence per area and level, and they decide the workshop score. Both layers must say the same thing; a gap between them is a defect, not a nuance.
- Describe the capability. A vendor or a methodology belongs in parentheses as an example, never as the bar.
- A threshold must not encode company size. A percentage of your own organisation scales; an absolute volume measures headcount.
- L1 describes a floor a team actually stands on, never an absence. Deficit wording is blocked by a test.
- Each capability has one owning area, so nobody scores twice for one practice.
Every edition is checked automatically against these rules before it ships: that the matrix and the assessment still say the same thing, that no capability is claimed by two areas, and that nothing is displayed which the assessment never actually scores. Previous editions are never edited, so a score from an earlier month stays comparable to what it meant at the time.
The level names promise a delegation and control arc, but CI/CD Pipeline and Build System still ladder on latency: their L4 criteria are wall-clock targets with no governance in them. Re-laddering those two areas would move existing workshop scores, so it is a decision about the model rather than a tidy-up.
What Didn't Change (and Why)
Sources
OpenAI: Separating Signal from Noise
SWE-Bench Pro recommendation formally retracted (~30% of tasks broken)
Cursor: reward hacking audit
63% of resolutions retrieved, not derived; 87.1% → 73.0% strict harness
Building to the Test (arXiv)
Agents ship dead code passing a visible 222-test oracle
Hugging Face: July security incident
The first runaway agent - eval model breached production via sandbox zero-day
Willison: the first known runaway AI agent
Analysis of the ExploitGym → HF breach
Pillar: The Week of Sandbox Escapes
Escapes plant files trusted host tools later read - four failure modes
Claude Code v2.1.200: default Manual
Default permission mode flipped from Auto to Manual (July 3)
GitHub: enterprise managed settings
'Any client outside the policy is a gap' - policies cover app + cloud agent
MCP 2026-07-28 release candidate
Stateless core, OAuth/OIDC auth, deprecation policy - biggest revision ever
Azure DevOps MCP flaw
Hidden PR comments hijack AI review agents (official Microsoft server)
AWS Kiro CVE-2026-10591
Poisoned web page makes the agent rewrite its own MCP config
Addy Osmani: Own the Outer Loop
Quality, Verdict, Answerability - the human's floor
Addy Osmani: Software Factories, Light and Dark
Back pressure and comprehension debt
Pragmatic Engineer: What is loop engineering?
The legitimizing deep-dive, with the skepticism kept in
Lilian Weng: Harness Engineering
Recursive self-improvement reframed as harness improvement
Fowler: Harness Engineering formalized
agents.md under 200 lines; the discipline gets its name
Cherny: Steps of AI Adoption
Gated → AI-native; each step breaks an org bottleneck
Bun: Rewriting Bun in Rust
535k LOC in 11 days for $165K; TS suite as conformance harness
Anthropic: new rules of context engineering
80%+ of the system prompt deleted; auto-memory over manual curation
RIT: agentic PR review study
78.9% of agentic PRs pass a single reviewer; more review changed nothing
curl: Summer of Bliss
Vulnerability intake suspended five weeks over AI slop
Microsoft rollout study (arXiv)
+24% merged PRs for CLI-agent adopters, sustained over four months
Agent PR conflicts (arXiv)
Cross-vendor pairs conflict at 41.7% vs 19.8% intra-vendor
CNBC: Fable 5 export controls lifted
Restored July 1 after a three-week worldwide blackout
Gibson Dunn: EU AI Act omnibus
High-risk duties deferred to Dec 2027 / Aug 2028
CNBC: the layoff reversal wave
Ford, CBA, IBM reverse AI cuts; 55% of leaders call them a mistake
InfoQ: AWS Claude Apps Gateway
Self-hosted identity/policy/telemetry control plane for Claude agents
Malicious skills defeat scanners
Static skill scanning evaded >90% of the time (HKUST)