October 2026 · v1.7 · October 1, 2026
What top engineers read about AI in September 2026
System One - The Cheap Brain and the Control Plane
The monthly reading list that pairs with the VISDOM AI maturity matrix. Essays, reports and papers that engineers actually read and argued about this month, and what they add up to. No chasing every model, no top-50 lists, no marketing. In September they add up to a split: the agent's brain came apart into a cheap fast half and an expensive slow one, control moved into a single layer of gateways and identities, and that layer became the thing attackers went for first.
On September 15 a company called TypeSafe AI came out of stealth with a $40M seed round and a model that refuses to write prose. You cannot ask Jev to summarise a document or draft an email. You hand it unstructured text and a schema you defined, and it hands back a typed answer (a Choice, a Score or a Boolean) with a calibrated probability attached, and that is all it will ever do.
I read the launch post expecting another "we built a smaller model" announcement and came away thinking it was the most useful reframing of the month, even though it cannot write a single sentence. Most of September's better reading turned out to be about the same thing from different angles: which decisions deserve an expensive brain, which ones do not, and who gets to control the switch.
The most interesting model of the month cannot write a sentence
TypeSafe calls Jev a "System One model", borrowing Kahneman's fast, intuitive System 1 for the kind of thinking it does. The target list is deliberately boring: classification, routing, scoring, extraction, safety gates. It costs $0.042 per million input tokens and nothing for output, which is a pricing decision you can only make when the output is one word long.
The founders are Diogo Almeida (ex-OpenAI, where he worked on RLHF and ChatGPT's instruction-following), Erik Gafni and Sasha Sheng, and the round was led by DCVC. TechCrunch ran it as a new kind of AI model from a ChatGPT inventor that is thrilling developers, which is roughly what happened, because within about ten days Jev was on Vercel's AI Gateway, on OpenRouter as Jev 1.13, on Cloudflare and on Databricks. Our radar picked up around fifty projects mentioning it between September 18 and 28: pydantic-ai added a TypeSafeModel, langfuse shipped decision-model evaluators, open-swe started using Jev for automated model routing, and gpt-researcher swapped embeddings for Jev as its context selector.
Now the number everyone repeated. TypeSafe claims Jev is "193.6x faster, 444.6x cheaper" than an average of two frontier LLMs, and 70-500 ms end to end. Those figures come from TypeSafe's own workflow evals and nobody has reproduced them independently, a point ts2 made in its headline so I do not have to. There is no SLA, there is one model version, and the rate limits may change. Vercel's CEO added that it runs "up to 18x faster (p95) and more accurate" than their GPT Luna-based safety reviewer, which is a partner talking about a product it now sells. Treat all of it as a direction.
What earned my trust was the other half of the launch post. TypeSafe publishes its own failure modes: counting and arithmetic, dates, and irrelevant context in the input. A vendor that tells you where its model breaks is rarer than a 444x speedup, and more useful.
The episode that made it click for me was Latent Space's Jev: System One models for Prod, not God. "Prod, not God" is the whole argument in four words. Much of what an agent does all day is not reasoning at all. Is this PR risky? Which tool comes next? Should I compact now? Is this output safe to show? We have been sending every one of those questions to the most expensive model in the building, like hiring a principal architect to sort the mail.
The shape that falls out is three tiers: a System One model decides, a cheap model executes, the frontier plans. The part I did not expect is that the decision layer also becomes auditable almost by accident. A typed, logged, millisecond call with a probability attached is something you can put in a dashboard and argue about in a post-mortem. A paragraph of chain-of-thought inside a frontier session is not.
That raises the obvious question. If the small decisions just got several hundred times cheaper, and the big models got cheaper too, did anyone's bill go down?
Vendors stopped selling tokens and started selling tasks
The frontier price sheet moved again in September, and quite a lot. OpenAI launched GPT-6 Sol and Luna on September 22. Sol is $2/$10 and the new Codex default, while Luna at $0.10/$0.50 sets a new floor for API pricing at less than half what GPT-5.6 Luna cost. The same day Anthropic shipped Claude Opus 5.5 at $4/$20, down from $5/$25, which Anthropic says runs at about Fable 5.1 level "on most work" for roughly 40% less than Opus 5. Six days later Sonnet 5.5 held at $2/$10 with, per Anthropic, up to 30% lower cost per task, and the planned September rise to $3/$15 was quietly cancelled. Top-end per-token prices fell 20-50% in a month.
Read the vendor copy closely, though, because the pitch has changed. Opus 5.5, Sonnet 5.5 and Meta's Muse Spark 1.3 were all sold on cost per task and fewer tokens, not only on price per token. That is not a marketing whim. Gemini 3.8 Flash went GA on September 2 at an introductory $0.75/$3.75, and its cost per task went up from $0.40 to $0.58 against 3.7 Flash, because it takes more reasoning steps and more tool calls to finish the same job. Same price per token, more expensive task. Anyone whose routing spreadsheet has a "price per million" column and nothing else learned that the hard way.
The tools moved the same way. GitHub gave Copilot Auto model selection three tiers, Efficiency, Balance and Intelligence, routing per prompt with a 10% discount for letting it choose. Codex's rate-limit prompt now suggests switching to Luna. Cursor's new Projects runs a coordinator agent that plans and writes no code. Everybody is building the three-tier split from the previous section, with or without a System One model in the first tier.
So the unit got cheaper and the vendors are now selling the task. The next reading answered what happens to the bill when you do both.
Cheaper tokens bought longer sessions, not smaller bills
Anthropic published the most honest usage chart of the month, and it was in a post about Opus 5.5 rather than a pricing page. Comparing Claude Code usage from March to September: Claude works 3.3x longer per prompt, makes 40%+ more model calls per prompt, and sends 2.6x more context per request. The input-to-output ratio went from 189:1 to 324:1. Interruptions fell 68%, and developers are about twice as likely to have a tool server connected or a skill in use.
That is Jevons' paradox, and the interesting part is where it shows up. It does not show up as more developers or more prompts. It shows up inside the session, where a cheaper token gets spent on a longer run, a fatter context and more calls per prompt. Nobody decided to spend more. The agent simply had room to keep going, so it did.
The Ramp AI Index for August puts the corporate side in numbers. The effective token price paid by Ramp's customers is down 41% since the March peak, from $1.15 to $0.68 per million. Median spend per employee at the top 1% of firms fell only 9.7%, from $7,976 to $7,205 a month. A 41% cheaper unit bought a 10% smaller bill, and the gap went into usage. The line I underlined, though, was that frontier models' share of tokens fell from 53% to 45%, with growth coming from mid-tier models, because companies are setting default policies that restrict frontier use. The bill moved when someone wrote a rule, not when the price dropped.
Uber supplied the case study in Running a Software Factory Efficiently at Uber Scale. More than 70% of their PRs now come from agents. Between February and August weekly agent requests grew 9.4x and weekly active users 7x, while spend has been flat since April. The techniques are not magic: cheap models for subagents, a 400K-token context cap with automatic compaction, and an MCP "code-mode" that cuts tokens by 50-90%. The Thoughtworks radar filled in the less flattering half of the story: Uber burned through a twelve-month AI budget in four months first, and the efficiency programme came after. Both things are true and the order matters. The budget blew, and only then did someone own the policy.
Put the three reads together and the conclusion is uncomfortable for anyone who treats model choice as personal taste. Cost per task bends only through routing, default policies and context caps, which makes the model-routing policy and the budget governance decisions, not something each developer settles in their own IDE. That leaves one question: where does a policy like that actually live?
Control moved out of the tools and into one layer
September's answer was the gateway, and the product announcements read like a single spec written by several companies.
Kong AI Gateway 2.0 went GA on September 1 with MCP Server Bundling: many MCP servers behind one route, with the tool list filtered per caller, so an agent cannot even see a tool it is not authorised to use. Its "Identity Principals" carry one identity across rate limiting, cost attribution and access control. ServiceNow's AI Control Tower gave every MCP server a lifecycle from intake through approval to deprecation, with a pause switch per server. Consolidation had already done its work: Palo Alto bought Portkey in May, and Stripe agreed to buy OpenRouter in July.
The coding tools grew the same controls from the inside. Claude Code 2.1.283 added availableModelsMatch: exact, which keeps a newly released model blocked until an admin lists it, a deniedModels list that wins even over the allowlist, and an opt-in x-claude-code-prompt-id header so a gateway can group every request behind one user prompt. That last one is small and very good, because it turns "cost per token" into "cost per thing a human asked for". Codex split managed configuration into requirements.toml for hard rules and managed_config.toml for defaults, pushed via MDM, and every admin needs to re-pin before gpt-5.5 retires on October 14. Model pinning has an expiry date now, so it needs an owner. GitHub gave organisations a "default policy for new features" with 28 days to choose before new Copilot features switch on by themselves.
The pattern is hard to miss. Last month the autonomy policy moved from the permission prompt into a written ruleset. This month the ruleset moved up a level, out of the individual tool and into a shared layer that decides which models exist, which tools an agent can see and whose budget it spends. It is the right move, and it has one obvious side effect: once everything goes through one layer, that layer is the most valuable thing in your stack to break.
The control plane became the target in the same month
On September 2, CISA added LiteLLM's CVE-2026-59822 to its Known Exploited Vulnerabilities catalogue. The bug is broken authentication on the MCP Streamable HTTP endpoint, so that any bearer token at all gets a fully authorised MCP session. It scores CVSS 8.8 and affects every version before 1.84.0. The fix shipped in May and disclosure came on July 22, so anyone who was hit in September was hit by a bug that had been fixed for months. It is reported as the first MCP implementation on KEV. Whether or not that framing survives, the gateway is now critical infrastructure and an attack surface at once, and it needs what critical infrastructure gets: a patch SLA, an SBOM and someone watching the CVE feeds.
Then came the read that changed how I think about agent credentials. GreyNoise's data, written up by VentureBeat under the headline AI agents breached 395 organizations using credentials your IAM policy still treats as human, describes the first mass-exploitation campaign run by AI agents. It hit 395 organisations in 48 countries through PaperCut, harvested Active Directory credentials at 280 of them, and made its first compromise in 26 seconds. Stolen tokens turned up on 5,871 machines, including Claude, Cursor and Gemini sessions, and one compromised agent dashboard burned $600k of model credits. The attackers were agents, and some of what they stole was agents.
Anthropic's September threat intelligence report closes the loop. An actor used prompt injection against an AI vendor's automated evaluation sandbox, which handed over production API keys for several providers, and the follow-on campaign reached about thirty AI companies in roughly four days. The report's advice is dull and correct: "treat AI keys and agent integrations like production credentials".
The plumbing had its own bad month too: GitSpawn used a repository's own .git/config to turn an agent's routine startup git status into code execution outside the sandbox, across eight agents. Last month the lesson was that opening a repository runs code. This month it is that the agent's credentials are worth more than the code it writes, and those credentials are not attached to anyone.
Every agent is getting a name
Identity vendors spent September answering exactly that, and Oktane was where it came together. Okta shipped Agent-to-Agent Connections to GA, which governs which agents may call which and makes the handoffs auditable, along with access certifications for agents. It also announced an Agent Gateway that sits between agent and tools as a virtual MCP server (GA planned for Q3) and a kill switch that revokes an agent's live tokens at the gateway (planned for Q4). The Blueprint Alliance with AWS, CrowdStrike and Salesforce comes with two slogans worth keeping: "every agent is a first-class identity" and "scope access to the task, no standing access". Okta's own figure is that only 34% of organisations apply the same controls to agents as to people.
CrowdStrike's Agentic Identity Provider went further on provenance: each agent registers automatically, gets a cryptographically verifiable identity and short-lived least-privilege credentials, and every action is chained back to a human or a workload. GitHub shipped enterprise-managed permissions for Copilot agent operations on September 9, so admins can block, require approval for or allow shell commands, file access and network domains, and users cannot override them. Two weeks later GitHub added "proof of presence", a re-authentication before high-impact actions such as creating a token, aimed squarely at stolen tokens and cookies.
A healthy dose of scepticism is in order, because a good share of the Oktane stage was roadmap. The kill switch is Q4 and the gateway is "planned". The parts that shipped are mostly permissions and certifications, which is the unglamorous half, and the half you can use on Monday. Still, the direction is not in doubt. An agent that holds a token is a principal and should be treated like one: named, scoped to the task, revocable, and logged. The 26-second compromise above is what happens when it is treated as a convenience feature of whichever developer happened to paste the key.
All of which is great on a slide. The last reading question of the month was whether anyone actually enforces what they wrote down.
Governance you cannot enforce is a wish
EY surveyed 202 US companies with more than $1B in revenue for their AI risk and governance report, and one pair of numbers is the month's whole argument. 98% have formal AI governance policies. 47% skipped them for urgent deployments. 91% use agentic AI, 49% have not updated their governance for it, and 26% cannot detect an unauthorised agent in their own estate. 36% have had a materially negative AI incident. The one encouraging figure: where formal reviews actually happen, 64% end up significantly changing a quarter or more of their AI systems. Reviews work when you run them, which is a low bar that half the sample still does not clear.
Marmelab's State of AI Harness Engineering 2026 found the same gap one level down, in the repository. 63% of large open-source projects now ship agent instructions, but only 4.4% of the security rules in public CLAUDE.md files are backed by a technical control. The other 95.6% are polite requests to a model. The rest of the report is a quiet demolition of "more is better": the same model scored anywhere from 68% to 88% depending on which of eight harnesses ran it, machine-generated context files did worse than none at all while costing 20%+ more, 60% of harnesses have no tests or evals, and Vercel removed 80% of an agent's tools and watched success go from 80% to 100%. Adding a second reviewer agent lowered success by 8%. The harness matters more than the model, and less of it often beats more.
Then Meta dropped AI usage from engineer performance reviews, ending about ten months of grading people on "AI-driven impact" and scrapping a token leaderboard that covered 85,000 employees. The Information put Meta's consumption at 60.2 trillion tokens in 30 days. Last month I named metering adoption by token spend as the anti-pattern of the season, and I did not expect the largest possible confirmation to arrive this quickly. Meta's replacement is budget controls and a central dashboard, planned for 2027, which is to say the same routing and default-policy machinery from the Ramp and Uber reads, arriving about a year late.
The filter I'm keeping
The September reads sort into one sentence. The brain split in two, the control moved into one layer, and that layer turned out to be both the only place the bill bends and the first place attackers go.
Three questions for October. Which decisions in your agent loop are being paid for at frontier prices when a typed yes-or-no would do, and would you be able to see them in a log if you moved them? Where does your routing and model policy live: in a gateway with an owner and a patch SLA, or in thirty developers' dotfiles? And for every rule you have written down, in a governance deck or in a CLAUDE.md, what is the technical control that makes it true when nobody is reading?
None of that needs a new model, which feels like the theme of the last three months. The capability curve keeps going up and to the right. The work that decides whether you benefit from it is plumbing, policy and names on things.
All of this went into the October edition of the VISDOM AI maturity matrix. The biggest change is at L4 of Coding Agent Usage, where routing is now three-tier (a decision model decides, a cheap model executes, the frontier plans) and re-costed per task rather than per token. The AI gateway, with its patch SLA, and the agent as a distinct identity in your IdP now run through Governance & Compliance and MCP & Tool Integration.
PS. Jev answers every question with a typed Boolean and a calibrated probability, and it publishes a list of the questions it is bad at. I have sat through design reviews that would have gone better with both.