Development
How developers work with AI day-to-day. From sidebar chat to fleet agents.
Maturity →
You don't have to figure this out alone.
Every level in this matrix has a path. Read the playbooks the teams that have climbed it wrote. Run the assessment with our consultants. Start where you are.
Book an AI Maturity Assessment session with your team.
We walk you through all four perspectives, score where you actually are, and leave you with a 90-day plan to climb in the dimensions that matter most.
The October 2026 zeitgeist is System One.
For two years the coding agent had one brain, and every decision went through it: which file to read, which tool to call, whether the context was full, whether the diff looked risky. In September that brain split. TypeSafe AI's Jev is a model that does not write prose at all - you hand it text and a schema, and it returns a typed choice, score or boolean with a calibrated probability, at $0.042 per million input tokens. Within ten days it sat behind four gateways, pydantic-ai had a model class for it, open-swe used it for automated model routing, and gpt-researcher swapped embeddings for it as a context selector. The shape that falls out is three-tier routing: a System One model decides, a cheap model executes, the frontier plans. GitHub made the same point from the product side with Copilot Auto's Efficiency / Balance / Intelligence tiers, routing per prompt. Treat the headline numbers with care, though: "193.6x faster, 444.6x cheaper" is TypeSafe's own benchmark, the company lists its own failure modes (counting, dates, irrelevant context), and there is no SLA. Put a decision model on your own eval set before you let it gate anything.
The second story is context, and it cuts against the instinct to write more of it. Uber's software factory post caps context at 400K tokens with auto-compaction and routes subagents to cheap models, which is how it held spend flat while agent requests grew 9.4x. Marmelab's State of AI Harness Engineering supplies the uncomfortable half: machine-generated context files did worse than having none, at 20%+ more cost, while human-written ones helped by about 4%. And only 4.4% of the security rules written in public CLAUDE.md files have a technical control behind them. A line in an instruction file that says "never touch production" is a request to the model. If it matters, it belongs in a deny rule, a sandbox or a credential the agent does not hold.
Which brings the autonomy question up a level. Last month the maturity signal was a written, version-controlled deny/ask ruleset in the repository. In September GitHub made enterprise-managed permissions for Copilot agent operations generally available: admins decide which shell commands, file operations and network domains an agent may use, and users cannot override them. Claude Code shipped exact-version model pinning in the same month. For an individual developer this feels like a loss of control. For the organisation it is the first time the rules are actually enforced, and that is the version of autonomy that survives an audit.