Production feedback → CI auto-adjusts test suite
The test suite stops being a static artifact: an incident in production generates its own regression test and adds it to CI automatically.
- Production feedback loop auto-adjusts the CI test suite (adds tests for observed failures, removes redundant tests)
- CI auto-scales runner capacity based on agent load (no manual capacity planning)
- CI provides sub-minute feedback for standard changes
- CI runner utilization stays between 50-80% (auto-scaling prevents both waste and queuing)
- Test suite evolution is auditable (each auto-added/removed test has a provenance record)
- CI run duration dashboard showing sub-minute median for standard changes
- Auto-scaling configuration and runner utilization metrics
- Test suite change log showing production-feedback-driven additions and removals
- Delivery L4 (CI/CD Pipeline) - sub-2-minute CI and ephemeral sandboxes must be operational
- Infrastructure L4 (Observability & Feedback Loop) - production telemetry required for feedback-driven test suite adjustment
What It Is
"Production feedback drives CI test suite adjustment" is an L5 pattern where the CI test suite is not a static artifact maintained by engineers but a dynamic system that evolves based on what's actually going wrong in production. When a production incident occurs, the system automatically: identifies the code path that failed, generates a regression test covering that failure mode, adds the test to CI, and ensures that path is covered on every future change that touches the affected code. The test suite grows in response to real failures, not anticipated failures.
The pattern works in the opposite direction too: when production telemetry shows that a code path has never caused a problem and has not been changed in 6 months, the system can flag the tests covering it as candidates for deprioritization in the fast CI path - moving them to a weekly validation suite rather than running them on every commit. The test suite's composition is continuously optimized: more coverage where production failures occur, less coverage of stable code paths that are rarely exercised in production.
This is a natural extension of the production telemetry patterns that already exist in mature engineering organizations (distributed tracing, error tracking, anomaly detection). The new element is the feedback loop: production signals automatically trigger test suite changes rather than requiring a human to analyze an incident, decide to write a regression test, implement it, and add it to CI. At L5, this loop runs continuously and automatically, with human review for the generated tests but no human requirement to initiate the process.
Thoughtworks made this an explicit invariant in its operating model for enterprise AI agent reliability, published 2026-08-14: every production failure becomes a permanent regression test. The word doing the work is "permanent". A regression test that can be deleted by the next agent tidying up a test file is not a control, it is a comment. In the Thoughtworks model each of six layers - terminology, routing, agent intent, semantic context, execution, result - has a named owning team and an executable "truth contract", and the regression suite is where those contracts accumulate. If you are building this loop, decide up front which team owns the generated test and what it takes to remove one.
The mechanism typically involves: error tracking (Sentry, Datadog) that identifies failing code paths in production; an agent that reads the error and the relevant source code and generates a regression test; a CI integration that adds the test to the affected module's test suite; and a review step (which can be automated to skip human review for straightforward regression tests with high confidence). The loop closes when the regression test is added to CI and future changes to the affected code path run against it automatically.
The same loop now has to answer a harder question: where verification happens at all. CircleCI data across 28 million workflows, analysed by Paul Stack (AI broke the assumptions behind CI, 2026-08-26), shows throughput up 59% year on year while main-branch success fell to 70.8%, a five-year low, and recovery time rose 13%. More changes are arriving, more of them are breaking main, and fixing main takes longer. Stack's argument is that CI was built on the assumption that a human ran things locally before pushing, and agents broke that assumption. His proposal is to move verification before the pull request opens - the agent runs the relevant tests, including every production-derived regression test for the code it touched, and attaches the results as evidence - and to have CI check that evidence rather than re-run everything from scratch. In a shadow run over 35 pull requests, nothing slipped through. Thirty-five PRs is a small sample, and the idea drew public pushback (Jeremy Miller replied on 2026-08-31), so treat it as a direction to pilot rather than a settled result. But it fits this pattern exactly: the production-derived suite is the most valuable thing an agent can run before it opens a PR, and CI's job shifts from repeating that run to verifying it happened, against the right commit, with the right tests.
This is also where the capacity problem lands. The martinfowler.com write-up "An Accidental Blackboard" (2026-09-02), in which ten engineers built an airline disruption-management system in four days with many agents, reports that the agents strained the build pipelines. A CI that re-runs every agent's full suite on every push will not keep up; a CI that checks signed evidence and re-runs only what is missing or suspicious has a chance.
Why It Matters
- Test coverage grows where it matters most - tests are added for code paths that fail in production, which is the most reliable signal of where coverage is needed; coverage grows in response to real risk rather than developer intuition
- Regression prevention becomes automatic - every production incident generates a regression test; the same bug cannot ship again without failing CI; the test suite becomes an automatically maintained safety net
- CI test suite stays relevant as codebases evolve - code paths that are never exercised in production and never changed are deprioritized in CI, keeping the test suite proportional to actual risk; the suite doesn't grow unboundedly
- Eliminates "we should write a regression test" backlog items - production incidents generate regression tests immediately and automatically; there's no backlog of "we should have a test for this" items because the system creates them
- Permanence is the whole point - Thoughtworks' operating model treats every production failure as a permanent regression test with a named owning team; a suite that agents can quietly prune is a suite that will regress to the mean
- Main is breaking more often as agent volume rises - across 28M CircleCI workflows, throughput rose 59% while main-branch success fell to a five-year low of 70.8% and recovery time grew 13%; catching failures before the PR opens, with production-derived tests, is cheaper than repairing main after
- Demonstrates that CI and production are a continuous loop, not separate stages - production feedback flowing back to CI represents the highest level of CI maturity: the pipeline learns from reality rather than being a static artifact
Getting Started
- Establish production error tracking with code path attribution - Before auto-generating tests, you need production errors attributed to specific code paths. Sentry, Datadog APM, and Honeycomb all provide stack trace analysis that maps errors to source code locations. Configure your error tracking to capture full stack traces and correlate them with your current deployed version's source map or symbol table.
- Build a "production failure to test spec" translation layer - Create a process (initially human-run, eventually agent-automated) that takes a production error with its stack trace and converts it into a test specification: what inputs triggered the error, what was the expected behavior, what was the actual behavior. This can start as a structured post-incident template and evolve toward automated generation.
- Implement an agent-based regression test generator - Use Claude Code or a similar agent to take the test specification (error description, failing code path, relevant context) and generate a regression test in the appropriate test framework. The agent needs access to: the error details, the source code of the failing function, examples of existing tests in the same file as style reference, and the test framework's assertion patterns.
- Add CI integration for automatically-generated tests - Generated tests should be added to the test file for the failing module with a metadata tag (a comment or test attribute) identifying them as production-feedback-generated. This allows tracking them separately and reviewing them as a group. Configure CI to run these tests with the same priority as manually written tests.
- Implement test coverage-to-production-path correlation - Use code coverage data (Istanbul/nyc for JavaScript, coverage.py for Python, JaCoCo for Java) correlated with production code paths (from APM traces) to identify over-tested stable code and under-tested risky code. This correlation is the basis for deprioritizing stable tests and prioritizing risky ones.
- Gate the generated tests on mutation strength, not on their existence - a generated regression test can pass, add a line to the coverage report and assert nothing meaningful. Sławomir Radzymiński's August write-up found one component at 100% line coverage and 61% mutation strength. Apply mutation testing selectively to the high-risk areas the generated tests land in - auth, business rules, payments - and treat the mutation score, not the test count, as the signal the loop is working.
- Make the generated test permanent and owned - tag it with the incident it came from, assign it to the team that owns the failing layer, and require an explicit justification to delete it. Otherwise the next tidy-up pass removes exactly the tests you paid an outage for.
- Move verification before the PR opens, and make CI check the evidence - require agents to run the affected tests, including every production-derived regression test for the touched code paths, before opening a PR, and to attach the results (commit hash, test list, outcomes) to it. Have CI verify that evidence matches the head commit and covers the required tests, and re-run only what is missing, stale or on a sampled audit basis. Pilot it in shadow mode first - run the old full suite alongside and count what the evidence-check would have let through - as Stack did over 35 PRs.
- Start with human review of generated tests before automation - For the first 3 months, have a human engineer review each auto-generated regression test before it's added to CI. This validates the quality of the generation and builds the team's confidence in the process. After reviewing 50-100 tests, you'll know the false positive rate well enough to decide which categories of auto-generated tests can be added without review.
The "production failure to regression test" loop is valuable even before full automation. A lightweight version - Sentry alerts an agent that drafts a regression test and opens a PR - delivers most of the value with minimal infrastructure. Start with the semi-automated version and automate the human review step once you've validated the test quality.
Common Pitfalls
Generating tests that test implementation details rather than behavior. Auto-generated tests that directly assert on internal function structure rather than observable behavior will break every time the implementation changes, even when behavior is correct. Validate generated tests against the behavior-specification principle: do they test inputs and outputs, not implementation?
Adding auto-generated tests without fixing the underlying bug. A regression test that fails on the production-reproducing input is valuable. But if it's added to CI without also fixing the bug that caused the production failure, every CI run will fail until someone fixes it. The workflow must be: fix the bug first, then add the regression test to prevent recurrence. Not the other way around.
Deprioritizing tests based on coverage data alone. Coverage data says which code paths are executed by tests, not which code paths are important. A path that has 100% test coverage but is never exercised in production (a rarely-used admin feature, a deprecated API endpoint) may be safe to deprioritize. But a path with 100% test coverage that handles critical production traffic should not be deprioritized even if it's stable. Use production traffic data alongside coverage data for deprioritization decisions.
Not attributing auto-generated tests in version control. Generated tests should be clearly marked as auto-generated in their file header or with test metadata. This allows engineers to understand the test's origin, helps during code review ("this test was auto-generated from production error XYZ"), and enables bulk operations (find all auto-generated tests for a module, review their quality as a batch).
Creating a feedback loop that generates tests faster than the team can maintain them. An organization with many production incidents generating automatic regression tests could accumulate thousands of tests quickly. Monitor test suite growth rate and set a threshold that triggers review: if more than X tests are generated per week, review the generation quality and consider tightening the generation criteria. SlopCodeBench's July 2026 run is the cautionary number here: ~29,000 source lines of which 51% were test code, with complexity and duplication rising at every checkpoint for every model. Test-suite mass is a warning sign, not a maturity signal.
Trusting agent-supplied evidence without checking it. Evidence-based CI only works if the evidence is bound to the exact commit and the exact test set, produced by a runner the agent cannot tamper with, and spot-checked by re-running a sample. An agent that reports "all tests passed" in a PR description has produced a claim, not evidence.
Reading a green auto-generated suite as proof of coverage. Mutation testing is the check that catches this, and it is now the recommended monitor for agent-written code generally - Birgitta Böckeler's August experiment on TDD inside the agent loop ends with exactly that recommendation, having found that agents told to TDD produced worse-judged solutions at 3-8x the tokens. Her result rests on five batches of greenfield business logic judged largely by a model, so carry it as a provocation rather than a settled finding; the mutation-testing advice stands on its own either way.
How Different Roles See It
Bob's team has been improving test coverage for months, but production incidents still occur in areas that have tests. Investigation shows that most production incidents are in code paths that have tests - but those tests don't cover the specific edge case that failed. Bob realizes the problem: developers write tests for the happy path and obvious edge cases, but production failures happen in the long tail of real-world inputs and conditions that no one anticipated during development.
Bob should propose the production-feedback-to-CI loop as the infrastructure investment that addresses the root cause. The argument: "we can't anticipate all production failure modes in advance, but we can automatically capture them after they occur and prevent their recurrence." Bob should fund a 2-sprint project to implement the initial version: Sentry alert triggers an agent draft, human engineer reviews and approves, test is added to CI. After 3 months of operation, Bob should review the results: how many regression tests were generated, how many would have caught the original incident if they'd existed, and what the recurrence rate of production issues is for code paths with auto-generated regression tests vs. without. That data validates the investment and justifies automation of the human review step.
Sarah has been tracking MTTR (mean time to resolution) for production incidents and notices that ~40% of incidents are regressions - bugs that were previously fixed and re-introduced. This is exactly the pattern that auto-generated regression tests would address. She has the data to make a compelling case: if 40% of incidents are regressions, and regression tests cost 30 minutes of engineer time per incident to write manually, the current manual process is costing the team N × 30 minutes per month, where N is the number of incidents per month.
Sarah should present the regression rate data and the math to Bob: "40% of our incidents are regressions, we have X incidents per month, writing regression tests manually costs Y engineer-hours per month, automating this process would recover Y hours per month at the cost of a 2-sprint implementation." She should also propose a leading indicator: "regression recurrence rate" (how often the same type of incident recurs) as a metric that the auto-test loop should reduce. After implementation, if the regression recurrence rate drops from 40% to 10%, that's a direct measurement of the infrastructure ROI.
Victor has already built a proof of concept: a webhook from Sentry that sends production error events to a Claude Code agent via an MCP tool. The agent reads the error, pulls the failing function from GitHub, and drafts a regression test using the existing test file in that module as a style reference. Victor reviews the draft, makes minor edits, and opens a PR. The whole process takes him 5 minutes instead of the 30 minutes a manually-written regression test would take.
Victor should automate the human review step for the cases where the generated test meets a quality bar he can define: the test covers a specific, named code path; the test uses only public function interfaces; the test has clear assertions; the test passes after the bug fix and fails before it. A fifth criterion is worth adding now: the test kills the mutants that the original defect would have produced. Tests that meet all criteria can be auto-merged without human review. Tests that fail one criterion go to the review queue. Victor should implement this quality check as part of the agent's test generation workflow and track the auto-merge rate over time. If 70% of generated tests pass the quality criteria and are auto-merged, he has a system that operates largely autonomously and generates high-quality regression tests at scale. Victor is also the right person to pilot pre-PR verification: have his agents run the production-derived tests for the code they touch before opening a PR, attach the results as evidence, and let CI verify rather than repeat the run - in shadow mode first, comparing against the full suite until the miss count is known.
Further Reading
How This Guide Changed
What each edition changed in this guide, newest first.
- V1.7October 2026LATEST
The item now adds that verification moves before the PR opens and CI checks the evidence instead of re-running it, prompted by CircleCI data across 28M workflows showing main-branch success at a five-year low of 70.8% as agent throughput rose. The guide added Paul Stack's evidence-based CI argument with its small 35-PR shadow run flagged as a pilot result, a step for pre-PR verification using the production-derived suite, and a pitfall about accepting agent-supplied evidence that is not bound to the commit.
- V1.6September 2026
The matrix item added "every production failure becomes a permanent regression test", and this guide now argues the case for that word "permanent" using Thoughtworks' August operating model, where each layer has a named owning team and its regression tests are the executable form of its truth contract. The bigger change is a new quality gate: generating a test is easy, generating one that asserts something is not, so the guide now routes generated tests through mutation testing rather than counting them, and carries SlopCodeBench's 51%-test-code finding as the reminder that suite mass is a warning sign rather than progress.
- V1.5August 2026
The summary was cut back to the mechanism itself: an incident generates its own regression test, and that test enters CI without anyone deciding to write it. The body, including the reverse direction where long-stable paths get demoted out of the fast lane, was untouched.
- V1.0March 2026
An L5 pattern from the first edition, built on the observation that test suites grow against imagined failures while production keeps a list of real ones. Both directions of the loop were described from the outset: coverage added where incidents actually happen, and coverage moved off the critical path where nothing has gone wrong in months.
Where does your team actually sit on this?
This guide describes one level of one area. Run the assessment to place your team across all 16 areas, see which gates you have passed, and get a report you can take to your stakeholders.