Skip to main content

Test-Driven Development (TDD) in the AI-Driven Development Life Cycle (AI-DLC): A Review of the Evidence and a Proposed Practice

· 27 min read
Nguyễn Huỳnh Minh Tiến
Middle Fullstack Developer @ Utop.vn
Số lượt xem trang
Summary

Abstract. Test-Driven Development (TDD) is a technique that repeats three steps: write a failing test (Red), write the minimum code that makes it pass (Green), then restructure the code while keeping all tests passing (Refactor). Empirical studies from before the AI era are mixed: a case study of four industrial teams recorded a 40–90% reduction in defect density at the price of 15–35% more initial development time, while a meta-analysis of 27 studies found only a small quality improvement. In the AI-Driven Development Life Cycle (AI-DLC), where source code is increasingly generated by large language models (LLMs), tests acquire an additional function: they act as executable specifications that clarify intent and verify AI output. Studies of test-driven code generation from 2023 to 2026 report that giving tests to an LLM improves results on benchmarks, and the newest approach lets the model generate and then jointly refine both tests and code. However, the evidence comes mainly from small-scale benchmark problems, the quality of the test suite itself caps any claim of correctness, and no evaluation yet exists at the level of the whole AI-DLC.

Keywords: TDD, AI-DLC, unit testing, large language models, executable specification, refactoring.

1. Problem statement​

As developers delegate a growing share of code writing to AI assistants (see the AI-Driven Development series, in Vietnamese), the central question of testing changes. It used to be "is the human-written code correct?" It is now "who is responsible for defining what correct means, when code is generated faster than anyone can read it line by line?"

TDD is a natural candidate for this question because it places the definition of correctness (the test) before the implementation. However, the classic evidence on TDD was collected when humans wrote both tests and code. This article addresses three questions:

  1. What does the available empirical evidence say about the effectiveness of TDD?
  2. Where can TDD be placed in AI-DLC, and how far do studies of LLM code generation support that placement?
  3. What does a workable practice look like, and which risks need to be controlled?

The literature covered consists of publications whose figures were checked against their abstracts at the source; the handling is detailed in Section 7.

2. Background​

2.1. The TDD loop​

Beck (2002), in Test-Driven Development: By Example, describes TDD through two rules: write new code only when an automated test is failing, and eliminate duplication. Together they form a three-phase loop:

PhaseActivityExit condition
RedWrite one test describing the desired behaviorThe test runs and fails for the expected reason
GreenWrite the minimum code, hard-coding if necessaryAll tests pass
RefactorImprove the structure of code and tests without changing behaviorAll tests still pass

TDD should be distinguished from three neighboring concepts. Test-first prescribes only the order of writing tests first; TDD adds disciplined refactoring and a rhythm of small steps. Unit tests are the product, whereas TDD is the process that creates them and shapes the design at the same time. ATDD/BDD place tests at the level of business behavior, typically forming an outer loop around unit-level TDD.

One condition is easily overlooked: the Red phase must fail for the expected reason. A test that is red because of an exception inside the test code itself proves nothing about the behavior to be built.

2.2. AI-DLC​

AI-DLC was proposed by Raja SP (AWS) and published on 31 July 2025. The method holds that AI acts as the primary executor while humans retain decision authority wherever business context and judgment are required, under the principle "AI Powered Execution with Human Oversight". The life cycle has three phases:

  • Inception: AI turns business intent into requirements, user stories and units of work through Mob Elaboration, in which a cross-functional team validates the AI's proposals and questions.
  • Construction: AI proposes the logical architecture, domain models, code and test suites through Mob Construction, with the team clarifying technical decisions in real time.
  • Operations: AI manages infrastructure as code and deployment, drawing on context accumulated in earlier phases.

In terms of vocabulary, the Bolt replaces the sprint (cycles measured in hours or days) and the Unit of Work replaces the epic. In this description, tests appear as an artifact continuously generated by AI during Construction, and the introduction mentions AI applying an organization's coding standards, design patterns and security requirements when generating test suites. Attaching TDD to AI-DLC in this article is therefore the author's proposal, not something prescribed by the original material.

3. Empirical evidence on TDD​

The studies below differ in subjects (students or professionals), design (experiment or case study) and measures, so they should be read as complementary slices rather than as a single figure.

StudyTypeMain result
Nagappan, Maximilien, Bhat, Williams (2008)Case study, 4 teams (3 Microsoft, 1 IBM)Pre-release defect density down 40–90% versus comparable projects; initial development time up 15–35%
Rafique & Mišić (2013)Meta-analysis, 27 studiesSmall improvement in external quality, little to no effect on productivity; industrial studies show both a larger quality gain and a larger productivity drop than academic ones
Fucci et al. (2017)Experiment, 39 professional developersOrder of writing tests and code had no important influence; quality and productivity were tied to the granularity and uniformity of steps
Causevic, Sundmark, Punnekkat (2011)Systematic reviewSeven factors limiting adoption, including increased development time, lack of TDD experience, lack of upfront design, domain- and tool-specific issues, and legacy code
George & Williams (2004)Experiment, 24 professional pair programmersThe TDD group passed 18% more functional black-box tests but took 16% more time; the control group often did not write the required automated tests after finishing the code
Romano et al. (2017)Multi-method (qualitative) study, novice and professional developersExamines the values, beliefs and assumptions of people applying TDD, i.e. how TDD is actually practiced rather than only what it yields

Three observations follow from the table:

  1. Quality benefits come with a cost. The 40–90% defect reduction cannot be separated from the 15–35% increase in initial time. The earlier experiment by George & Williams (2004) shows the same trade-off structure: 18% higher functional quality for 16% more time. The net value depends on the cost of defects in the specific environment.
  2. Context amplifies both benefit and cost. Rafique & Mišić find that industrial studies show both a larger quality improvement and a larger productivity drop than academic ones. The productivity drop is also larger when the TDD group invests significantly more test effort than the control group.
  3. Mechanism matters more than order. Fucci et al. found no important effect of writing tests before or after code; good results were tied to small, steady steps. The authors suggest the benefit comes from "fine-grained, steady steps that improve focus and flow". This point is especially relevant for Section 4.

Karac & Turhan (2018) also examine how far TDD has met the expectations placed on it, stressing that TDD is more than writing tests first. This article does not cite a quantitative conclusion from that publication. A controlled experiment published in Information and Software Technology in 2011, comparing TDD with test-last development in small increments, is also often invoked in the productivity debate; because its results could not be verified, it is not included in the table above (see Section 7).

4. TDD in the context of AI-DLC​

4.1. Tests as executable specifications​

When code is generated by an LLM, a natural-language requirement typically contains ambiguities that the model will resolve in its own way. A test removes that ambiguity with a statement that can be run. Several recent studies examine this approach:

StudyDesignMain result
Fakhoury et al. (2024), TiCoder, IEEE TSEInteractive workflow using tests to clarify intent; user study with 15 programmers; 4 LLMs, 2 Python datasetsAverage absolute improvement of 45.97% in pass@1 within 5 interactions; significantly lower task-induced cognitive load
Liang et al. (2026), ClassEval-TDDIterative TDD-style framework for class-level generation, 8 LLMsCorrectness up 12–26 percentage points over direct generation; up to 71% fully correct classes
Piya & Sullivan (2023), LLM4TDDChatGPT on LeetCode problems, tests presented incrementallyExamines the effect of test, prompt and problem attributes; no specific figures cited here
Mathews & Nagappan (2024)Tests supplied alongside the problem statement to GPT-4 and Llama 3; MBPP and HumanEval benchmarksAdding tests to the prompt led to more successful problem solving; the authors regard TDD as a promising way to ensure LLM-generated code captures the requirements
Cui (2025), Tests as PromptWebApp1K benchmark, 1,000 challenges across 20 domains, 19 frontier LLMs; tests serve as both prompt and verificationInstruction following and in-context learning matter more than raw coding ability; identifies instruction loss in long prompts
Yu et al. (2026), TDD-Agent (arXiv preprint)The model generates tests first, then refines code and tests jointly using execution feedback; evaluated on LiveCodeBench and RepoEvalImproves over reasoning-, retrieval- and agent-based baselines; refined tests show higher pass rates, coverage and mutation scores
Cassieri et al., ACM TOSEMA laboratory study, a controlled experiment with graduate students and three industry qualitative studies on generative AI for TDDDetailed results not verified; recorded as evidence that the topic is being studied with multiple methods

The results point the same way: supplying tests to a model, and letting it iterate on feedback, substantially improves accuracy over generating directly from a description. Three limitations apply. First, the experiments use benchmark problems (single functions or classes) and do not reflect multi-component systems. Second, "accuracy" is measured by the very test suites involved, so it depends on their quality (see 4.2). Third, the TiCoder user sample is small (15 people). In addition, TDD-Agent is an arXiv preprint that has not been peer reviewed. The result of Cui (2025) carries a practical implication, at the level of inference: if models lose instructions in long prompts, introducing tests in small increments suits better than supplying one large test suite at once, which also coincides with the small-steps recommendation in Section 3.

4.2. Test quality bounds correctness​

Liu et al. (2023), with EvalPlus, expanded the HumanEval test suite 80-fold and re-evaluated 26 LLMs. Pass rates fell by up to 19.3–28.9%, and the ranking among models changed: two open-source models outperformed ChatGPT on the expanded suite but not on the original. The authors conclude that insufficient tests can lead to mis-ranking.

This result does not concern TDD as such, but it has a direct consequence when AI-DLC lets AI generate both code and tests. If a model writes tests based on its own reading of the requirement and then writes code to pass them, "all tests pass" demonstrates only the model's internal consistency, not fitness for the business intent. This is the author's inference; no study that measures this phenomenon directly within AI-DLC was found.

This argument should be balanced against TDD-Agent (Yu et al., 2026): the system lets the model itself generate tests and then refine both tests and code from execution feedback, and reports improved pass rates, coverage and mutation scores for the tests. That indicates AI-generated tests are not worthless, and that letting AI refine tests is a direction worth studying. However, these evaluations rely on benchmarks, do not measure fitness to a specific organization's business intent, and do not evaluate the role of human review. This article therefore keeps a cautious stance in 4.3: let AI help draft tests, but have humans confirm them. Broader risks of delegating code to AI are analyzed in part 3 of the AI-Driven Development series, and ways to feed project knowledge to AI agents are discussed in the post on agent skills (both in Vietnamese).

4.3. Proposed division of roles​

From these two observations, the article proposes a division of roles mapped onto AI-DLC phases:

AI-DLC phaseCorresponding TDD activityProposed role
Inception (Mob Elaboration)Turn acceptance criteria into concrete examples and behavior-level testsAI drafts, humans validate, since this is where business intent is fixed
Construction (Mob Construction), RedWrite the unit test describing the next behaviorHumans write or approve each test before AI writes code
Construction, GreenMinimal implementation to pass the testAI performs it; the result is verified by the tests, not by skimming
Construction, RefactorRestructure while all tests passAI proposes, humans review; tests are the safety net

The rationale for placing authority in the Red phase is that it is where correctness is defined, which matches AI-DLC's principle that decision authority rests with humans. The rationale for assigning Green to AI is that it has the clearest automated feedback, exactly the kind of task that the studies in 4.1 show models handle better when guided by tests.

The rhythm is also compatible. Fucci et al. tie good outcomes to small, steady steps, and AI-DLC uses the Bolt (hours or days) in place of the sprint. Each Red-Green-Refactor cycle can be seen as a finer unit nested within a Bolt.

4.4. Research gap​

Two bodies of literature currently exist side by side. The first is TDD for LLM code generation (4.1). The second looks at AI across the whole software development life cycle: Guimaraes & Nascimento (2025), in the FSE Companion, examine AI's impact across the phases of the life cycle and outline a research agenda; Bhati (2026) describes the shift from code completion to agentic systems working at the level of a repository, a feature or an algorithm, proposes a six-layer reference architecture and names five open problems (evaluation, governance, technical debt, skill redistribution and the economics of attention). Bhati also compiles time savings of 13.6–55.8% across controlled studies of AI-assisted coding; this is a figure for AI in general, not for TDD specifically.

Within the scope of the author's search, no study was found that evaluates an AI-driven life cycle designed around TDD at industrial project scale. Joining the two bodies of work to conclude something about the combined process is an inference that requires verification.

Proposed study design (the author's proposal, not carried out). A controlled experiment could compare two groups: group A uses AI in the conventional way (requirement, AI writes code, then tests), group B uses AI with TDD (requirement, tests, AI writes code, tests, refactor). This design does not require assuming that AI is better than developers. Measurable indicators:

DimensionMetric
CorrectnessTest pass rate
Test robustnessMutation score
QualityDefect density
MaintainabilityComplexity, coupling
ProductivityTime per task
AI efficiencyTokens per task, agent task success rate
Human effortNumber of interventions
ReliabilityRegression rate

As for the process framework, one can imagine a test-centered variant of AI-DLC: human requirement, AI-drafted requirement analysis and acceptance criteria, tests (Red), implementation (Green), refactoring, code review, mutation, integration and end-to-end testing, human approval, operation, and then feedback from the real environment returning as new tests. This is a design sketch, not a validated framework.

5. Practical illustration​

The assumed problem, small enough to follow: orders of 1,000,000 or more receive a 10% discount; VIP customers receive an additional 5%, stacked; a negative subtotal is invalid input. Following the division of roles in 4.3, the developer writes the tests and an AI assistant proposes the implementation.

5.1. Cycle 1: Red, then Green by faking it​

The first test picks the simplest case:

public class OrderPricingTests
{
[Fact]
public void Total_UnderThreshold_NoDiscount()
{
var pricing = new OrderPricing();

Assert.Equal(500_000m, pricing.Total(500_000m, isVip: false));
}
}

The class OrderPricing does not yet exist, so the code does not compile; this is itself a form of Red. An empty class whose method throws NotImplementedException should be created so the test runs and fails for the expected reason, followed by just enough code to pass:

public class OrderPricing
{
public decimal Total(decimal subtotal, bool isVip) => subtotal;
}

This is Beck's Fake It technique: return exactly the value the test needs, without generalizing yet. With an AI assistant this phase matters even more: the developer should observe the test actually failing before asking the AI to write code.

5.2. Cycle 2: triangulation​

Add a test that forces the code to actually compute (triangulation):

[Fact]
public void Total_AtThreshold_Gets10PercentOff()
{
var pricing = new OrderPricing();

Assert.Equal(900_000m, pricing.Total(1_000_000m, isVip: false));
}
public decimal Total(decimal subtotal, bool isVip)
=> subtotal >= 1_000_000m ? subtotal * 0.90m : subtotal;

The boundary value (exactly at the threshold) is where > versus >= errors tend to appear. An AI assistant may pick the wrong boundary if the requirement is described only in words; a boundary test rules this out.

5.3. Cycle 3: VIP, invalid input and refactoring​

[Fact]
public void Total_VipUnderThreshold_Gets5PercentOff()
{
var pricing = new OrderPricing();

Assert.Equal(475_000m, pricing.Total(500_000m, isVip: true));
}

[Fact]
public void Total_VipAtThreshold_StacksTo15Percent()
{
var pricing = new OrderPricing();

Assert.Equal(850_000m, pricing.Total(1_000_000m, isVip: true));
}

[Fact]
public void Total_NegativeSubtotal_Throws()
{
var pricing = new OrderPricing();

Assert.Throws<ArgumentOutOfRangeException>(
() => pricing.Total(-1m, isVip: false));
}

Once all tests pass, the Refactor phase removes the nested conditions and names the business constants:

public class OrderPricing
{
const decimal BulkThreshold = 1_000_000m;
const decimal BulkRate = 0.10m;
const decimal VipRate = 0.05m;

public decimal Total(decimal subtotal, bool isVip)
{
if (subtotal < 0)
throw new ArgumentOutOfRangeException(nameof(subtotal));

var rate = 0m;
if (subtotal >= BulkThreshold) rate += BulkRate;
if (isVip) rate += VipRate;

return subtotal * (1 - rate);
}
}

The stacking rule now lives in one place. If a model or a human later switches to multiplicative discounts, the StacksTo15Percent test fails immediately. This is the "safety net" function of a test suite, and it matters more when code is changed in bulk by AI.

Verification note. The expected values (475,000; 850,000; 900,000) were computed by hand from the rules. The C# code in this article was not compiled or run during writing. Run dotnet test before use.

5.4. Common deviations​

  • Writing many tests at once before asking for code. The short feedback loop, a factor associated with good results (Fucci et al., 2017), is lost. With AI the risk grows because a model easily produces a large block of code that is hard to review.
  • Skipping Refactor. Doing only Red-Green yields working but poorly structured code, and the tests gradually become a maintenance burden.
  • Tests tied to implementation rather than observable behavior. Changing the implementation breaks the tests even though behavior is unchanged.
  • Letting AI write both tests and code without reviewing the tests (see 4.2).

6. Limitations and scope of application​

Schools. When code has dependencies, TDD splits into two approaches. The Chicago (classicist) school checks resulting state, uses real objects or simple fakes, and works from the inside out. The London (mockist) school checks interactions between objects using mocks, and works from the outside in (Fowler, 2007; Freeman & Pryce, 2009). For pure business logic such as the example in Section 5, the classicist approach is usually less coupled to internal structure. Whether a dependency can be swapped for a test double depends on a design following dependency inversion, covered in the Clean Architecture lesson; for repositories specifically, the Unit of Work and Repository lesson explains why integration tests often remove the need to mock (both in Vietnamese).

Poorly suited situations.

  • Exploratory code (spikes). When it is unclear what to build, writing tests first only fixes an assumption that may be wrong. Explore, extract the insight, discard the code and rewrite it with TDD.
  • UI and visual effects. Results need to be judged by human eyes; snapshot-style tests give low benefit relative to maintenance cost.
  • Infrastructure integration (databases, queues, networks). Mocking infrastructure easily creates false comfort; deliberate integration tests are more reliable (see the API Testing lesson on WebApplicationFactory and Testcontainers, in Vietnamese).
  • Legacy code without seams. Dependencies must be broken first, following Feathers (2004), before TDD can be applied.

Threats to validity of this review. The figures in Section 3 largely predate AI assistants; those in 4.1 and 4.2 come from small-scale problems and benchmarks. Section 4.3 is a proposal that has not been empirically validated.

7. Handling of sources​

The figures of Nagappan et al., Rafique & Mišić, Fucci et al., Causevic et al., George & Williams, Fakhoury et al., Liang et al., Liu et al., Mathews & Nagappan, Cui, Yu et al. and Bhati were checked against the published abstracts at the source (publisher, arXiv) or through search results quoting the abstract. Some publisher pages block automated access, so for George & Williams and Romano et al. only secondary sources could be checked. For Piya & Sullivan only the abstract could be checked and it contains no quantitative results; for Karac & Turhan, Guimaraes & Nascimento and Cassieri et al. only publication details and a description of the design could be checked, without detailed results. The 2011 experiment published in Information and Software Technology (Pančur & Ciglarič, 53(6), 557–573) was left out because the secondary summaries found describe its results inconsistently. TDD-Agent is an arXiv preprint that has not been peer reviewed. When citing for academic purposes, open the original paper to verify sample size, measures, confidence intervals and experimental conditions. Statements about AI-DLC rest on the method author's own introduction, not on independent evaluation.

Adoption checklist​

Before confirming an AI-assisted TDD process

  • •Each new test is run and seen failing for the expected reason before AI writes code
  • •Humans write or approve every test; tests drafted by AI are reviewed against the requirements
  • •Each Red-Green-Refactor cycle is small enough to review the AI's output
  • •The Refactor phase actually happens, not skipped under pressure
  • •Tests describe observable behavior, not implementation details
  • •Boundary values (threshold, empty, negative, null) are written as tests early
  • Scope is appropriate: pure business logic, not forced onto UI or exploratory code

Frequently asked questions​

Frequently asked questions about TDD and AI-DLC

Does TDD really reduce defects?

There are signs that it does, but the magnitude depends on context. The case study of four industrial teams by Nagappan et al. (2008) recorded a 40–90% drop in defect density versus comparable projects not using TDD. The meta-analysis of 27 studies by Rafique and Mišić (2013) found only a small improvement in external quality, although the improvement in industrial studies was larger than in academic ones.

How much does TDD slow development down?

Nagappan et al. reported initial development time rising by about 15–35%. This cost is expected to be recovered in bug fixing and maintenance, so it is usually reasonable for long-lived products and less so for short-lived experiments.

What is AI-DLC, and does it require TDD?

AI-DLC is a software development method proposed by Raja SP (AWS) in 2025, with three phases (Inception, Construction and Operations) in which AI executes while humans retain decision authority. The introduction to the method describes tests as an artifact generated by AI during Construction and does not prescribe TDD. Combining TDD with AI-DLC in this article is the author's proposal.

Why not let AI write both the tests and the code?

If the same model misreads a requirement, the tests and the code will be wrong consistently and still pass. The EvalPlus study (Liu et al., 2023) shows that an insufficient test suite can inflate pass rates by up to 19.3–28.9% and mis-rank models. Systems such as TDD-Agent (Yu et al., 2026, preprint) show that letting AI generate and refine tests can improve benchmark results, but they do not measure fitness to business intent. AI should therefore help draft tests while humans review them, because that is where correctness is defined.

Does giving tests to an LLM really improve code generation?

Benchmark studies point the same way. Mathews and Nagappan (2024) found that adding tests to the problem statement helped GPT-4 and Llama 3 solve more problems on MBPP and HumanEval. Fakhoury et al. (2024) recorded an average absolute pass@1 gain of 45.97% with a test-based interactive workflow. The limitation is that the experiments use small benchmark problems and are measured by the same test suites, so they do not transfer directly to industrial systems.

Is writing tests after the code worse than TDD?

Not necessarily. The experiment by Fucci et al. (2017) with 39 professional developers concluded that the order of writing tests and code had no important influence; quality and productivity were tied to working in small, steady steps. However, tests written afterward tend to mirror the existing code and skip the hard cases.

Conclusion​

TDD is not a universal solution, and the empirical evidence is not strong enough to treat it as a mandatory principle. What the studies support is broader than the "TDD" label: work in small steps, get automated feedback after each step, and refactor regularly.

In the context of AI-DLC, the value of TDD shifts from helping developers write correct code to allowing humans to define correctness as an executable specification and delegate the implementation to AI. Three main conclusions:

  1. Costs and benefits are both real. Fewer defects in exchange for up-front time; weigh them against the product's lifetime.
  2. Small steps have the most supporting evidence. With AI-generated code, small steps also keep the volume to be reviewed within what humans can control.
  3. Tests are where human authority is retained. The quality of the test suite caps every claim about the correctness of AI-generated code; the Red phase should therefore be owned by humans.

The next research step is to measure this combined process at industrial project scale, where no direct evidence currently exists.

References​


Last updated: October 2026

Vietnamese-language pages on this site: