Test-Driven Development (TDD) in the AI-Driven Development Life Cycle (AI-DLC): A Review of the Evidence and a Proposed Practice
Abstract. Test-Driven Development (TDD) is a technique that repeats three steps: write a failing test (Red), write the minimum code that makes it pass (Green), then restructure the code while keeping all tests passing (Refactor). Empirical studies from before the AI era are mixed: a case study of four industrial teams recorded a 40–90% reduction in defect density at the price of 15–35% more initial development time, while a meta-analysis of 27 studies found only a small quality improvement. In the AI-Driven Development Life Cycle (AI-DLC), where source code is increasingly generated by large language models (LLMs), tests acquire an additional function: they act as executable specifications that clarify intent and verify AI output. Studies of test-driven code generation from 2023 to 2026 report that giving tests to an LLM improves results on benchmarks, and the newest approach lets the model generate and then jointly refine both tests and code. However, the evidence comes mainly from small-scale benchmark problems, the quality of the test suite itself caps any claim of correctness, and no evaluation yet exists at the level of the whole AI-DLC.
Keywords: TDD, AI-DLC, unit testing, large language models, executable specification, refactoring.
1. Problem statement
As developers delegate a growing share of code writing to AI assistants (see the AI-Driven Development series, in Vietnamese), the central question of testing changes. It used to be "is the human-written code correct?" It is now "who is responsible for defining what correct means, when code is generated faster than anyone can read it line by line?"
TDD is a natural candidate for this question because it places the definition of correctness (the test) before the implementation. However, the classic evidence on TDD was collected when humans wrote both tests and code. This article addresses three questions:
- What does the available empirical evidence say about the effectiveness of TDD?
- Where can TDD be placed in AI-DLC, and how far do studies of LLM code generation support that placement?
- What does a workable practice look like, and which risks need to be controlled?
The literature covered consists of publications whose figures were checked against their abstracts at the source; the handling is detailed in Section 7.
2. Background
2.1. The TDD loop
Beck (2002), in Test-Driven Development: By Example, describes TDD through two rules: write new code only when an automated test is failing, and eliminate duplication. Together they form a three-phase loop:
| Phase | Activity | Exit condition |
|---|---|---|
| Red | Write one test describing the desired behavior | The test runs and fails for the expected reason |
| Green | Write the minimum code, hard-coding if necessary | All tests pass |
| Refactor | Improve the structure of code and tests without changing behavior | All tests still pass |
TDD should be distinguished from three neighboring concepts. Test-first prescribes only the order of writing tests first; TDD adds disciplined refactoring and a rhythm of small steps. Unit tests are the product, whereas TDD is the process that creates them and shapes the design at the same time. ATDD/BDD place tests at the level of business behavior, typically forming an outer loop around unit-level TDD.
One condition is easily overlooked: the Red phase must fail for the expected reason. A test that is red because of an exception inside the test code itself proves nothing about the behavior to be built.
2.2. AI-DLC
AI-DLC was proposed by Raja SP (AWS) and published on 31 July 2025. The method holds that AI acts as the primary executor while humans retain decision authority wherever business context and judgment are required, under the principle "AI Powered Execution with Human Oversight". The life cycle has three phases:
- Inception: AI turns business intent into requirements, user stories and units of work through Mob Elaboration, in which a cross-functional team validates the AI's proposals and questions.
- Construction: AI proposes the logical architecture, domain models, code and test suites through Mob Construction, with the team clarifying technical decisions in real time.
- Operations: AI manages infrastructure as code and deployment, drawing on context accumulated in earlier phases.
In terms of vocabulary, the Bolt replaces the sprint (cycles measured in hours or days) and the Unit of Work replaces the epic. In this description, tests appear as an artifact continuously generated by AI during Construction, and the introduction mentions AI applying an organization's coding standards, design patterns and security requirements when generating test suites. Attaching TDD to AI-DLC in this article is therefore the author's proposal, not something prescribed by the original material.
3. Empirical evidence on TDD
The studies below differ in subjects (students or professionals), design (experiment or case study) and measures, so they should be read as complementary slices rather than as a single figure.
| Study | Type | Main result |
|---|---|---|
| Nagappan, Maximilien, Bhat, Williams (2008) | Case study, 4 teams (3 Microsoft, 1 IBM) | Pre-release defect density down 40–90% versus comparable projects; initial development time up 15–35% |
| Rafique & Mišić (2013) | Meta-analysis, 27 studies | Small improvement in external quality, little to no effect on productivity; industrial studies show both a larger quality gain and a larger productivity drop than academic ones |
| Fucci et al. (2017) | Experiment, 39 professional developers | Order of writing tests and code had no important influence; quality and productivity were tied to the granularity and uniformity of steps |
| Causevic, Sundmark, Punnekkat (2011) | Systematic review | Seven factors limiting adoption, including increased development time, lack of TDD experience, lack of upfront design, domain- and tool-specific issues, and legacy code |
| George & Williams (2004) | Experiment, 24 professional pair programmers | The TDD group passed 18% more functional black-box tests but took 16% more time; the control group often did not write the required automated tests after finishing the code |
| Romano et al. (2017) | Multi-method (qualitative) study, novice and professional developers | Examines the values, beliefs and assumptions of people applying TDD, i.e. how TDD is actually practiced rather than only what it yields |
Three observations follow from the table:
- Quality benefits come with a cost. The 40–90% defect reduction cannot be separated from the 15–35% increase in initial time. The earlier experiment by George & Williams (2004) shows the same trade-off structure: 18% higher functional quality for 16% more time. The net value depends on the cost of defects in the specific environment.
- Context amplifies both benefit and cost. Rafique & Mišić find that industrial studies show both a larger quality improvement and a larger productivity drop than academic ones. The productivity drop is also larger when the TDD group invests significantly more test effort than the control group.
- Mechanism matters more than order. Fucci et al. found no important effect of writing tests before or after code; good results were tied to small, steady steps. The authors suggest the benefit comes from "fine-grained, steady steps that improve focus and flow". This point is especially relevant for Section 4.
Karac & Turhan (2018) also examine how far TDD has met the expectations placed on it, stressing that TDD is more than writing tests first. This article does not cite a quantitative conclusion from that publication. A controlled experiment published in Information and Software Technology in 2011, comparing TDD with test-last development in small increments, is also often invoked in the productivity debate; because its results could not be verified, it is not included in the table above (see Section 7).
4. TDD in the context of AI-DLC
4.1. Tests as executable specifications
When code is generated by an LLM, a natural-language requirement typically contains ambiguities that the model will resolve in its own way. A test removes that ambiguity with a statement that can be run. Several recent studies examine this approach:
| Study | Design | Main result |
|---|---|---|
| Fakhoury et al. (2024), TiCoder, IEEE TSE | Interactive workflow using tests to clarify intent; user study with 15 programmers; 4 LLMs, 2 Python datasets | Average absolute improvement of 45.97% in pass@1 within 5 interactions; significantly lower task-induced cognitive load |
| Liang et al. (2026), ClassEval-TDD | Iterative TDD-style framework for class-level generation, 8 LLMs | Correctness up 12–26 percentage points over direct generation; up to 71% fully correct classes |
| Piya & Sullivan (2023), LLM4TDD | ChatGPT on LeetCode problems, tests presented incrementally | Examines the effect of test, prompt and problem attributes; no specific figures cited here |
| Mathews & Nagappan (2024) | Tests supplied alongside the problem statement to GPT-4 and Llama 3; MBPP and HumanEval benchmarks | Adding tests to the prompt led to more successful problem solving; the authors regard TDD as a promising way to ensure LLM-generated code captures the requirements |
| Cui (2025), Tests as Prompt | WebApp1K benchmark, 1,000 challenges across 20 domains, 19 frontier LLMs; tests serve as both prompt and verification | Instruction following and in-context learning matter more than raw coding ability; identifies instruction loss in long prompts |
| Yu et al. (2026), TDD-Agent (arXiv preprint) | The model generates tests first, then refines code and tests jointly using execution feedback; evaluated on LiveCodeBench and RepoEval | Improves over reasoning-, retrieval- and agent-based baselines; refined tests show higher pass rates, coverage and mutation scores |
| Cassieri et al., ACM TOSEM | A laboratory study, a controlled experiment with graduate students and three industry qualitative studies on generative AI for TDD | Detailed results not verified; recorded as evidence that the topic is being studied with multiple methods |
The results point the same way: supplying tests to a model, and letting it iterate on feedback, substantially improves accuracy over generating directly from a description. Three limitations apply. First, the experiments use benchmark problems (single functions or classes) and do not reflect multi-component systems. Second, "accuracy" is measured by the very test suites involved, so it depends on their quality (see 4.2). Third, the TiCoder user sample is small (15 people). In addition, TDD-Agent is an arXiv preprint that has not been peer reviewed. The result of Cui (2025) carries a practical implication, at the level of inference: if models lose instructions in long prompts, introducing tests in small increments suits better than supplying one large test suite at once, which also coincides with the small-steps recommendation in Section 3.
4.2. Test quality bounds correctness
Liu et al. (2023), with EvalPlus, expanded the HumanEval test suite 80-fold and re-evaluated 26 LLMs. Pass rates fell by up to 19.3–28.9%, and the ranking among models changed: two open-source models outperformed ChatGPT on the expanded suite but not on the original. The authors conclude that insufficient tests can lead to mis-ranking.
This result does not concern TDD as such, but it has a direct consequence when AI-DLC lets AI generate both code and tests. If a model writes tests based on its own reading of the requirement and then writes code to pass them, "all tests pass" demonstrates only the model's internal consistency, not fitness for the business intent. This is the author's inference; no study that measures this phenomenon directly within AI-DLC was found.
This argument should be balanced against TDD-Agent (Yu et al., 2026): the system lets the model itself generate tests and then refine both tests and code from execution feedback, and reports improved pass rates, coverage and mutation scores for the tests. That indicates AI-generated tests are not worthless, and that letting AI refine tests is a direction worth studying. However, these evaluations rely on benchmarks, do not measure fitness to a specific organization's business intent, and do not evaluate the role of human review. This article therefore keeps a cautious stance in 4.3: let AI help draft tests, but have humans confirm them. Broader risks of delegating code to AI are analyzed in part 3 of the AI-Driven Development series, and ways to feed project knowledge to AI agents are discussed in the post on agent skills (both in Vietnamese).
4.3. Proposed division of roles
From these two observations, the article proposes a division of roles mapped onto AI-DLC phases:
| AI-DLC phase | Corresponding TDD activity | Proposed role |
|---|---|---|
| Inception (Mob Elaboration) | Turn acceptance criteria into concrete examples and behavior-level tests | AI drafts, humans validate, since this is where business intent is fixed |
| Construction (Mob Construction), Red | Write the unit test describing the next behavior | Humans write or approve each test before AI writes code |
| Construction, Green | Minimal implementation to pass the test | AI performs it; the result is verified by the tests, not by skimming |
| Construction, Refactor | Restructure while all tests pass | AI proposes, humans review; tests are the safety net |
The rationale for placing authority in the Red phase is that it is where correctness is defined, which matches AI-DLC's principle that decision authority rests with humans. The rationale for assigning Green to AI is that it has the clearest automated feedback, exactly the kind of task that the studies in 4.1 show models handle better when guided by tests.
The rhythm is also compatible. Fucci et al. tie good outcomes to small, steady steps, and AI-DLC uses the Bolt (hours or days) in place of the sprint. Each Red-Green-Refactor cycle can be seen as a finer unit nested within a Bolt.
4.4. Research gap
Two bodies of literature currently exist side by side. The first is TDD for LLM code generation (4.1). The second looks at AI across the whole software development life cycle: Guimaraes & Nascimento (2025), in the FSE Companion, examine AI's impact across the phases of the life cycle and outline a research agenda; Bhati (2026) describes the shift from code completion to agentic systems working at the level of a repository, a feature or an algorithm, proposes a six-layer reference architecture and names five open problems (evaluation, governance, technical debt, skill redistribution and the economics of attention). Bhati also compiles time savings of 13.6–55.8% across controlled studies of AI-assisted coding; this is a figure for AI in general, not for TDD specifically.
Within the scope of the author's search, no study was found that evaluates an AI-driven life cycle designed around TDD at industrial project scale. Joining the two bodies of work to conclude something about the combined process is an inference that requires verification.
Proposed study design (the author's proposal, not carried out). A controlled experiment could compare two groups: group A uses AI in the conventional way (requirement, AI writes code, then tests), group B uses AI with TDD (requirement, tests, AI writes code, tests, refactor). This design does not require assuming that AI is better than developers. Measurable indicators:
| Dimension | Metric |
|---|---|
| Correctness | Test pass rate |
| Test robustness | Mutation score |
| Quality | Defect density |
| Maintainability | Complexity, coupling |
| Productivity | Time per task |
| AI efficiency | Tokens per task, agent task success rate |
| Human effort | Number of interventions |
| Reliability | Regression rate |
As for the process framework, one can imagine a test-centered variant of AI-DLC: human requirement, AI-drafted requirement analysis and acceptance criteria, tests (Red), implementation (Green), refactoring, code review, mutation, integration and end-to-end testing, human approval, operation, and then feedback from the real environment returning as new tests. This is a design sketch, not a validated framework.
5. Practical illustration
The assumed problem, small enough to follow: orders of 1,000,000 or more receive a 10% discount; VIP customers receive an additional 5%, stacked; a negative subtotal is invalid input. Following the division of roles in 4.3, the developer writes the tests and an AI assistant proposes the implementation.
5.1. Cycle 1: Red, then Green by faking it
The first test picks the simplest case:
public class OrderPricingTests
{
[Fact]
public void Total_UnderThreshold_NoDiscount()
{
var pricing = new OrderPricing();
Assert.Equal(500_000m, pricing.Total(500_000m, isVip: false));
}
}
The class OrderPricing does not yet exist, so the code does not compile; this is itself a form of Red. An empty class whose method throws NotImplementedException should be created so the test runs and fails for the expected reason, followed by just enough code to pass:
public class OrderPricing
{
public decimal Total(decimal subtotal, bool isVip) => subtotal;
}
This is Beck's Fake It technique: return exactly the value the test needs, without generalizing yet. With an AI assistant this phase matters even more: the developer should observe the test actually failing before asking the AI to write code.
5.2. Cycle 2: triangulation
Add a test that forces the code to actually compute (triangulation):
[Fact]
public void Total_AtThreshold_Gets10PercentOff()
{
var pricing = new OrderPricing();
Assert.Equal(900_000m, pricing.Total(1_000_000m, isVip: false));
}
public decimal Total(decimal subtotal, bool isVip)
=> subtotal >= 1_000_000m ? subtotal * 0.90m : subtotal;
The boundary value (exactly at the threshold) is where > versus >= errors tend to appear. An AI assistant may pick the wrong boundary if the requirement is described only in words; a boundary test rules this out.
5.3. Cycle 3: VIP, invalid input and refactoring
[Fact]
public void Total_VipUnderThreshold_Gets5PercentOff()
{
var pricing = new OrderPricing();
Assert.Equal(475_000m, pricing.Total(500_000m, isVip: true));
}
[Fact]
public void Total_VipAtThreshold_StacksTo15Percent()
{
var pricing = new OrderPricing();
Assert.Equal(850_000m, pricing.Total(1_000_000m, isVip: true));
}
[Fact]
public void Total_NegativeSubtotal_Throws()
{
var pricing = new OrderPricing();
Assert.Throws<ArgumentOutOfRangeException>(
() => pricing.Total(-1m, isVip: false));
}
Once all tests pass, the Refactor phase removes the nested conditions and names the business constants:
public class OrderPricing
{
const decimal BulkThreshold = 1_000_000m;
const decimal BulkRate = 0.10m;
const decimal VipRate = 0.05m;
public decimal Total(decimal subtotal, bool isVip)
{
if (subtotal < 0)
throw new ArgumentOutOfRangeException(nameof(subtotal));
var rate = 0m;
if (subtotal >= BulkThreshold) rate += BulkRate;
if (isVip) rate += VipRate;
return subtotal * (1 - rate);
}
}
The stacking rule now lives in one place. If a model or a human later switches to multiplicative discounts, the StacksTo15Percent test fails immediately. This is the "safety net" function of a test suite, and it matters more when code is changed in bulk by AI.
Verification note. The expected values (475,000; 850,000; 900,000) were computed by hand from the rules. The C# code in this article was not compiled or run during writing. Run
dotnet testbefore use.
5.4. Common deviations
- Writing many tests at once before asking for code. The short feedback loop, a factor associated with good results (Fucci et al., 2017), is lost. With AI the risk grows because a model easily produces a large block of code that is hard to review.
- Skipping Refactor. Doing only Red-Green yields working but poorly structured code, and the tests gradually become a maintenance burden.
- Tests tied to implementation rather than observable behavior. Changing the implementation breaks the tests even though behavior is unchanged.
- Letting AI write both tests and code without reviewing the tests (see 4.2).
6. Limitations and scope of application
Schools. When code has dependencies, TDD splits into two approaches. The Chicago (classicist) school checks resulting state, uses real objects or simple fakes, and works from the inside out. The London (mockist) school checks interactions between objects using mocks, and works from the outside in (Fowler, 2007; Freeman & Pryce, 2009). For pure business logic such as the example in Section 5, the classicist approach is usually less coupled to internal structure. Whether a dependency can be swapped for a test double depends on a design following dependency inversion, covered in the Clean Architecture lesson; for repositories specifically, the Unit of Work and Repository lesson explains why integration tests often remove the need to mock (both in Vietnamese).
Poorly suited situations.
- Exploratory code (spikes). When it is unclear what to build, writing tests first only fixes an assumption that may be wrong. Explore, extract the insight, discard the code and rewrite it with TDD.
- UI and visual effects. Results need to be judged by human eyes; snapshot-style tests give low benefit relative to maintenance cost.
- Infrastructure integration (databases, queues, networks). Mocking infrastructure easily creates false comfort; deliberate integration tests are more reliable (see the API Testing lesson on WebApplicationFactory and Testcontainers, in Vietnamese).
- Legacy code without seams. Dependencies must be broken first, following Feathers (2004), before TDD can be applied.
Threats to validity of this review. The figures in Section 3 largely predate AI assistants; those in 4.1 and 4.2 come from small-scale problems and benchmarks. Section 4.3 is a proposal that has not been empirically validated.
7. Handling of sources
The figures of Nagappan et al., Rafique & Mišić, Fucci et al., Causevic et al., George & Williams, Fakhoury et al., Liang et al., Liu et al., Mathews & Nagappan, Cui, Yu et al. and Bhati were checked against the published abstracts at the source (publisher, arXiv) or through search results quoting the abstract. Some publisher pages block automated access, so for George & Williams and Romano et al. only secondary sources could be checked. For Piya & Sullivan only the abstract could be checked and it contains no quantitative results; for Karac & Turhan, Guimaraes & Nascimento and Cassieri et al. only publication details and a description of the design could be checked, without detailed results. The 2011 experiment published in Information and Software Technology (Pančur & Ciglarič, 53(6), 557–573) was left out because the secondary summaries found describe its results inconsistently. TDD-Agent is an arXiv preprint that has not been peer reviewed. When citing for academic purposes, open the original paper to verify sample size, measures, confidence intervals and experimental conditions. Statements about AI-DLC rest on the method author's own introduction, not on independent evaluation.
Adoption checklist
Before confirming an AI-assisted TDD process
- •Each new test is run and seen failing for the expected reason before AI writes code
- •Humans write or approve every test; tests drafted by AI are reviewed against the requirements
- •Each Red-Green-Refactor cycle is small enough to review the AI's output
- •The Refactor phase actually happens, not skipped under pressure
- •Tests describe observable behavior, not implementation details
- •Boundary values (threshold, empty, negative, null) are written as tests early
- Scope is appropriate: pure business logic, not forced onto UI or exploratory code
Frequently asked questions
Frequently asked questions about TDD and AI-DLC
Does TDD really reduce defects?
There are signs that it does, but the magnitude depends on context. The case study of four industrial teams by Nagappan et al. (2008) recorded a 40–90% drop in defect density versus comparable projects not using TDD. The meta-analysis of 27 studies by Rafique and Mišić (2013) found only a small improvement in external quality, although the improvement in industrial studies was larger than in academic ones.
How much does TDD slow development down?
Nagappan et al. reported initial development time rising by about 15–35%. This cost is expected to be recovered in bug fixing and maintenance, so it is usually reasonable for long-lived products and less so for short-lived experiments.
What is AI-DLC, and does it require TDD?
AI-DLC is a software development method proposed by Raja SP (AWS) in 2025, with three phases (Inception, Construction and Operations) in which AI executes while humans retain decision authority. The introduction to the method describes tests as an artifact generated by AI during Construction and does not prescribe TDD. Combining TDD with AI-DLC in this article is the author's proposal.
Why not let AI write both the tests and the code?
If the same model misreads a requirement, the tests and the code will be wrong consistently and still pass. The EvalPlus study (Liu et al., 2023) shows that an insufficient test suite can inflate pass rates by up to 19.3–28.9% and mis-rank models. Systems such as TDD-Agent (Yu et al., 2026, preprint) show that letting AI generate and refine tests can improve benchmark results, but they do not measure fitness to business intent. AI should therefore help draft tests while humans review them, because that is where correctness is defined.
Does giving tests to an LLM really improve code generation?
Benchmark studies point the same way. Mathews and Nagappan (2024) found that adding tests to the problem statement helped GPT-4 and Llama 3 solve more problems on MBPP and HumanEval. Fakhoury et al. (2024) recorded an average absolute pass@1 gain of 45.97% with a test-based interactive workflow. The limitation is that the experiments use small benchmark problems and are measured by the same test suites, so they do not transfer directly to industrial systems.
Is writing tests after the code worse than TDD?
Not necessarily. The experiment by Fucci et al. (2017) with 39 professional developers concluded that the order of writing tests and code had no important influence; quality and productivity were tied to working in small, steady steps. However, tests written afterward tend to mirror the existing code and skip the hard cases.
Conclusion
TDD is not a universal solution, and the empirical evidence is not strong enough to treat it as a mandatory principle. What the studies support is broader than the "TDD" label: work in small steps, get automated feedback after each step, and refactor regularly.
In the context of AI-DLC, the value of TDD shifts from helping developers write correct code to allowing humans to define correctness as an executable specification and delegate the implementation to AI. Three main conclusions:
- Costs and benefits are both real. Fewer defects in exchange for up-front time; weigh them against the product's lifetime.
- Small steps have the most supporting evidence. With AI-generated code, small steps also keep the volume to be reviewed within what humans can control.
- Tests are where human authority is retained. The quality of the test suite caps every claim about the correctness of AI-generated code; the Red phase should therefore be owned by humans.
The next research step is to measure this combined process at industrial project scale, where no direct evidence currently exists.
References
- Beck, K. (2002). Test-Driven Development: By Example. Addison-Wesley. https://dl.acm.org/doi/10.5555/579193
- Raja SP (2025, 31 July). AI-Driven Development Life Cycle: Reimagining Software Engineering. AWS DevOps & Developer Productivity Blog. https://aws.amazon.com/blogs/devops/ai-driven-development-life-cycle
- Nagappan, N., Maximilien, E. M., Bhat, T., Williams, L. (2008). Realizing quality improvement through test driven development: results and experiences of four industrial teams. Empirical Software Engineering, 13(3), 289–302. https://www.microsoft.com/en-us/research/wp-content/uploads/2009/10/Realizing-Quality-Improvement-Through-Test-Driven-Development-Results-and-Experiences-of-Four-Industrial-Teams-nagappan_tdd.pdf
- Rafique, Y., Mišić, V. B. (2013). The effects of test-driven development on external quality and productivity: a meta-analysis. IEEE Transactions on Software Engineering, 39(6), 835–856. https://ieeexplore.ieee.org/document/6197200
- Fucci, D., Erdogmus, H., Turhan, B., Oivo, M., Juristo, N. (2017). A dissection of the test-driven development process: does it really matter to test-first or to test-last? IEEE Transactions on Software Engineering, 43(7), 597–614. https://doi.org/10.1109/TSE.2016.2616877
- Karac, I., Turhan, B. (2018). What do we (really) know about test-driven development? IEEE Software, 35(4), 81–85. https://doi.org/10.1109/MS.2018.2801554
- Causevic, A., Sundmark, D., Punnekkat, S. (2011). Factors limiting industrial adoption of test driven development: a systematic review. ICST 2011, 337–346. https://doi.org/10.1109/ICST.2011.19
- George, B., Williams, L. (2004). A structured experiment of test-driven development. Information and Software Technology, 46(5), 337–342. https://www.sciencedirect.com/science/article/pii/S0950584903002040
- Romano, S., Fucci, D., Scanniello, G., Turhan, B., Juristo, N. (2017). Findings from a multi-method study on test-driven development. Information and Software Technology, 89, 64–77. https://doi.org/10.1016/j.infsof.2017.03.010
- Mathews, N. S., Nagappan, M. (2024). Test-driven development for code generation. arXiv:2402.13521. https://arxiv.org/abs/2402.13521
- Cui, Y. (2025). Tests as prompt: a test-driven-development benchmark for LLM code generation. arXiv:2505.09027. https://arxiv.org/abs/2505.09027
- Yu, H., Li, K., Li, J., Chai, H., Yuan, Y., He, R., Wei, J. (2026). TDD-Agent: test-driven reasoning for code generation. arXiv:2608.16742 (preprint, not peer reviewed). https://arxiv.org/abs/2608.16742
- Cassieri, P., Romano, S., Lenarduzzi, V., Taibi, D., Scanniello, G. A multi-study evaluation into generative artificial intelligence for test-driven development. ACM Transactions on Software Engineering and Methodology. https://doi.org/10.1145/3828553
- Guimaraes, E., Nascimento, N. (2025). AI in the software development lifecycle: insights and open research questions. FSE Companion 2025, 1353–1357. https://doi.org/10.1145/3696630.3730538
- Bhati, H. (2026). Agentic AI in the software development lifecycle: architecture, empirical evidence, and the reshaping of software engineering. arXiv:2604.26275. https://arxiv.org/abs/2604.26275
- Fakhoury, S., Naik, A., Sakkas, G., Chakraborty, S., Lahiri, S. K. (2024). LLM-based test-driven interactive code generation: user study and empirical evaluation. IEEE Transactions on Software Engineering, 50(9), 2254–2268. https://arxiv.org/abs/2404.10100
- Liang, Y., Ying, R., Ni, S., Cui, Z. (2026). Scaling test-driven code generation from functions to classes: an empirical study. arXiv:2602.03557. https://arxiv.org/abs/2602.03557
- Piya, S., Sullivan, A. (2023). LLM4TDD: best practices for test driven development using large language models. arXiv:2312.04687. https://arxiv.org/abs/2312.04687
- Liu, J., Xia, C. S., Wang, Y., Zhang, L. (2023). Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation. arXiv:2305.01210. https://arxiv.org/abs/2305.01210
- Fowler, M. (2007). Mocks Aren't Stubs. martinfowler.com. https://martinfowler.com/articles/mocksArentStubs.html
- Freeman, S., Pryce, N. (2009). Growing Object-Oriented Software, Guided by Tests. Addison-Wesley.
- Feathers, M. (2004). Working Effectively with Legacy Code. Prentice Hall.
Last updated: October 2026
Related posts
Vietnamese-language pages on this site:
- AI-DD: AI-Driven software development, a comprehensive series: general context on AI-assisted development.
- AI-DD - Part 3: Figures, field experience and risks: risks of delegating code to AI.
- Installing skills and coding on: the mechanism and the knowledge left behind: feeding project knowledge to AI agents.
- 9.8 - API Testing: integration testing with WebApplicationFactory and Testcontainers.
- 13.9 - Unit of Work and Repository Pattern: when to mock a repository and when not to.
- 16.3 - Clean Architecture: dependency inversion and testability.
- Types of API testing: 9 types, how they differ and when to run them: a map of test levels beyond the unit level.
- All posts on testing
