TL;DR
We’ve spent decades treating documentation like that boring homework you copy-paste at 11 PM before the deadline. Turns out, that was a mistake. And now that AI agents are starting to write our code, it’s about to become a catastrophic one. Five decades of academic research shows that good specs reduce bugs, accelerate delivery, and now serve as the primary interface between what you want and what the machine builds. If you thought documentation was optional before, welcome to the era where it’s the difference between shipping fast and debugging AI-generated chaos.
Introduction
Here’s a question that’s been haunting software engineering since the 1960s: if documentation is so valuable, why does everyone treat it like digital garbage?
Every team I’ve worked with has the same pattern. Someone writes code, ships it, and then maybe, if the PM is really annoying about it, adds some half-hearted comments three weeks later. By then, the code has already changed twice and the documentation is lying to you. We all know it’s lying. We read it anyway, then immediately check the source code to see what’s actually happening.
Academics have been studying this dysfunction for fifty years. The research is clear and it’s been clear forever: good documentation correlates with fewer bugs, faster delivery, and teams that don’t hate each other. Spec-first approaches like Cleanroom methodology achieved a tenfold reduction in defects compared to the „write code first, ask questions later“ approach that most of us still use.[1]
But here’s the twist. AI is about to make this academic debate very, very practical. Because when an AI agent writes your code, it doesn’t have the luxury of reading your mind or asking clarifying questions over Slack. It has exactly one input: your specification. If that spec is vague, outdated, or missing, you’re about to generate a whole lot of confidently wrong code.
Let’s talk about what five decades of research actually tells us, why we’ve been ignoring it, and why ignoring it now might be professionally suicidal.
The Academic World Has Been Screaming At Us For Decades
Software documentation has an identity crisis.
A 2017 systematic study analyzed 270 papers on documentation taxonomies and found something hilarious: 97% of them targeted completely unique subject matters.[2] Translation: we can’t even agree on what documentation is supposed to be, let alone how to do it well. Some researchers obsess over inline comments. Others focus on architectural diagrams. A few weirdos are really into README files.
The closest thing we have to a unified framework comes from Carnegie Mellon’s Software Engineering Institute. Their book, Documenting Software Architectures: Views and Beyond, is basically the Bible for people who care about this stuff.[3] It introduces the idea that documentation isn’t one monolithic thing. It’s multiple perspectives: module views, component-and-connector views, allocation views. You wouldn’t use a single blueprint to build a house, wire the electricity, and landscape the yard. Same logic applies here.
At the code level, things get even messier. A 2019 study classified over 40,000 lines of Java comments and found 16 distinct categories.[4] Some comments summarize what code does. Others explain why a decision was made. Some mark incomplete work with TODOs. Others warn you about deprecated methods that will explode if you touch them. A follow-up study in 2023 identified 11 types of „comment smells,“ basically documentation anti-patterns that make things worse instead of better.[5]
And READMEs? A study of 4,226 README sections from GitHub repos found that most developers obsess over the „what“ and „how“ while completely ignoring the „why.“[6] You get a wall of installation instructions and API examples, but no explanation of what problem the thing actually solves or why you’d use it instead of the twelve other libraries that do something similar.
The taxonomy problem isn’t just academic navel-gazing. It reflects a real issue: we treat „documentation“ like it’s one thing, so we apply one-size-fits-all advice like „write more comments.“ But inline comments serve a completely different purpose than architectural decision records, which serve a completely different purpose than API references. Telling someone to „document better“ without specifying what kind is like telling them to „build better“ without saying whether you mean a skyscraper or a sandwich.
Developers Use Documentation But Don’t Trust It (And They’re Right)
The foundational study on how engineers actually use documentation came out in 2003.[7] Researchers surveyed, interviewed, and observed developers, and the findings have held up for over twenty years:
Engineers read documentation, but they don’t believe it. They verify everything against the source code.
Why? Because documentation lies. Not maliciously. It just decays. Code changes fast. Documentation changes slow. The gap between them grows until the docs are basically fanfiction.
A 2015 industrial case study at a GPS company found that developers use different documentation for different tasks.[8] If you’re implementing a new feature, you need architectural overviews and design rationale. If you’re fixing a bug, you need inline comments and error-handling notes. If you’re onboarding, you need high-level context and examples. One-size-fits-all documentation fails all of these use cases.
The most systematic research on documentation problems comes from a pair of studies in 2019 and 2020.[9][10] Researchers mined 878 documentation-related complaints from mailing lists, Stack Overflow, issue trackers, and pull requests. Then they surveyed 146 practitioners to confirm what they found.
The top problems: insufficient content, outdated information, and ambiguous phrasing. The primary cause: lack of time. Developers know documentation matters. They just never have time to do it properly. And when they do have time, they’re updating the code instead of the docs.
Here’s the kicker: for API documentation, ambiguity is the worst failure mode.[11] Developers surveyed in 2015 said that unclear API docs were worse than missing docs. At least with missing docs, you know you’re on your own. Ambiguous docs give you false confidence, so you build something that breaks in production.
Six out of ten documentation problems were classified as „blockers,“ meaning developers gave up on the API entirely rather than try to figure it out. If your documentation is bad enough, people just won’t use your tool. That’s not a documentation problem anymore. That’s a business problem.
Documentation Rots In Predictable Ways (And We’ve Measured It)
Code and comments don’t evolve together. They evolve near each other and hope for the best.
The largest study on this problem analyzed 1.3 billion code changes from 1,500 Java projects.[12] They built an AST-level diff tool to track how comments changed relative to code, and they found something depressing: refactoring operations constantly create stale comments. You rename a variable, extract a method, inline a function. The code is fine. The comment still refers to the old structure.
A 2009 study found that 97% of comment changes happen in the same commit as the code change.[13] That sounds good until you realize what it implies: 3% of the time, comments are updated separately. And when they’re updated separately, they’re usually updated later, which means there’s a window where the code and docs are out of sync. For API documentation specifically, that window can be weeks or months.
The really scary finding came in 2012: inconsistencies between code and comments correlate with bugs.[14] When a function’s logic changes but the comment doesn’t, the probability of introducing a defect spikes. This makes intuitive sense. The comment represents the developer’s mental model. If the comment is wrong, the mental model is wrong, and the next person to touch that code is working from faulty assumptions.
This is the documentation decay cycle:
- Developer writes code and comments together
- Code evolves faster than comments
- Comments become misleading
- Next developer trusts the comments
- Next developer introduces bug
- Someone fixes bug but not comment
- Repeat
We’ve known about this cycle for over a decade. We still haven’t fixed it. And now we’re about to feed all of this rotting documentation into AI models and ask them to write production code.
The ROI Of Documentation Is Obvious But Unmeasured
Here’s an embarrassing fact: we have no idea what documentation actually costs.
A 2015 systematic review searched the entire academic literature for studies measuring documentation costs and benefits.[15] They found tons of papers proposing new documentation techniques. They found plenty of papers with anecdotal evidence about documentation value. But systematic cost models? Methods for measuring ROI? Nope.
We know documentation debt is expensive. A 2019 study on technical debt found that outdated or missing documentation significantly increases maintenance costs.[16] McKinsey estimates that tech debt accounts for 40% of IT balance sheets, with documentation debt adding 10 to 20% in extra project costs.[17][18] But those are rough estimates, not rigorous measurements.
The closest we get to hard numbers is in specific contexts. For onboarding, a 2024 systematic review found that structured documentation guides newcomers through first contributions and removes barriers to productivity.[19] A Microsoft case study found that onboarding tasks like fixing and updating documentation correlate with long-term job satisfaction and performance.[20] But even here, we’re talking about correlation, not controlled experiments.
API documentation has the strongest empirical support. A controlled experiment in 2020 found that developers working with optimized API docs made fewer errors and were faster at both planning and execution.[21] A field study of 440 Microsoft developers found that the most severe obstacles to learning APIs were documentation-related.[22]
So we know documentation helps. We just don’t know how much, or how much it costs to produce, or what the breakeven point is. This makes it really hard to justify documentation time when you’re in a sprint planning meeting and the PM is asking why you’re not writing features.
But here’s the thing: the ROI calculation is about to get a lot clearer. Because if AI-generated code quality depends on documentation quality, then every hour you spend writing good specs translates directly into fewer hours debugging AI mistakes. That’s a measurable ROI. We just haven’t measured it yet.
Spec-First Development Has Fifty Years Of Evidence Behind It
The idea that you should write specifications before writing code is not new. It’s not even middle-aged. It’s old enough to collect Social Security.
Floyd, Hoare, and Dijkstra established the mathematical foundations of program correctness in the late 1960s and early 1970s.[23] They treated specifications as formal logic statements from which you could prove that code behaved correctly. This work spawned entire communities around formal methods like Z notation, VDM, and the B Method.[24][25]
The practical version came from Bertrand Meyer in 1992 with „Design by Contract.“[26] Meyer’s idea: software development is a series of contracts between modules. Each contract has preconditions (what the caller promises), postconditions (what the function promises), and invariants (what’s always true). If everyone honors their contracts, the system works. If someone breaks a contract, you know exactly who to blame.
Meyer argued against defensive programming, which he saw as adding complexity without adding clarity. Instead of checking for conditions that shouldn’t happen, write contracts that specify what should happen and let the system enforce them. This sounds abstract until you’ve debugged a null pointer exception at 2 AM and realized the function assumed its input wouldn’t be null but never actually said so.
David Parnas took this further. His 1972 paper on information hiding is one of the most cited papers in software engineering.[27] His 1985 paper with Clements argued that the design process should produce documentation as its primary artifact, not code.[28] Parnas returned to this theme in 2011: „The prime cause of the sorry state of the art in software development is our failure to produce good design documentation.“[29]
The synthesis of these ideas came in 2004 with „Agile Specification-Driven Development.“[30] The authors combined Test-Driven Development with Design by Contract, arguing that tests and contracts are complementary types of specifications. Tests show examples. Contracts state invariants. Together, they define behavior more precisely than either alone. Their punchline: SDD is better than TDD or DbC individually.[31]
Cleanroom Proved Spec-First Works (We Just Ignored It)
The strongest empirical evidence for spec-first development comes from Cleanroom Software Engineering, developed at IBM in the 1980s and 1990s.[32]
Cleanroom was radical. You write formal specifications first. Then you write code. Then you review the code in teams to verify correctness. You explicitly do not debug. Instead, you use statistical testing to certify quality. The whole process is incremental and under rigorous quality control.
The results were absurd. Over 90% of defects were found before testing, compared to 60% in conventional approaches. HP projects achieved 1 defect per thousand lines of code versus an industry average of 5 to 10. That’s a five to tenfold improvement.[32]
A NASA project in 1990 processed 34,000 lines of FORTRAN using Cleanroom. Productivity: 4.9 lines per staff-hour versus a typical 2.9. That’s a 70% improvement. Error rate: 3.3 per thousand lines versus 6.0 typical.[32]
For formal methods in safety-critical systems, the evidence is even more extreme. The Paris Metro used the B Method to formally verify 9,000 lines of code for a train control system. That system carried 800,000 passengers daily on trains running two minutes apart. Zero defects in the verified code.[33]
So why didn’t everyone adopt this? Because it’s hard. Writing formal specs takes time and skill. Reviewing code without running it takes discipline. Statistical testing requires statistical expertise. And most projects don’t build metro systems where a bug kills hundreds of people.
But the research is clear: if you invest in specifications and verification before testing, you catch defects earlier, build faster, and ship more reliable code. We’ve known this since the 1980s. We just decided it wasn’t worth the effort.
Now AI is making it worth the effort. Because AI doesn’t need you to write perfect formal proofs. It just needs you to be clear about what you want. And if being clear reduces defects by 5x, that’s a pretty good trade.
AI Code Generation Lives Or Dies By Specification Quality
LLMs are really good at writing code. They’re terrible at guessing what you meant.
The first systematic study on this was Self-Planning Code Generation in 2024.[34] Researchers found that if you make the model produce a high-level plan before writing code, quality improves. The plan acts like a functional spec. The model thinks through the steps, identifies edge cases, and then generates code that matches the plan. Without the planning step, the model jumps straight to implementation and makes mistakes.
This isn’t subtle. A 2024 paper literally titled „Requirements Are All You Need“ argues that high-quality requirements are now the primary bottleneck in software development.[35] The better your spec, the better the generated code. The worse your spec, the more iteration you need. The paper introduces „Progressive Prompting,“ which is just a fancy term for breaking requirements into smaller, clearer pieces and feeding them to the model stepwise.
Prompt engineering research confirms this. Different specification formats require different prompting strategies, and choosing the right strategy can improve accuracy by up to 1.9% while reducing token usage by 74.8%.[36] AlphaCodium argues that code generation is fundamentally harder than natural language tasks because code requires exact syntax, edge case handling, and attention to tiny details in the spec.[37]
The performance gap between benchmarks and reality is telling. GPT-4 achieves 96% on HumanEval but only 76% on HumanEval Pro.[38] The difference? HumanEval Pro has more complex, ambiguous specifications. NaturalCodeBench, which uses real-world scenarios, drops GPT-4 to around 53% pass rate.[39] The models aren’t getting dumber. The specifications are getting harder.
Agentic AI Turns Specs Into Coordination Protocols
Multi-agent systems make this even more critical.
Systems like MetaGPT simulate an entire software company.[40] One agent plays Product Manager and writes requirements. Another plays Architect and designs the system. Another plays Engineer and writes code. Another plays QA and writes tests. They pass artifacts to each other in sequence.
The whole thing only works if the artifacts are clear. If the Product Manager agent writes a vague requirement, the Architect agent builds the wrong design, the Engineer agent implements the wrong feature, and the QA agent writes tests that pass but don’t validate the actual requirement. Garbage in, garbage out, but now with four layers of garbage.
SWE-bench is the standard benchmark for autonomous coding agents.[41] Agents are given a GitHub issue (the specification) and a codebase, and they have to fix the issue. Top agents now hit 71% on the verified set. But here’s the problem: 32.67% of successful patches had „solution leakage,“ meaning the answer was literally in the issue description.[42] The agents aren’t solving hard problems. They’re pattern-matching against specs that accidentally contain the solution.
The DORA 2024 report found that a 25% increase in AI adoption correlates with a 7.5% improvement in documentation quality.[43] This is correlation, not causation, but the pattern makes sense. Teams that adopt AI invest in better specs because they realize the AI is only as good as its inputs. Teams that don’t adopt AI keep writing vague specs because humans can tolerate vagueness.
Industry research backs this up. Capgemini found that 85% of software professionals expect to use generative AI by 2026, up from 46% in 2024.[44] Gartner predicts 33% of enterprise applications will include agentic AI by 2028.[45] Microsoft research found that explicit specifications reduced iterative refinement by 68%.[46]
Translation: the better your spec, the less you argue with the AI. That’s a productivity gain you can measure in hours saved per sprint.
Documentation Is Becoming The Primary Interface Layer
Here’s the paradox we started with: documentation has always been valuable, but engineers have always undervalued it.
Lethbridge’s 2003 study showed that engineers use documentation but verify it against code.[7] Aghajani’s 2020 study found that lack of time is the primary barrier to good documentation.[10] This gap has persisted for decades.
AI resolves the paradox by changing the audience. Humans can tolerate vague or outdated docs because we can read code, ask questions, and infer intent. AI can’t. It needs the documentation to be correct, complete, and unambiguous.
RAG systems (Retrieval-Augmented Generation) work by pulling documentation into context when generating code.[47] If your API docs are wrong, the generated code is wrong. If your docs are incomplete, the generated code has gaps. If your docs are ambiguous, the generated code makes arbitrary choices.
GitHub Copilot for Docs explicitly pulls from official library documentation and cites sources.[47] The model defers to the docs as the source of truth. This is good if your docs are accurate. It’s catastrophic if they’re not.
For OpenAPI specs, RAG systems chunk by endpoint so each chunk contains all the information for a specific API call.[48] For code, systems generate natural language descriptions because embeddings alone don’t capture semantic meaning.[48] Documentation becomes the bridge between human intent and machine execution.
Meta’s React team found that documentation-primed prompts resulted in code that was 78% more likely to follow recommended patterns.[46] If AI systems amplify documentation quality into generated code at scale, the stakes for getting specs right just went exponential.
Conclusion
Let’s summarize what fifty years of research actually tells us:
Documentation isn’t one thing. It’s a whole taxonomy of artifacts, each serving different purposes. Treating it as a monolith guarantees you’ll do it wrong.
Developers use documentation but don’t trust it because documentation decays. Code evolves fast. Docs evolve slow. The gap grows until the docs are lying to you. We’ve measured this. It’s predictable. We still haven’t fixed it.
The ROI of documentation is obvious but unmeasured. We know it helps with onboarding, API learning, and maintenance. We just don’t have rigorous cost models. AI is about to make the ROI measurable: better specs mean less debugging.
Spec-first development has fifty years of evidence behind it. From formal methods to Design by Contract to Cleanroom’s tenfold defect reduction, the research is unanimous. Writing specifications before code reduces bugs and accelerates delivery.
AI code generation depends acutely on specification quality. LLMs can’t read your mind. They can’t ask clarifying questions. They have exactly one input: your spec. If it’s vague, the code is wrong. If it’s clear, the code is right. The performance gap between benchmarks and reality proves this.
Agentic systems turn specs into coordination protocols. When multiple AI agents collaborate, the specs are how they communicate. Bad specs cascade through the entire pipeline, creating compounding errors.
The paradigm is shifting. Documentation used to be for humans who could tolerate ambiguity. Now it’s for AI systems that can’t. This doesn’t diminish the developer’s role. It elevates it. You’re no longer writing code. You’re writing the intent from which code is generated.
Microsoft found that explicit specifications reduce iteration by 68%. Top agentic systems solve 71% of real-world issues. Documentation quality correlates with AI adoption. The transition is already happening.
The question isn’t whether this shift is coming. The question is whether you’ll treat documentation as the afterthought it’s always been, or as the primary interface it’s about to become. Because in the agentic era, your AI is only as good as your specs. And your specs, historically, have been pretty bad.
Time to fix that.
References
[1] R. Linger and C. Trammell, „Cleanroom Software Engineering Reference,“ Technical Report CMU/SEI-96-TR-022, Software Engineering Institute, Carnegie Mellon University, 1996. PDF
[2] M. Usman, R. Britto, J. Börstler, and E. Mendes, „Taxonomies in software engineering: A Systematic mapping study and a revised taxonomy development method,“ Information and Software Technology, vol. 85, pp. 43-59, 2017. Link
[3] P. Clements, F. Bachmann, L. Bass, D. Garlan, J. Ivers, R. Little, P. Merson, R. Nord, and J. Stafford, „Documenting Software Architectures: Views and Beyond,“ 2nd Edition, Addison-Wesley Professional, 2010. ISBN 978-0-321-55268-6.
[4] L. Pascarella, M. Bruntink, and A. Bacchelli, „Classifying code comments in Java software systems,“ Empirical Software Engineering, vol. 24, no. 3, pp. 1499-1537, 2019. DOI
[5] E. Jabrayilzade et al., „Taxonomy of inline code comment smells,“ Empirical Software Engineering, 2024. Link
[6] G. A. A. Prana, C. Treude, F. Thung et al., „Categorizing the Content of GitHub README Files,“ Empirical Software Engineering, vol. 24, pp. 1296-1327, 2019. Link
[7] T. C. Lethbridge, J. Singer, and A. Forward, „How software engineers use documentation: the state of the practice,“ IEEE Software, vol. 20, no. 6, pp. 35-39, 2003. DOI
[8] Garousi, G., Garousi-Yusifoğlu, V., Ruhe, G. et al., „Usage and usefulness of technical software documentation: An industrial case study,“ Information and Software Technology, vol. 57, no. 1, pp. 664-682, 2015. Link
[9] E. Aghajani, C. Nagy, O. L. Vega-Márquez, M. Linares-Vásquez, L. Moreno, G. Bavota, and M. Lanza, „Software Documentation Issues Unveiled,“ in Proceedings of the 41st International Conference on Software Engineering (ICSE 2019), pp. 1199-1210. DOI
[10] E. Aghajani, C. Nagy, M. Linares-Vásquez, L. Moreno, G. Bavota, M. Lanza, and D. C. Shepherd, „Software Documentation: The Practitioners‘ Perspective,“ in Proceedings of the 42nd International Conference on Software Engineering (ICSE 2020), pp. 590-601. Link
[11] G. Uddin and M. P. Robillard, „How API documentation fails,“ IEEE Software, vol. 32, no. 4, pp. 68-75, 2015. Link
[12] F. Wen, C. Nagy, G. Bavota, and M. Lanza, „A Large-Scale Empirical Study on Code-Comment Inconsistencies,“ in Proceedings of the 27th International Conference on Program Comprehension (ICPC 2019). DOI
[13] B. Fluri, M. Würsch, E. Giger et al., „Analyzing the co-evolution of comments and source code,“ Software Quality Journal, vol. 17, pp. 367-394, 2009. Link
[14] W. M. Ibrahim, N. Bettenburg, B. Adams, and A. E. Hassan, „On the relationship between comment update practices and software bugs,“ Journal of Systems and Software, vol. 85, no. 10, pp. 2293-2304, 2012.
[15] J. Zhi, V. Garousi-Yusifoğlu, B. Sun, G. Garousi, S. Shahnewaz, and G. Ruhe, „Cost, benefits and quality of software development documentation: A systematic mapping,“ Journal of Systems and Software, vol. 99, pp. 175-198, 2015. Link
[16] L. Mendes, C. Cerdeiral, and G. Santos, „Documentation Technical Debt,“ in Proceedings of the XXXIII Brazilian Symposium on Software Engineering (SBES 2019), pp. 447-451, 2019. DOI
[17] McKinsey & Company, „Tech debt: Reclaiming tech equity,“ October 2020. Link
[18] McKinsey & Company, „Tech debt: Reclaiming tech equity,“ October 2020. Link
[19] I. Santos, K. R. Felizardo, I. Steinmacher, and M. A. Gerosa, „Software Solutions for Newcomers‘ Onboarding in Software Projects: A Systematic Literature Review,“ August 2024. arXiv:2408.15989
[20] P. Rodeghero, T. Zimmermann, B. Houck, and D. Ford, „Please Turn Your Cameras On: Remote Onboarding of Software Developers during a Pandemic,“ in Proceedings of ICSE-SEIP 2021. Link
[21] M. Meng, S. M. Steinhardt, and A. Schubert, „Optimizing API Documentation: Some Guidelines and Effects,“ in Proceedings of the 38th ACM International Conference on Design of Communication (SIGDOC 2020), 2020. DOI
[22] M. P. Robillard and R. DeLine, „A field study of API learning obstacles,“ Empirical Software Engineering, vol. 16, pp. 703-732, 2011. DOI
[23] R. W. Floyd, „Assigning meaning to programs,“ in Mathematical aspects of computer science, AMS, pp. 19-32, 1967; C. A. R. Hoare, „An Axiomatic Basis for Computer Programming,“ Communications of the ACM, 1969; E. Dijkstra, „A discipline of programming,“ Engelwood Cliffs: Prentice-Hall, 1976.
[24] Z Notation, developed at Oxford University Programming Research Group, became BSI standard in 1981. Wikipedia
[25] J.-R. Abrial, „The B-Method,“ developed in the 1980s-1990s. Wikipedia
[26] B. Meyer, „Applying ‚design by contract‘,“ IEEE Computer, vol. 25, no. 10, pp. 40-51, October 1992. DOI
[27] D. L. Parnas, „On the Criteria to Be Used in Decomposing Systems into Modules,“ Communications of the ACM, vol. 15, no. 12, pp. 1053-1058, December 1972. Link
[28] D. L. Parnas and P. C. Clements, „A Rational Design Process: How and Why to Fake It,“ IEEE Transactions on Software Engineering, vol. SE-12, no. 2, February 1986. Link
[29] D. L. Parnas, 2011 quote on design documentation.
[30] J. S. Ostroff, D. Makalsky, and R. F. Paige, „Agile Specification-Driven Development,“ in Extreme Programming and Agile Processes in Software Engineering (XP 2004), Lecture Notes in Computer Science, vol. 3092, pp. 104-112, Springer, 2004. Link
[31] J. S. Ostroff, D. Makalsky, and R. F. Paige, „Agile Specification-Driven Development,“ XP 2004.
[32] R. Linger and C. Trammell, „Cleanroom Software Engineering Reference,“ Technical Report CMU/SEI-96-TR-022, 1996.
[33] Paris Metro SAET-METEOR, B Method formal verification, 1989-1998. Link
[34] Self-Planning Code Generation, ACM Transactions on Software Engineering and Methodology, 2024. DOI
[35] „Requirements are All You Need: From Requirements to Code with LLMs,“ arXiv, June 2024. arXiv:2406.10101
[36] C.-Y. Wang et al., „Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity (PET-Select),“ arXiv:2409.16416, September 2024. arXiv
[37] T. Ridnik, D. Kredo, and I. Friedman, „Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering,“ arXiv:2401.08500, January 2024. arXiv
[38] Z. Yu, Y. Zhao, A. Cohan, and X.-P. Zhang, „HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation,“ in Findings of ACL 2025, pp. 13253-13279, Vienna, Austria. Link
[39] S. Zhang et al., „NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts,“ arXiv:2405.04520, May 2024. arXiv
[40] „A Survey on Code Generation with LLM-based Agents,“ arXiv:2508.00083, 2024. arXiv
[41] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan, „SWE-bench: Can Language Models Resolve Real-World GitHub Issues?“ in Proceedings of ICLR 2024. Link
[42] „SWE-Bench+: Enhanced Coding Benchmark for LLMs,“ arXiv:2410.06992, October 2024. arXiv
[43] DORA, „Accelerate State of DevOps Report 2024,“ Google Cloud, 2024. Link
[44] Capgemini, „Generative AI in organizations 2024,“ 2024. Link
[45] Gartner, „Agentic AI Predictions,“ 2024-2028. Link
[46] Microsoft Research, „New Future of Work Report 2024,“ December 2024. PDF
[47] GitHub, „What is retrieval-augmented generation, and what does it do for generative AI?“ The GitHub Blog. Link
[48] „Advanced System Integration: Analyzing OpenAPI Chunking for Retrieval-Augmented Generation,“ arXiv:2411.19804, November 2024. arXiv; K. Tamai, „Requirements specification quality IEEE 2007,“ in Proceedings of the 15th IEEE International Requirements Engineering Conference (RE’07), pp. 69-78, October 2007. DOI