Skip to main content
AI-Developer/Agentic Coding

How to Build Production Software With AI Without Drowning in Technical Debt

An AI coding agent can write 300 lines in twenty seconds and hand you an unmaintainable codebase in two weeks. Here is the physics of why it breaks down, and the rules that keep you in control.

October 11, 202615 min read
#Agentic Coding#AI Agents#Context Window#Technical Debt#LLM#Software Architecture

Part 1 of 1 in the Agentic Coding seriesView all parts →

The Agentic Engineering Mindset

A coding agent can write three hundred lines in twenty seconds. It can also turn your codebase into an unmaintainable pile in two weeks. The difference is not the model — it is whether you understand the machine you are driving.

Primary Objective
Learn the physics of foundation models: why their capability is jagged, why the context window is finite working memory, and how to keep every task inside the zone where the model is sharp.

Open your terminal. Give an agent a prompt. It writes three hundred lines in twenty seconds. You spin up the local server, click around, and it works. You feel like you just gained superpowers.

Then week two arrives.

You ask for a small feature — a new billing tier or an extra user role — and the authentication flow breaks. You ask the model to fix authentication, and it rewrites your database queries. You ask it to fix the queries, and now thirty tests fail in files you did not know existed. Before long you are staring at a 1,500-line pull request full of odd abstractions, duplicated helpers, and code you cannot explain to anyone on your team.

You did not build software. You built technical debt at lightspeed.

None of this is a story about AI being bad. The same model that produced that pull request can produce clean, tested, production-grade code. The developers who get the second outcome are not better at writing prompts. They understand what the model actually is, why it fails the way it does, and which decisions a human must never delegate.

Grab your coffee. We are going to look past the hype and at the machine.

Phase 1 of 6

To stop writing messy AI code, start by understanding what an LLM actually is inside your workflow.


The Model Is an Operating System

An LLM in a coding agent is not autocomplete, and it is not a chat box. It behaves more like a new kind of processor: a context-driven execution engine that runs computation over digital information and mutates the state of your machine.

Once you frame it that way, the three parts of an agent session map onto hardware you already understand:

  • Your prompt is the instruction set. It defines what the machine should do right now.
  • The context window is working memory (RAM). It is everything the engine can currently see.
  • The model is the execution engine. It interprets the context and produces an action: a file edit, a command, a diff.

For the first time, natural language is the instruction set, and the context window is the system's memory. That single reframing explains almost every failure we are about to cover — because most people treat that working memory like a messy group chat instead of the finite resource it is.

Three Eras of Execution
The Evolution of Software Execution
1

Manual Logic

1950s–2000s
Input
Human Developer
Engine
Exact if/else rules
Output
CPU

Where it breaks · Any edge case you never wrote a rule for crashes the system.

2

Machine Learning

2010s
Input
Data Engineer
Engine
Trained weights
Output
Neural Network

Where it breaks · It learns statistical patterns, but only from the past data you curate.

3

The Context Engine

Now
Input
Human Intent
Engine
Context window (RAM)
Output
Foundation Model

Where it breaks · The context window is finite. Treat it like a group chat and output quality collapses.

Phase 2 of 6

The model is powerful. Now let us look at the thing it is quietly terrible at.


The Intern With No Street Smarts

Why does an agent write code that looks clean on the surface and then falls apart a week later? Because a coding agent behaves like a hyper-fast, slightly eccentric intern.

That intern has read every API doc, every manual, and every forum thread in human history. Their recall of syntax is superhuman. But they have zero company context, zero street smarts, and no long-term accountability. When you give a loose prompt — add payment checkout to our dashboard — the agent has one goal: produce an answer that satisfies the text pattern right now. It does not care what your codebase looks like in six months.

Here is a concrete way that destroys a codebase. You have an app where users sign in with Google but pay through Stripe. Both services expose an email field. Left alone, an agent will often match an incoming payment to a user by comparing the two email strings — because string matching is fast, easy, and passes the quick test on your laptop, where you happen to use the same address for both.

In production, people sign in with a personal account and pay with a corporate card. The string match fails, the payment lands on an unlinked record, credits are never granted, and support drowns. You never link financial transactions by mutable text. You generate an immutable user ID, pass it into the checkout metadata, and verify the webhook against that ID.

Production Failure
The Mutable Identity Bug
A user signs in with a personal Google account, then pays with a corporate card. Two systems, two email fields, one broken assumption.
The shortcut the agent takes
Stripe email
DB lookup by email string
No row matches
Payment captured. Zero credits assigned. The support inbox fills up.
The architecture you must specify
User UUID
usr_8f3a…
metadata.user_id at checkout
Webhook verifies UUID
Payment and credits are bound to one immutable identity.

The agent knows how to write both versions. It is not that it cannot do the right thing — it is that, left to itself, it defaults to the shortcut. The rule that follows from this is the spine of the whole series: the agent is the labor, but the human owns the architecture. You never delegate identity schemas, security boundaries, or data models to an intern.

Phase 3 of 6

So the model defaults to shortcuts. Why is it still so capable at some things and so clumsy at others?


Why AI Is Brilliant at Code and Foolish About the Car Wash

The answer is the reward loop the model was trained inside. Frontier models are shaped by reinforcement learning: the model attempts a solution, an automated verifier decides whether it was right, and the model receives a signal. Where a verifier exists, capability compounds. Where it does not, the model never gets better.

Code is verifiable. You can run a compiler, execute a test suite, and read an exit code. A wrong answer announces itself instantly, so the labs can run millions of trials a day and the model becomes superhuman at writing code that runs. Common sense is not verifiable. There is no unit test for street smarts, so there is no signal to optimize against.

Reinforcement Learning
The Verifiability Engine
Verifiable · the model flies
Compiling code
Running unit tests
Mathematical proofs
Formal board games
Automated reward signal
A verifier scores every attempt automatically. Millions of trials per day, per task.
Unverifiable · the model stumbles
Physical common sense
Architectural taste
Real-world context
Unwritten business rules
Automated reward signal
No unit test exists for street smarts. The reward loop never closes, so the model never improves here.

This produces what researchers call jagged intelligence: towering peaks of skill right next to deep canyons of ignorance. A frontier model can find a zero-day in a Linux kernel and then fail a question a child would catch:

The car wash is fifty meters down the street. Should I walk or drive?

The model answers, with total confidence, that you should walk — it is close, it is good exercise, and you save gas. It understands distance, health, and fuel economy, and completely misses that the point of going to a car wash is to wash the car. We are not building animals shaped by gravity and hunger. We are summoning statistical circuits shaped by text and training environments, and they have never had to learn what a car wash is for.

The practical consequence for you: if you want high performance, give the model a way to verify its own work. A test suite is not just quality insurance. It is the reward signal that turns a guess into a fact.

Phase 4 of 6

Verifiability explains why capability varies. It also explains exactly where it drops off.


The Capability Cliff

You cannot treat an LLM as a uniform genius across your codebase. Its skill is concentrated wherever the labs pointed their reinforcement learning — and that is almost always standard, abundant, public code.

If your task lands inside those circuits — idiomatic TypeScript, a React component, a SQL query, a Python data script — the model flies. If your task depends on a proprietary internal SDK, an undocumented endpoint, or a trade-off that only makes sense given your team's history, you are outside the training distribution. The model will hallucinate methods that do not exist, invent config keys, and confidently break things.

Jagged Intelligence
The Capability Boundary
SuperhumanStruggles
training cliff
Inside the lab RL loop
TypeScript and ReactSQL queriesPython data scriptsAlgorithm puzzles
Outside the training distribution
Your internal SDKsProprietary business rulesUndocumented APIsLong-horizon architecture

This is the wall most teams hit and mis-diagnose. When the model stumbles on your domain, the instinct is to prompt harder — to add more emphasis, more instructions, more "please be careful." That does not work. You cannot prompt your private context into a model that has never seen it. You have two real options: write an explicit specification, or build an automated test the model can check itself against. Both move the task back toward something verifiable.

Phase 5 of 6

You now know why capability is uneven. The final piece is why a session that started strong gets worse hour by hour.


Context Is Finite Working Memory

Watch what happens inside a single long session. It starts fast, focused, and correct. Forty-five minutes later, something shifts. The model begins ignoring negative constraints — things you explicitly told it not to do. It introduces circular bugs: fixing line 40 breaks line 12, and fixing line 12 breaks line 40 again. It apologizes, then repeats the same broken code.

You did not break the model. You ran into the attention degradation curve. In a transformer, every token must weigh its relationship against every other token, so the attention cost grows with the square of the sequence length. Early in a session, the window is nearly empty and the signal between your instructions and the code is sharp. As the window fills with drafts, stack traces, corrections, and dead snippets, that signal is stretched thin. The model can no longer tell a rule you set at the start from a throwaway hack you tried fifteen turns ago.

O(N²) Attention
The Attention Degradation Curve
SMART ZONEattention intact · sharp signalDUMB ZONEattention diluted~40%0100k200k+CONTEXT FILL · TOKENS IN WINDOWREASONING QUALITY
Smart Zone · 0–40%Dumb Zone · 60%+Hard reset boundary

Be skeptical of the million-token marketing. There is a real difference between retrieval and reasoning. Large windows are excellent at retrieval — scanning a thousand pages to find one invoice number. They are poor at active reasoning — tracking state, holding constraints, and generating valid code. Dump a four-hundred-thousand-token repository into a prompt and ask for a feature, and the model does not become smarter. It simply starts the session already deep inside the low-attention region.

Phase 6 of 6

If long sessions decay, the obvious fix sounds reasonable: summarize and keep going. It is one of the worst things you can do.


Compaction Is Memento Syndrome

Many tools offer compaction: when the context fills, ask a model to summarize the conversation, delete the history, and paste the summary at the top. The problem is that you are asking a model to summarize its own messy, iterative mistakes — and summaries are lossy in exactly the wrong places.

An agent wakes up with total amnesia each turn and knows only what is written in front of it. When a summary replaces the history, it drops the specific schema change, the subtle edge case, and the exact file path, and it quietly generalizes a precise migration into a vague bullet. The next turn guesses the missing details to keep moving, and those guesses are then treated as fact. That is conversational sediment, and it is how a codebase drifts without a single obvious mistake.

Conversational Sediment
The Compaction Corruption Cycle
1Messy session
A hundred turns of debugging, errors, and dead ends live in the window.
Subtle constraints are already buried.
2Auto-summarize
The tool summarizes the mess and deletes the original history.
Schema details and exact file paths disappear.
3Fill the gaps
On the next turn the agent guesses the missing details.
Guesses get accepted as fact.
4Silent drift
The summary is now the source of truth, and it is wrong.
Hidden bugs ship to production.
Every compaction compounds the drift
A constraint that was summarized away never comes back. You cannot un-lose it.

Elite AI engineers almost never compact. They use the cold reset. The moment a task is done — or the moment the agent starts circling the same bug — they commit the working code, wipe the window, and reload only the high-signal context they actually need.

text
MESSY WAY:   prompt → debug → error → compacted summary → hallucination
CLEAN WAY:   prompt → task done → /clear → fresh context → task done

A fresh agent holding three thousand tokens of sharp specification will beat an exhausted agent dragging a hundred and fifty thousand tokens of conversation every single time.


The Disappearing Middle Tier

The fastest way to add technical debt to an AI project is to have the agent write code that the model has already made obsolete. Historically, any real feature required a heavy middle tier of glue: services that parse, store, connect, and render.

Picture an app for restaurant diners. You photograph a text-only menu, and the app shows realistic pictures of each dish. Built the traditional way, that is a pipeline of services, and every service is a permanent maintenance cost with its own failure mode — an OCR step that breaks on cursive fonts, a regex parser that breaks on odd price columns, a database schema that needs migrations, an image-generation API, and a frontend that stitches it all together.

Architecture
The Disappearing Middle Tier
5 services → 1 call
Traditional · glue pipeline
OCR service
Extract raw text from the image
Backend parser
Regex out dish names and prices
Relational database
Match items against ingredient tables
Image generation API
Text prompts to food pictures
Frontend canvas
Render layout and match images to text
Modern · single inference
Multimodal model
Reads the pixels, generates the dish images, and renders them into the menu margins in one pass.

Now do it the modern way. You hand the raw image straight to a multimodal model with one instruction: identify the dishes, generate realistic depictions of them, and render the result into the blank margins of this image. The model reads the pixels, interprets the language, generates the imagery, and draws the output in a single inference. There is no OCR server, no schema, and no layout engine to keep alive.

The lesson generalizes. Before you instruct an agent to build parsers, databases, and API orchestrators, stop and ask whether that layer needs to exist — or whether a foundation model can map the input directly to the output. Most of the middle tier you used to hand-write is now a prompt.


The Rule of 40%

Everything so far becomes a single daily habit. Never code with an agent blind. If you do not know how full your context window is, you are driving without a speedometer and you will not notice the crash until the code is already broken.

Whether you use Claude Code, Cursor, Aider, OpenHands, or a custom terminal agent, make sure your status line shows the live session token count — something like 42k / 200k. Then hold yourself to three zones.

Interactive · Rule of 40%
Make Your Own Call
34%context filled
Caution Zone
fresh session30%50%window full
Reasoning quality92%
Operating instruction · 30–50%
Caution Zone

Finish the current task only. No new scope. Run the tests.

  1. Green, 0–30%. The model is at its sharpest. Use this window for architectural decisions, core logic, and the first version of new files.
  2. Caution, 30–50%. History is accumulating and small shortcuts appear. Finish the current task and run the tests. Add no new scope.
  3. Hard stop, 50%+. Do not ask for new features. Commit the passing code to a branch, run /clear, open a fresh session, and reload only the clean specification and the specific files the next task needs.

The ritual matters more than the exact thresholds. A clean session and a cold reset are cheap. Debugging a corrupted one is not.


Common Misconceptions

❌ The Myth
A bigger context window means a smarter agent.
✅ The Reality
A large window helps retrieval, not reasoning. Past a point, more context means more noise, diluted attention, and a higher chance the model loses the one constraint that mattered.
❌ The Myth
If the model gets the domain wrong, prompt it harder.
✅ The Reality
You cannot prompt your private context into a model that has never seen it. When you are outside the training distribution, provide an explicit specification or an automated test the model can verify against.

Key Takeaways

✓What You Learned
  • ✓
    An LLM is a context-driven execution engine: the prompt is the instruction set and the context window is RAM.
  • ✓
    The agent is fast labor with no company context; the human owns architecture, identity, and security.
  • ✓
    Capability is jagged because reinforcement learning only compounds where an automated verifier exists.
  • ✓
    Attention degrades as the window fills, so long sessions get worse; a cold reset beats a compacted summary.
  • ✓
    Most of the traditional glue tier is now obsolete: check whether a foundation model can map input directly to output.
  • ✓
    Keep every task inside the Green Zone, finish in Caution, and hard-reset at 50% context.

Try It Yourself

Next time you open an agent, run these three experiments and watch the physics in action:

  1. The 50% drill. Commit early, run /clear, and restart with a written spec. Compare the diff quality against a session you let run past 50%.
  2. The invisible constraint. Tell the agent not to use a specific library, let the session run long, then check whether it still respects the rule. That is attention degrading in real time.
  3. The middle-tier audit. Take one service in your stack and ask whether a single model call could replace it. Do not build it yet — just count how many layers could disappear.

Up Next in the Series

💡
Series Roadmap

Part 2 — Choosing the Squad: Now that you understand the physics, the next problem is organization. We break down the five real-world agent topologies — linear pipelines, autonomous swarms, role-based squads, production state graphs, and visual canvases — and give you the decision tree for picking one.

MH

Mohamed Hamed

20 years building production systems — the last several deep in AI integration, LLMs, and full-stack architecture. I write what I've actually built and broken. If this was useful, the next one goes to LinkedIn first.

Follow on LinkedIn →