BIRKEY IA

ABOUT  RSS  PROJECTS  ARCHIVE


03 Oct 2026

Everything I Build Around the LLM Becomes a Distraction

I've been using LLMs for software development for a bit more than a year now. One thing I can say for sure: the landscape changes very fast.

There are a lot of trendy things that, at the time, seemed like a good idea. Over time they turned out to be more signal than the real entropy underneath. I want to come back to that in a separate post, because it deserves more than a few sentences.

The one thing that did not change

If there's a single takeaway, it's this: you have to actually use it, see it, and understand it. Understand what the benefit is, what the trade-offs are, what you're losing. And then you have to adapt to that.

Methodologies change. Strategies change. Every few months something new comes along and the best practice from last quarter looks naive. But underneath all that churn there is one thing I keep coming back to, again and again: the core of the problem, and the core mental model of understanding it, has always been the same.

Yes, the LLM helps. But if you don't use it carefully, it actually generates more content, more features, more code, more of everything, that you have no understanding whatsoever. Even when it seemingly works. Even when it's sort of tested and whatnot. There's no orthogonal consistency given the complexity of the problem we're solving.

What I want to do here is walk through the journey itself. What I actually did, what worked, what didn't, and where I ended up.

Confined problems

Early on, I mostly gave it a very confined problem. Refactoring a code base. Simplifying a certain module. And they did a pretty good job at it. The reason is simple: the problem is already solved. It's just a matter of refactoring, which is a well suited shape of task for an LLM. That was the early stage of LLM assisted development for me.

The context rabbit hole

Then, as the models got more powerful and started to get into it, I realized how critical the context you give an LLM is. I spent a lot of time refining it, trying to get as much relevant information as possible in there.

The way I did it was by experimenting. Starting with a simple grab search tool, whether it was syntactic or semantic, to pull more relevant information from the code base, or from the web, or from whatever literature is relevant. Getting the right context in front of the model. That's where we started, and the repository was the first step.

It was a mixed result overall. The lesson I took from it is that context engineering is a kind of dark art. Even refining the context engineering, you can get into a deep rabbit hole. The LLM keeps changing, and then a seemingly pretty solid methodology just degrades. The model updates underneath you. So you start a new eval to refine it. And then you keep refining. The more you refine, the more you realize the quality doesn't really improve in proportion to the time you put in.

The amount of time you spend on context engineering versus the benefit you actually get back doesn't justify it. Eventually what you end up doing is starting over. There's no way of prompting, no way of giving more context, that fixes it.

Context management and compaction

The next issue that hits you immediately is context management, meaning compaction and things like that. That also turned out not to be an easy problem to solve. There are many different methodologies you can try.

The easiest would be to just ask an LLM to summarize it and write a handoff. But then you have the risk of: how do we compact it, how do you summarize it? The more you work on it, in terms of summarizing after the context is full versus doing summarization after each successful request and response as a background process running async, the more you end up creating. You get into a rabbit hole of managing different processes and coordinations. At the end of the day, it's more about state continuity. And because of the stochastic nature of LLMs, there's just no way to consistently, deterministically find a better way to do summarization and handoff.

The best way I've actually been using right now is: no automatic handoff, no automatic summarization. You summarize manually. You ask a certain model for the conversation history so far, and then you verify it before handing it off.

Which seems to me that maybe in the future this will just be done automatically by the model provider on the server side. But on the client side, it seems to me that it's a drag on your development productivity. Including so-called agentic coding, coding agents. Because it distracts you from actually solving the problem, versus actually creating a non-deterministic handoff flow that you keep just changing, which is never ending. That's also not a good way of managing the context problem.

Understanding the tool you are using

The third thing is using coding agents that are developed by someone else. Even though you end up customizing it, you end up putting a lot of tools and plugins and things like that in it. Eventually it gets to such a complex state that it's just hard to understand the tool you're using.

For example, when I say the tools you're using, it's hard to understand what the current context is, what the current prompt is that the LLM is seeing, what the actual in-memory instruction set is, the in-memory context the coding agent is actually working out of. And yes, there are tools to inspect this and that, but it becomes a job of its own. It distracts you from the problem you're trying to solve.

That dictates that you start either using a very bare bone, very minimal coding agent, which you can then customize to your workflow. Or, in my case, even after experimenting with the Pi coding agent, which had so many different choices and plugins and kept evolving so fast, I eventually decided to write my own. With a particular constraint that I put on it, and the particular workflow I can actually wrap my head around.

Whether you use someone's coding agent or do it your own, the learning is: you really have to understand the tools you're using, and be able to master them, to take the best advantage of the coding models.

Skills

Another thing I've changed, another thing I've learned, is using skills. I'm still very on board with them, because they have this nice property: they furnish the context on demand, only when it's needed. When it's not needed, it's not going to be used. That's one of the really nice things about it. It's dynamically loaded, so it doesn't cover the context the whole time.

The tradeoff is that it's not a deterministic instruction set. It's more of an advisory to the LLM. The LLM more or less follows it, but there's no guarantee.

The nice part, though, is that you can take some ad hoc prompt you've been repeating over and over and turn it into a really well structured skill. And you get better results. You can express a lot, the majority of workflows that you have, with a skills file. I still use them. I don't know what the future holds for this, but so far it gives me more ROI than not. You can describe them in a very example based approach, and then most of the time, almost all of the time, LLMs are pretty good at following those instructions exactly, the way you'd like them to approach solving a problem.

Spec driven development

In terms of the other aspects, it's spec driven development. It starts out well, especially when you have a well-defined spec, which is in and of itself a complicated problem.

But as time goes by you generate too much documentation. Doc, doc, doc, doc. Because of the nature of specs, over time it drifts. It introduces contradictory, non-orthogonal ideas. Because you probably forgot the spec you wrote about a few days ago, let alone a month ago, and then the LLM starts to introduce something else. Eventually the documentation explosion is going to be a big problem.

What I've learned over time is that the spec should be sparingly used. A very high level description of how a specific module, or program, or feature works. But in terms of the detail, I rely more on executable code. That's the direction I landed on.

Scripts instead of prose

One of the other things that works really well for me is this. When there's a certain kind of repetitive thing I always ask it to do, instead of writing skills and whatnot, I ask LLM to create a script. An executable script. And then combine that with a skill that says: hey, this is the script you run to accomplish certain things, and then use the output to go on within the problem we're trying to solve.

Start with an executable prototype

And one last thing. After all of those experiences, one thing I've been having good results with is to always start with an executable prototype of the actual flow. An actual problem.

You can even start from the UI. You don't even have to start with the database. You don't even have to have all of the server running. It's just one executable program. Or, if you're doing more of a web app, a consumer facing app, it's more like starting with one HTML file with everything just embedded in there. That executable, it is code.

That's one thing I believe is helping me way more than anything else right now. It's still yet to be seen in terms of the results, but so far the iteration loop works like this. Say you have a flow, some kind of onboarding flow. You can ask the LLM to generate an onboarding flow based on whatever your requirement is. You can even simulate it. You can even let it test it. And the requirement is just one HTML file, everything is embedded, no framework, nothing.

That gives you a prototype you can actually iterate over until you say that this is a fleshed out flow. And then you can use that as a blueprint for the actual app. That has been paying pretty well for me.

I hope some of this rambling is useful.

Tags: AI engineering coding-agent