Everything I Build Around the LLM Becomes a Distraction
I've been using LLMs for software development for a bit more than a year now. One thing I can say for sure: the landscape changes very fast.
There are a lot of trendy things that, at the time, seemed like a good idea. Over time they turned out to be more signal than the real entropy underneath. I want to come back to that in a separate post, because it deserves more than a few sentences.
The one thing that did not change
If there's a single takeaway, it's this: you have to actually use it, see it, and understand it. Understand what the benefit is, what the trade-offs are, what you're losing. And then you have to adapt to that.
Methodologies change. Strategies change. Every few months something new comes along and the best practice from last quarter looks naive. But underneath all that churn there is one thing I keep coming back to, again and again: the core of the problem, and the core mental model of understanding it, has always been the same.
Yes, the LLM helps. But if you don't use it carefully, it actually generates more content, more features, more code, more of everything, that you have no understanding whatsoever. Even when it seemingly works. Even when it's sort of tested and whatnot. There's no orthogonal consistency given the complexity of the problem we're solving.
What I want to do here is walk through the journey itself. What I actually did, what worked, what didn't, and where I ended up.
Confined problems
Early on, I mostly gave it a very confined problem. Refactoring a code base. Simplifying a certain module. And they did a pretty good job at it. The reason is simple: the problem is already solved. It's just a matter of refactoring, which is a well suited shape of task for an LLM. That was the early stage of LLM assisted development for me.
The context rabbit hole
Then, as the models got more powerful and started to get into it, I realized how critical the context you give an LLM is. I spent a lot of time refining it, trying to get as much relevant information as possible in there.
The way I did it was by experimenting. Starting with a simple grab search tool, whether it was syntactic or semantic, to pull more relevant information from the code base, or from the web, or from whatever literature is relevant. Getting the right context in front of the model. That's where we started, and the repository was the first step.
It was a mixed result overall. The lesson I took from it is that context engineering is a kind of dark art. Even refining the context engineering, you can get into a deep rabbit hole. The LLM keeps changing, and then a seemingly pretty solid methodology just degrades. The model updates underneath you. So you start a new eval to refine it. And then you keep refining. The more you refine, the more you realize the quality doesn't really improve in proportion to the time you put in.
The amount of time you spend on context engineering versus the benefit you actually get back doesn't justify it. Eventually what you end up doing is starting over. There's no way of prompting, no way of giving more context, that fixes it.
Context management and compaction
The next issue that hits you immediately is context management, meaning compaction and things like that. That also turned out not to be an easy problem to solve. There are many different methodologies you can try.
The easiest would be to just ask an LLM to summarize it and write a handoff. But then you have the risk of: how do we compact it, how do you summarize it? The more you work on it, in terms of summarizing after the context is full versus doing summarization after each successful request and response as a background process running async, the more you end up creating. You get into a rabbit hole of managing different processes and coordinations. At the end of the day, it's more about state continuity. And because of the stochastic nature of LLMs, there's just no way to consistently, deterministically find a better way to do summarization and handoff.
The best way I've actually been using right now is: no automatic handoff, no automatic summarization. You summarize manually. You ask a certain model for the conversation history so far, and then you verify it before handing it off.
Which seems to me that maybe in the future this will just be done automatically by the model provider on the server side. But on the client side, it seems to me that it's a drag on your development productivity. Including so-called agentic coding, coding agents. Because it distracts you from actually solving the problem, versus actually creating a non-deterministic handoff flow that you keep just changing, which is never ending. That's also not a good way of managing the context problem.
Understanding the tool you are using
The third thing is using coding agents that are developed by someone else. Even though you end up customizing it, you end up putting a lot of tools and plugins and things like that in it. Eventually it gets to such a complex state that it's just hard to understand the tool you're using.
For example, when I say the tools you're using, it's hard to understand what the current context is, what the current prompt is that the LLM is seeing, what the actual in-memory instruction set is, the in-memory context the coding agent is actually working out of. And yes, there are tools to inspect this and that, but it becomes a job of its own. It distracts you from the problem you're trying to solve.
That dictates that you start either using a very bare bone, very minimal coding agent, which you can then customize to your workflow. Or, in my case, even after experimenting with the Pi coding agent, which had so many different choices and plugins and kept evolving so fast, I eventually decided to write my own. With a particular constraint that I put on it, and the particular workflow I can actually wrap my head around.
Whether you use someone's coding agent or do it your own, the learning is: you really have to understand the tools you're using, and be able to master them, to take the best advantage of the coding models.
Skills
Another thing I've changed, another thing I've learned, is using skills. I'm still very on board with them, because they have this nice property: they furnish the context on demand, only when it's needed. When it's not needed, it's not going to be used. That's one of the really nice things about it. It's dynamically loaded, so it doesn't cover the context the whole time.
The tradeoff is that it's not a deterministic instruction set. It's more of an advisory to the LLM. The LLM more or less follows it, but there's no guarantee.
The nice part, though, is that you can take some ad hoc prompt you've been repeating over and over and turn it into a really well structured skill. And you get better results. You can express a lot, the majority of workflows that you have, with a skills file. I still use them. I don't know what the future holds for this, but so far it gives me more ROI than not. You can describe them in a very example based approach, and then most of the time, almost all of the time, LLMs are pretty good at following those instructions exactly, the way you'd like them to approach solving a problem.
Spec driven development
In terms of the other aspects, it's spec driven development. It starts out well, especially when you have a well-defined spec, which is in and of itself a complicated problem.
But as time goes by you generate too much documentation. Doc, doc, doc, doc. Because of the nature of specs, over time it drifts. It introduces contradictory, non-orthogonal ideas. Because you probably forgot the spec you wrote about a few days ago, let alone a month ago, and then the LLM starts to introduce something else. Eventually the documentation explosion is going to be a big problem.
What I've learned over time is that the spec should be sparingly used. A very high level description of how a specific module, or program, or feature works. But in terms of the detail, I rely more on executable code. That's the direction I landed on.
Scripts instead of prose
One of the other things that works really well for me is this. When there's a certain kind of repetitive thing I always ask it to do, instead of writing skills and whatnot, I ask LLM to create a script. An executable script. And then combine that with a skill that says: hey, this is the script you run to accomplish certain things, and then use the output to go on within the problem we're trying to solve.
Start with an executable prototype
And one last thing. After all of those experiences, one thing I've been having good results with is to always start with an executable prototype of the actual flow. An actual problem.
You can even start from the UI. You don't even have to start with the database. You don't even have to have all of the server running. It's just one executable program. Or, if you're doing more of a web app, a consumer facing app, it's more like starting with one HTML file with everything just embedded in there. That executable, it is code.
That's one thing I believe is helping me way more than anything else right now. It's still yet to be seen in terms of the results, but so far the iteration loop works like this. Say you have a flow, some kind of onboarding flow. You can ask the LLM to generate an onboarding flow based on whatever your requirement is. You can even simulate it. You can even let it test it. And the requirement is just one HTML file, everything is embedded, no framework, nothing.
That gives you a prototype you can actually iterate over until you say that this is a fleshed out flow. And then you can use that as a blueprint for the actual app. That has been paying pretty well for me.
I hope some of this rambling is useful.
GitHub's Recent Crisis Has a Simple Fix
Introduction: The Outages Aren't Bugs; They're Symptoms
GitHub is down again? Oh boy, it has a scalability issue, you might say. I say your coding agent may be to blame, and thus may be you should take responsibility? The platform is buckling under a weight it wasn't designed to carry. We've seen downtime, API timeouts, and "maintenance mode" messages since its inception, but we have not seen them as frequently as we have since November 2025. (Does this month ring a bell? It is when Anthropic released Claude Opus 4.5, which wowed all of us with its complex code generation capabilities.)
The mainstream narrative blames "infrastructure scaling" or "unexpected coding party traffic spikes." But the real culprit is staring us in the face: the explosion of agentic AI.
We are no longer just seeing human-paced code commits. We are seeing AI agents generating thousands of API calls per second—opening PRs, commenting on issues, and pushing commits at machine speed. GitHub, the world's primary collaboration tool, is being swamped by a flood of low-value, automated noise.
The "AI Slop" Problem
Before we talk about engineering, we need to talk about quality. "AI slop" isn't just a buzzword; it's a systemic degradation of our codebases and our learning processes.
Agentic AI makes it too easy to generate working demos. It's incredibly easy to spin up a project that looks impressive in a 30-second TikTok video. But when that project hits a real codebase, it becomes unmaintainable vaporware.
The internet is clogging up with boilerplate PRs, automated dependency updates, and "AI-written" refactorings that no human actually reviewed. We are seeing a shift from writing software to reviewing AI output, and the latter is a fundamentally passive, low-value activity.
The Technical Fix: Intelligent Backoff Queues
So, how do we fix GitHub's backend crisis?
A naive solution is simple rate limiting: "You can only make 100 calls per hour." But smart AI agents will just hit that limit faster or spin up more parallel threads to compensate. Throttling a dumb agent just creates a "thundering herd" effect.
The real solution is organization-level flow control via hard queueing.
GitHub needs to implement an intelligent queue system per organization. Every API call (PR creation, issue comment, commit push) should go into a queue that enforces a thinking delay.
- Burst Limits: "Your team just pushed 50 commits. The next 50 must wait 10 seconds."
- Exponential Backoff: "You've generated 5 PRs in a minute. The next one must wait 10 minutes."
This isn't just about server load; it's about forcing a pause.
When an AI agent tries to "slop" a project, the queue forces it to sit on its hands. It breaks the agentic loop. It turns a 2-second generation cycle into an overnight batch process. It forces developers to check the logs, read the diffs, and actually collaborate before the next burst of commits is allowed through.
The Philosophical Win: Forcing the "Human Pause"
Here is the counterintuitive truth: We need GitHub to be slower.
When you write code manually, you understand the architecture. You feel the pain points. You learn the library. When AI writes it for you, you just review it. Reviewing is passive; writing is active.
By enforcing hard queues and flow control, GitHub can become the "Great Filter." It can force a friction point between generating code and shipping code.
- Real Software vs. Demos: Vaporware thrives on speed. Real, maintainable software thrives on deliberate thought. Slowing down the API forces the agent (and its human overseer) to think before they commit.
- Learning vs. Generating: If the platform forces you to wait 10 minutes between bursts, you might as well use that time to read the documentation. You might as well write the next function yourself to see where the AI got stuck.
Conclusion: The Library Guardian Metaphor
Let us face it. GitHub is the world's source control. It is the Library of Alexandria for the software age. It is supposed to be the place where the best ideas in code are preserved, reviewed, and improved.
When AI floods the platform with generative noise, we aren't just filling servers; we are diluting the signal. We are turning a library of human knowledge into a warehouse of automated spam.
We need GitHub to be the library guard. The one who looks at the flood of AI-generated PRs and says:
"Slow down. Read this first. Understand what you're doing. And for the love of code, don't push until you've actually thought about it."
By forcing a slowdown, GitHub doesn't just save its infrastructure. It saves the soul of software development. It forces us to stop slopping and start building maintainable, working software. Maybe GitHub will not do that, but someone should. If they do, I will migrate.
OneLoop: Local First Coding agent written in Rust with nix tooling
I have been working on my own personal coding agent for the last few months. As a matter of fact, my initial commit is on 18th of April 2026 and it went through quite a number of iterations. The git history is quite telling about my experimentation of features such as auto compacting, memory, and multi-model debate/consensus/judgment etc. It supported a few big name model providers on top of OpenRouter initially. After using it for real production work for some time, I am quite happy with the following state of OneLoop [Projects]:
- It defaults to local model (llama.cpp serving Qwen3.6-35B-A3B-Q4KM.gguf).
- It features 4 tools, on demand skill loading, append only session history with an option to reset via /clear command anytime.
- It has been battle tested, which includes developing OneLoop using OneLoop.
- It supports OpenRouter with web/link tools using /model command
Please note that this coding agent is designed by me for me and I do not recommend you use it for production projects since it lacks any security/permission guardrails. However, I'd recommend you clone the repo and just have fun hacking it. It is super easy to get going if you happen to use Nix on a machine at least with 32GB of ram:
git clone git@github.com:oneness/OneLoop.git ./ols # builds llama.cpp for your CPU/GPU, downloads Qwen model ./ol # run in another terminal and start interacting with the agent echo "Give me a mental model how this code base works" | ./ol
I'd like to share a few things I learned along the way building a coding agent tailored to my own workflow. First let me start with things to watch out when you are building your own coding agent, which I highly recommend.
- Keep your coding agent as dumb/minimal/simple as possible.
- Prompt engineering requires you to have deep knowledge about the problem domain. So optimize coding agent to teach you using research/learning skills that it can load on demand to instruct agent to help you get clarity.
- Context management is hard. I am thinking on this. Please reach out if you have novel ideas on how to do this better. I do have a few ideas but not settled yet.
- When you are in doubt of a feature you want to add, follow point 1.
- Use local model if possible. It forces you to slow down, be more cognizant of ever-growing context so you can breakdown problems for better solution.
Now, let me talk about how I use OneLoop. Following laws keeps myself and coding agent in check. I retrofitted Asimov's Laws of Robotics [Wikipedia link] as follows:
- A Coding Agent is not allowed to run in auto pilot.
- A Coding Agent must be made to obey the SOLID engineering principles.
- A Coding Agent is allowed to have side effects as long as it does not violate the first and the second law.
Ok, you might say that above is too vague or hand-wavy but trust me, they have served me and OneLoop well. It gave me clarity when I am developing OneLoop in terms of adding or removing features. Here are basic steps I utilize OneLoop for my day-to-day work now.
- Use OneLoop to get clear understanding of the problem I am trying to solve. This means that I ask OneLoop things like this: "Give me a hello world example of how x y z works". I ask it to generate one off scripts that I can run myself to exercise the code. I change here and there, add print statement and run it to get an intuition of how things work.
- Then I ask OneLoop to read codes from my repo and give me three options on how to add or integrate that feature.
- I review all options, ask further questions, ask OneLoop to generate one off standalone scripts so I can run them myself and iterate again from 1.
That is pretty much how I use OneLoop. It supports all I need such as clearing context to start new, summaries of long sessions so I can decide to use or not use in my ongoing conversation, switch to different model if I need (when I am running OneLoop from some low end pocket laptop) etc.
Remember, Coding Agent should serve you, optimize for your understanding, and helping you get better at your craft. If you have been just pressing `Enter` so far, please do yourself a favor and ask your coding agent to ban that command until you confirm that you understand what you are doing. Have fun coding!
Engineered, Not Vibe Coded
People call coding with AI "vibe coding". Sloppy, they say. Not real engineering. I am not interested in having that debate. Whether it is good or bad in the abstract is not an interesting question to me.
What interests me is what the current advances in AI, especially coding agents, actually enable: rigorous engineering. And by rigorous I do not mean lines of code. Lines of code matter, but they are not the point. I mean being very intentional about the feature set you want, starting from a clear spec and a clear direction, and keeping things as simple as possible so you get to the core of the problem you are solving.
I have a concrete case: a few months ago I started writing VoiceToText1, and I used Claude Code to write it. It has been my daily driver ever since. Whatever we end up calling this way of working, the result is a fairly robust tool that I rely on every day. I was in the driver's seat the whole time.
Why I built my own
Before writing my own, I used Wispr Flow and then the Willow Voice subscription service. Both work fine. I have nothing against either of them functionally. But the frequent updates kept making them more clever in ways I really did not want, and the feature set never quite met my needs. So I started writing my own.
A daily driver
I use VoiceToText for far more than I expected when I started. I compose messages and emails with it. I talk to coding agents with it. I dictate blog post drafts with it.
In fact, this very post started as a dictation. I spoke my raw thoughts into VoiceToText, read the transcript, and proofread it over and over. Then I asked Claude Code to clean it up, with strict instructions: make it idiomatic, give it structure, and do not generate anything I did not say. Which is pretty much the same way I drove it to write the code in the first place.
What kind of vibe is this?
I think we have been abusing the word "vibe". It is short for vibration, from the Latin vibrare, to shake. Somewhere along the way it came to mean coding by feel: wave your hands at the model, accept whatever comes back, move on.
That is not what I did. This tool was spec'd and engineered, with me guiding every step of the way. I chose the features. I wrote the spec. I set the direction. I rejected code that did not fit the structure and the idiomatic style I wanted, and I drove the AI until it did. None of it was autopilot. I was intentional about the specs, disciplined about the structure, and disciplined about the tests.
What drives a process like that is not a vibe. It is fundamentals, solid understanding, and long years of experience. The engineer is the one driving: setting the direction and managing the entire process. Call it vibe engineering if you like. Whatever the word, the distinction that matters is who is driving.
So is it sloppy? Is it vibe coded? I do not know what those words mean anymore. What I know is that I used a coding agent to build a robust tool that meets my needs exactly, with just the features that I want and none of the distraction or cleverness that I did not ask for. That, to me, is intentional planning and engineering.
Along the way I also used OneLoop, the small coding agent I built myself2. How I used it here is a separate post, and I will write about it at some point as a way of showing how I use coding agents in real work.
Footnotes:
VoiceToText: press a hotkey, speak, and the text appears where you are typing. Overview: https://www.birkey.co/VoiceToText/, source: https://github.com/oneness/VoiceToText.
I wrote about building OneLoop and what it taught me in The Agent Is Not the Point.
Ten Books That Shaped How I Think
I was recently asked for book recommendations and realized I have never written down the ones that actually changed how I think. Not the ones that taught me a language or a framework. The ones that shifted something in me as an engineer. Here they are, in no particular order.
1. A Philosophy of Software Design — John Ousterhout
Short, opinionated, and right about most things. The core idea is Deep Modules: hide complexity behind simple interfaces. This is ETC from another angle. If your module has a complex surface, every consumer pays that tax forever. Make the surface small and the internals can evolve freely. I read it in a weekend. I have been thinking about it for years.
2. Thinking in Systems: A Primer — Donella Meadows
I already thought in systems before reading this. What Meadows gave me was the vocabulary. Stocks and flows. Feedback loops. Leverage points. The places to intervene in a system, ranked from least to most effective. When I write about NixOS or Emacs or OneLoop, I am reaching for these ideas even when I do not name them directly. Short and elegant.
3. The Art of Doing Science and Engineering — Richard Hamming
Hamming spent decades at Bell Labs studying what separates great scientists from merely good ones. His central question: why do so few people make significant contributions while so many are forgotten? The answer is mostly about how you allocate your attention and whether you have the courage to work on important problems instead of safe ones. This maps directly to what I wrote about transmittable vs. untransmittable knowledge. Hamming is rigorous, contrarian, and deeply practical.
4. Designing Data-Intensive Applications — Martin Kleppmann
The best systems engineering book of the last decade. Kleppmann builds from simple primitives (logs, hashes, SSTables) all the way up to distributed systems, always grounding abstractions in the physical realities they paper over. Even if you never build a distributed database, the way Kleppmann thinks will change how you design any system. This book is first principles applied to data.
5. Programmers at Work — Susan Lammers
Interviews with 19 legendary programmers from the 1980s. Bill Gates, Andy Hertzfeld, Charles Simonyi, and others. What makes it timeless is the focus on how they think about problems, not what they built. You start to see patterns across how these minds approach complexity.
6. The UNIX Programming Environment — Kernighan & Pike
I live the Unix philosophy every day. Reading it from the source is
different than absorbing it osmotically. Kernighan and Pike do not
just describe the philosophy. They demonstrate it, building real
tools from shell pipelines, showing how composability creates power
that no monolith can match. Every time I write curl | jq or compose
tools in Emacs, I am living this book. My OneLoop agent (read, write,
edit, bash as four primitives) is an example of Unix way of doing things.
7. Structure and Interpretation of Computer Programs — Abelson & Sussman
The spiritual ancestor of everything I love about Lisps. Homoiconicity. Functions as first-class citizens. Building abstractions upward from the smallest possible primitives. SICP teaches you to think about computation itself, not any particular language. If you are a Clojure programmer and you have not read it cover to cover, fix that.
8. The Design of Everyday Things — Don Norman
Not a software book, and that is the point. Norman's central idea is that good design is invisible and bad design blames the user. This applies to everything from Makefile help targets to agent tool interfaces. Your self-documenting system? Your one obvious path? Your inspectability obsession? Norman gives you the design vocabulary to articulate why something feels wrong and how to fix it.
9. Computer Power and Human Reason — Joseph Weizenbaum
Weizenbaum wrote ELIZA in 1966, then spent the rest of his career warning about delegating human judgment to machines. He is not anti-computer. He is pro-human. He argues that some things should not be computed, not because they cannot be, but because doing so corrupts the thinking. Reading this in 2026, in the middle of the AI hype cycle, feels like finding a kindred spirit from fifty years ago. This is the philosophical depth behind "the agent is not the point."
10. The Mythical Man-Month — Frederick Brooks
The original essay on why adding people to a late project makes it later. But it is really about something deeper: the irreducible complexity of communication and the necessity of conceptual integrity in software. Brooks' concept of conceptual integrity (a system should reflect one set of design ideas, not a committee's compromises) is my oneness principle by another name. The anniversary edition includes "No Silver Bullet," which is even more relevant in the age of "AI will solve everything."
—
These are not the only books I have read. They are the ones I keep coming back to. The ones where I will catch myself mid-conversation reaching for an idea and realizing it came from one of these pages. If you have read them, you know what I mean. If you have not, start with 1, 2, and 9.