LLM Use Cases: Code Reviews

October 3, 2026  |  LLMs  ·  Programming  ·  BTC Map

Previous post

Crossing the Rubicon

BTC Map Android v1.2.0 shipped on September 30, roughly six months after v1.1.0. The previous release was 100% hand-written, and this one started the same way, with zero LLM involvement.

Early in the cycle I introduced OpenCode, starting with some models provided by PayPerQ and quickly switching to M2.7 Highspeed (2026-04-09) for better latency and reliability. It turned out to be counterproductive, but amusing.

Then came M3 (2026-06-01), which felt value neutral, more of a toy than a tool. Finally I tried DeepSeek V4.1 Flash (2026-09-10), and it crossed the Rubicon of actual, indisputable LLM usefulness on real tasks. By the time I cut the release, almost every recent change was LLM-accelerated.

Many new features were still hand-written early in the dev cycle. That notwithstanding, before cutting the release, I put every module, new and old, through an LLM code review and kept iterating until there were no findings of medium severity or higher left.

My Code Review Setup

Before doing anything else, I created a bash helper script for managing the emulator lifecycle. The goal was to make sure the LLM knew exactly what to do and would not waste time struggling with the dev environment.

My workflow was simple. I would point the model at one module, ask for a review with findings ranked by severity (low, medium, high, critical), fix every finding at medium severity or higher, then re-run the same review. Repeat until the module comes back clean, then move on to the next one. Some modules took one pass, others took a dozen.

In my experience, DeepSeek V4.1 Flash is a fantastic reviewer. What sold me most was its speed. It is fast, thorough, blunt, and cheap enough to re-run the whole app a gazillion times. That matters a lot, because quantity beats quality in code reviews. Fixing one issue reshapes the surrounding code and tends to expose the next one.

More Code Is a Liability

The obvious pushback is that most findings are edge cases nobody will ever hit, and every fix that adds code makes the codebase bigger and harder to maintain. There is truth to that. A reviewer, human or not, is rewarded for finding something, and “delete this and make it simpler” is a much rarer suggestion than “add one more null check for an impossible state”.

I did not fix everything. I routinely ignored low-severity findings on purpose, and I rejected my share of medium-severity findings too, precisely for the reason stated above. The severity threshold is what made my workflow net-positive. Without it, an eager clanker will happily bury a small app under defensive code, which is a net negative even if every individual suggestion is technically correct.

What the Reviews Actually Found

Some findings were genuinely useful, in ways that justify the effort.

  • Real bugs. Races, stale state, leaked resources, and mishandled empty or error responses. These are the ones that are hard to reproduce and easy to ship unnoticed.
  • Consistency. The same screen doing the same thing in three different ways, or error handling drifting between modules. Boring, but if you’re autistic enough, you’d appreciate the value.
  • Tests. A lot of findings were cheap to turn into regression tests instead of yet more defensive code, and the test suite grew noticeably because of it.
  • Architecture. Repeated reviews and revisions pushed the code toward less coupling and clearer separation of concerns, which is a benefit that no single finding captures.

That last point is the one I underrated. A reviewer that reads every module over and over starts to notice structural problems, not just local ones, and that nudges the architecture in a good direction.

Conclusion

The previous, fully manual release shipped six months ago. I am very happy with how 1.2.0 turned out, and it feels much more stable than v1.1.0.

This workflow is not free, and it is not magic. Reviewing is a loop with a real time cost, and the temptation to over-engineer is something you need discipline to suppress.

An LLM is also not a substitute for tests or for understanding your own code. It does not know your product and will cheerfully propose correct, useless changes. But as an extra pair of eyes that never gets bored, applied to a codebase you wrote and understand, code review is one of the better LLM use cases I have found so far.