Second Half Software — a rising sun
Second Half
Software LLC
A lifelong love for engineering.

From a chat window to a team of agents: notes from building an app

Hi, I’m Sam – the founder and sole human in Second Half Software. I was in the software industry for 25 years before I retired. I then took a few years off from software development, and have recently started again by exploring what AI can help with. I mention the 25 years only as context: writing software was familiar to me, and working this way was not.

What I am building is a mobile app for iPhone and Android with a small server behind it. It isn’t published yet, but I’ll update the blog when it is.

I’m building it with Claude, Anthropic’s AI model, and the way I work with it has changed five times. I’m writing the five stages down so that another developer can skip some of my mistakes. I’m not an expert at this. It’s one person’s record of one project, and the last stage is new enough that I can’t tell you whether it works.

1. Exploring the idea on the web

This was an early exploration. I’d heard about Claude but hadn’t used it much, so the first stage was just getting familiar with it: using it on the web to explore the idea.

At first I wasn’t sure which model to use. I quickly found that the latest model was the most helpful.

What I’d tell another developer: if you’re just starting, begin with the latest model. It was the most helpful one for me.

2. Claude Desktop, and my brief career as a clipboard

Next I installed Claude Desktop on my MacBook Pro and started learning what it could do: interact with the computer, drive the browser through its extension, and so on.

Then I started setting up an environment in the terminal. I was using Claude Code in the desktop app, but in a cloud project instead of a local one, so it couldn’t see my terminal. I found I was copying and pasting back and forth so that it could see the console. That’s what led to the next stage.

What I’d tell another developer: if you notice you’ve become the clipboard between two windows, take that as the signal to move the assistant to where the work is.

3. Claude Code in the terminal

So I installed Claude Code in the terminal, started working there and have been ever since. Claude runs inside the project: it reads the files, runs the commands and the tests, and sees the console itself, so nothing has to be pasted back and forth. It can even drive the browser for me through the Claude extension in Chrome.

For a sense of scale: the project’s first commit is dated August 19, 2026. By October 9, seven weeks later, it had a little over a thousand commits, and nearly all of them name Claude as co-author. The app is React Native, the server runs on Cloudflare Workers, and the UI is tested with Maestro on simulators and on real phones. A commit count measures activity, not progress, so please don’t be impressed by it.

The most useful thing I did in this stage came ten days in. I added a file called CLAUDE.md, which Claude Code reads automatically at the start of every session. Think of it as a working agreement between the two of us. Mine opens like this:

This is a starter draft — edit it down hard. It loads on every turn, so it’s a signal budget, not a wiki. Keep lines that change behavior; delete anything that reads as a platitude. Grow it from real friction: when you correct me the same way twice, that’s a line here.

Its first instruction is to keep it short. It’s now seven times longer. Neither of us listened.

But it did grow from friction, and that’s the part I’d recommend. It’s been changed in more than eighty commits, and most of its rules carry the story of the mistake that produced them. Here are four:

  • “Change the code base, not the line.” The classic failure on this project is a fix that’s correct where it was made and wrong everywhere else: a small change made without reading its neighbours.
  • “A diagnosis you didn’t watch fail is a guess.” Two notes explained the same test failure in two different, equally confident ways. Both were wrong, because both came from reading the test and not from running it. The real cause took one screenshot.
  • “Assert the thing under test, not a proxy for it.” A test that looked for a word on the screen passed on the wrong screen, because the same word was sitting in a progress bar on every step.
  • “A bug is a class, not an instance.” Every fix ends with a search for the same mistake elsewhere, and the answer goes in the commit message even when it’s “checked, no others”.

Two habits sit beside that file. Design decisions go into documents in the repository, about three dozen of them now, and CLAUDE.md points at them and doesn’t repeat them. And Claude keeps short dated notes between sessions, because each session starts with no memory of the last one.

What I’d tell another developer: create a CLAUDE.md, and grow it from mistakes that actually happened, with the mistake written next to the rule. Then try harder than I did to keep it short.

4. Several agents at once

Early in October, I added a rule called “Fan out, then coordinate.” A test run had come back with ten failures. They were diagnosed as five clusters at the same time, each by its own agent, and not one after another.

I call this stage semi-structured because nothing about it was designed up front. These are the rules that emerged:

  • Split the work by independent piece.
  • Never two hands in one file. Each agent works in its own copy of the repository, or is given a named set of files that is its alone.
  • Every change gets an adversarial review by an agent that didn’t write it.
  • A reviewer must demonstrate a finding, or mark it as reasoned and not shown.
  • A reviewer removes the key lines of a fix one at a time, to see which tests go red.

Here’s one day of it. On October 9 I found a serious sign-in bug on my own phone that no test had caught. One agent diagnosed it from the app’s analytics and the code. A second wrote the fix in its own copy of the repository. Then came four review rounds, each by a fresh agent, and each found something real that the round before had missed. In one, the fix itself had introduced a loop of requests. In another, the fix had quietly opened a path I’d decided against that same morning.

And how did the bug escape in the first place? My test phone had no passcode. My real phone does. Guess which one found the bug.

The thing to know about agents is that they report in the same confident voice whether they’re right or wrong. I have worked with people like this. That same day, Claude’s first diagnosis of a small visual bug was wrong, and the agent that went and looked properly corrected it. Another agent refused to edit CLAUDE.md on a different agent’s say-so, which was exactly the right call.

Working this way felt like a big jump in how fast the engineering moved. I haven’t measured it, so take that as the view from my chair. It also came naturally, because I spent years managing teams, and I really like working this way.

What I’d tell another developer: have every change reviewed by an agent that didn’t write it, and ask the reviewer to show you the failure, not describe it.

5. In progress: a team of agents

This stage started on October 9, the day I’m writing this, so it’s a design and not a result.

As I’ve grown used to working with agents, I found myself curious about their roles and how they work. So I asked Claude what agents were operating on the project. Claude responded and it got me thinking: this is just like a software team except without formal roles (and I’m deeply familiar with those formal roles).

So I proposed ten roles and asked Claude for its opinion. Claude refined the list: two renames, a release role, a support role, and privacy added to security. We ended with twelve: Project Lead (the main assistant, which coordinates and brings decisions to me), Developer, Code Reviewer, Architect, Tester, Release Engineer, Security and Privacy Expert, UI Designer, Product Manager, Marketer, Document Writer and Support Specialist. The Code Reviewer runs on a different model from the Developer, so that the two have different blind spots.

The protocols matter more than the names:

  • A role is a contract: what it reads first, what it may touch, what it must hand back, and which tools it has.
  • Every role starts clean each time. It learns through notes kept in the repository. A role proposes a note, the lead writes it, and each note is dated and marked as watched, measured or reasoned.
  • Collaboration runs through the lead, and few roles change code.
  • The lead runs as much of the work in parallel as it can.
  • After every release to testers, the team holds a blameless retrospective from a log of the round, with the same few numbers each time and at most three improvements, each checked at the next round.

All of this is new and unproven. Nothing has been measured yet. But this seems like a natural extension of how to work with Claude.

What I’d tell another developer: Nothing yet about the agent team. Ask me after a few rounds :). But what I do know is that working with Claude is an absolute delight and it works best when you approach it as a partner. Get curious. Ask it questions. Learn from it, challenge it. Ask it to challenge you.

What I don’t know yet

I don’t know whether the team is better than the looser arrangement it replaces. I don’t know whether notes in the repository will keep a role from repeating a mistake, or whether a retrospective’s improvements will survive to the round after.

Once the team has some mileage I’ll write a second post that says what worked and what didn’t. If I’ve missed something here, I’d like to hear it: [email protected].