All posts
Product Thinking

I Let My AI Agents Ship for Six Months. I Still Couldn't Say What I Did.

Share

Somewhere in the last year I stopped being the one who wrote most of the code. Cursor writes the first draft. Claude Code finishes the refactor while I'm in a meeting. A third agent runs in CI and opens PRs I sometimes review the same day, sometimes three days later.

This is, by every metric I'm supposed to care about, a win. I ship more. The backlog moves faster than it ever did when I typed every line myself.

But ask me what I did last month and watch me stall.

Not because nothing happened. Because too much did, and none of it left the kind of trace I used to leave on myself. When I wrote the code, writing it was the record. I remembered the migration because I fought with it for two days. Now the agent fought with it, merged it in eleven minutes, and moved to the next ticket before I'd finished my coffee. The work is real. My memory of doing it is not, because I didn't.

Multiply that by four tools. Cursor knows what it did in Cursor. Claude Code knows what it did in its own sessions. Whatever runs in the CI pipeline knows nothing except “merged.” None of them talk to each other, and none of them talk to the promotion rubric that still, somehow, asks me to describe my technical impact this half.

So I started doing the thing I did with brag docs: paste the PR list into an LLM at the end of the month and ask it to remind me what I'd been doing. Which is a strange sentence. I was asking a machine to help me remember work that a machine did.

That's when it stopped feeling like a productivity story and started feeling like an accountability gap.

A coworker put it better than I did. She runs three agents on a good day and told me, half joking, “I trust them more than I trust my own memory of Tuesday.” She wasn't joking. Performance review season came and she had a git log with two hundred commits and genuinely could not tell her manager, in a sentence, what any five of them meant for the team.

Here's the part that actually worries me. If an agent can open a PR, it can also claim credit for one it barely touched, or quietly skip the review step, or describe a change it never actually verified runs. Nobody is watching that gap because there's no ledger to check it against. We accepted the output speed and forgot to ask for the paper trail that used to come free when a human did the typing.

The work isn't disappearing. The evidence is. And evidence is the only thing a promotion committee, a manager, or future-you actually cares about, months later, when the details are gone and only the artifact remains.

So we're building the ledger to go with it.

Perfolio already has an MCP server your coding agent can talk to directly. When an agent finishes real work, it can log a session straight into your review inbox, tied to the actual commit, PR, or file it touched, with a human role attached: did you direct it, review it, pair with it, or did it work solo. Nothing gets in on a claim alone. The server refuses a write unless it names a real artifact, and every session still lands behind the same accept or reject gate as anything you typed yourself. The agent can move fast. It still can't invent your record.

The point isn't to slow the agents down. It's to stop losing the thread the moment they start doing the work for you. You should be able to ship faster than ever and still walk into a review with a specific answer, not a shrug and a git log nobody has time to read.

If your agents are shipping more than you can remember, you don't have a productivity problem. You have a ledger problem.


Share

Ready to stop rebuilding your case from memory?

Get Started with Perfolio