Series · 14Tokyo · Incoming Cambridge HSPS
Five months of running an AI coding agent, measured.
14 posts, published on 17 September 2026. Each one starts from a number I counted in my own transcripts, logs or ledgers while running small businesses with one coding agent. Where a measurement contradicted something I had said, the post says so.
14Posts
68 minReading
2026-09-17Published
S-012026-09-17Prose rules did not bend the curveFive months of running a coding agent against a rulebook that grew to 2,159 rules and 54 hooks, and the measurement that says only the rules a guard refuses in the path ever stopped recurring.6 min · what it measured→S-022026-09-17Assume the instrument is wrong before the world isFive months of running twelve web properties and a coding agent taught me that a zero, a round number and a clean series are claims about the meter, and each one needs a second meter before it gets published.6 min · what it measured→S-032026-09-17A score the agent gives itself measures its detectorsI told my coding agent to punish itself for repeating corrected mistakes, it built a 0 to 100 score, and eleven days of the ledger's own data showed the score tracked its detectors rather than my dissatisfaction.6 min · what it measured→S-042026-09-17The complaint distribution has no headI logged 5,270 complaints about my coding agent over five months, found 2,720 distinct behaviours with no dominant one, and learned that per-behaviour guards cannot cover a tail that shape.6 min · what it measured→S-052026-09-17Well-formed and wrong: shape improved, relevance did notI spent two months fixing the format of my coding agent's replies, first-attempt compliance went from zero to 43 per cent, and over the same months the share of my next messages that were corrections tripled.5 min · what it measured→S-062026-09-17Seven hours a day with a coding agent, and never a usage limitI measured my own agent use from the transcripts: nearly eight hours on a weekday median, 4,175 distinct prompts of my own since July, no limit hit, and the reason is routing rather than restraint.6 min · what it measured→S-072026-09-17The sessions that feel hostile are the ones that compactAcross 21 days of coding-agent sessions on my machine, the ones I described as hostile and forgetful were the ones with the most context compactions and the most guard refusals, and both scale with how much work the session was doing.6 min · what it measured→S-082026-09-17A cap is a floor. Raise it and re-run.A number that lands exactly on a limit is measuring the limit, so raise the cap and run it again before anything downstream of it gets believed.3 min · what it measured→S-092026-09-17A 404 on a guessed URL is not evidenceMy agent invented four API paths, got four routing errors, told me the product had no API, and so missed a crawler setting that was wrong on twelve of twelve zones.3 min · what it measured→S-102026-09-17Merged is not deployed is not servingThree links stand between a commit and the page a customer loads, each one fails in its own way, and each one needs its own proof.3 min · what it measured→S-112026-09-17The gate refused its own author three times before it workedA guard my agent wrote to stop unplanned sprawl fired on its own author three times in an hour, and each refusal named a different thing I had assumed about how these controls read the world.3 min · what it measured→S-122026-09-17A filter cannot report what it excludedA date window compared in the wrong timezone made a family business thread read as empty, and my coding agent reported the empty result as a finding, which is why message stores are now read end to end with nothing filtered.3 min · what it measured→S-132026-09-17Fifty-nine per cent of the log was the tool talking to itselfSixteen reader agents went through five months of my coding-agent transcripts one window each, and found that 9,039 of 15,270 logged prompts were machine-fired, that the agent's own tooling ate most of the real ones, and that several of the rules I trusted had never fired.6 min · what it measured→S-142026-09-17Ten changes after five months with a coding agentSixteen reader agents went through five months of my transcripts, and the ten rules they produced each carry the count that forced it.6 min · what it measured→