Language models write most of the code in my projects now. I describe a change, an agent writes it, runs it and opens a pull request. Ask the same question twice and you get two different answers. Both arrive in the same confident tone.
In 2021 Emily Bender, Timnit Gebru and their co-authors called these systems stochastic parrots: models that string together plausible language, sampled from a probability distribution, with no grounding in whether the result is true. The models have become far more capable since then. The sampling is still there. Every token is a draw.
That leaves a practical question. If the thing writing my code is non-deterministic, where does my trust come from?
My answer after a year of working this way: I trust the checks around the model, and those checks have to be deterministic. Same input, same verdict, every time, whoever or whatever wrote the code.
Green is a claim
First, some failures. People caused all of these. Agents produce the same kind of failure, faster and in larger numbers.
The blog that disappeared. The posts on this site come from an API. The loader wrapped that call in catch { return [] }. When the API started rejecting the site's token, the blog page rendered its "no posts yet" state and every post URL returned a 404. Build, tests and monitoring all stayed green. An empty page looks deliberate.
The release that shipped old code. A deploy went through with the right image tag and healthy pods. The backend inside that image had been built seven commits behind the tag it carried. Every signal I checked said the release was live. Seven commits of it were missing.
Five weeks of backups of the wrong database. The nightly backup job exited with code 0 every night. It was dumping a database that had stopped receiving writes after a migration. The exit code answered whether the dump command ran. I wanted to know whether the dump held this morning's data.
A type checker that never ran. The API is bundled with esbuild, which strips TypeScript types without checking them. Build, lint and tests passed while about 1,370 type errors sat in the tree, and new ones shipped with each release.
Each of these is a green check answering a narrower question than the one I thought I had asked. An agent has the same blind spot, and it will tell you, fluently, that the work is done.
Gates the model cannot talk its way past
You can tell an agent "never push without running the tests". That sentence goes into a context window and becomes one input among thousands to a sampler. Most of the time it holds. "Most of the time" is a probability.
So the rule lives in a pre-push hook. Every push runs build, lint, typecheck and tests, and a failure blocks the push. The agent instructions still say never to use --no-verify. That instruction is the polite version of the rule. The hook is the enforced version. It gives the same answer to me, to an agent, and to an agent at three in the morning that has decided the failing test is unrelated to its change.
The general shape: every rule you care about should exist as a program that returns a verdict. Prompts and guidelines describe intent. Programs enforce it.
Shift left, into the author's own loop
The earlier a check runs, the cheaper the fix. Developers have known this for decades: a red squiggle in your editor costs seconds, and the same bug found by a colleague in review costs a conversation. Agents make it sharper. A test that fails on the agent's machine, before the push, lands in its context as plain output: file, line, expected, actual. The agent reads it, tries again, and I never see the mistake. The same failure found in CI costs a round trip. Found in staging, it costs an investigation. Found by a user, it costs trust.
The checks are the same for both kinds of author. The hook that stops an agent stops me too, and an error message that tells an agent how to fix something tells the next developer the same thing. So the checks move as close to the edit as they can. Type errors, lint and unit tests run before a push. Tests sit next to the code they cover, so whoever changes a file, developer or agent, finds its test in the same directory. Rules that used to live in review comments, like "don't swallow this error" or "log through the structured logger", became CI checks with an error message that says what to do. A check that explains its own fix can be handled by whoever tripped it, without waiting for a reviewer.
Shift-left has a limit: a pre-push hook sees the code, and real users still take paths no test describes.
A ratchet on the debt
The type checker could not become a gate overnight, because gating on zero errors meant weeks of cleanup first. So it became a ratchet. A committed baseline file records how many errors each file may have. A file that gains errors fails. A clean file that starts producing errors fails. A file with fewer errors than its allowance also fails, until someone commits the lower number.
Without that third rule, fixing five errors leaves room for five new ones and CI stays green either way. With it, every improvement becomes the new floor. The debt is frozen, visible, and can only go down.
Type errors were the first. The same shape now covers most of the numbers I care about, across my repositories: file length, function length, cyclomatic complexity, duplicated code, lint findings and unit test coverage. Coverage runs the other way, so its floor can only rise. Every threshold lives in a committed file. In SpanBarn the gate refuses to run when a threshold is set in the environment, so COMPLEXITY=99 make quality-gate fails, and the only way to move a number is a diff someone reviews. Each gate also runs its own tests first, because a ratchet that has quietly stopped failing looks exactly like a pass.
Ratchets suit agents well. They turn a vague goal like "improve the types" into a comparison against a number in the repository, and the diff of the baseline shows exactly which files got better or worse.
Verify the artifact
After the stale release, the deploy check changed: it now looks inside the running bundle for something the release introduced. After the backup incident, the backup check counts rows in the dump and compares them with the live database. When two ingresses claimed the same hostname and the router picked one per request, a plain curl against the public URL became a coin flip, so deploy checks now talk to the pod or the service directly.
The pattern repeats. Measure the thing you care about, at the place where it runs. A status code, a tag name or an exit code is a proxy, and proxies drift.
Make failure loud and stable
Every error path in my projects logs at error level through a structured logger and reports to BugBarn, the error tracker I run myself. The return [] in the blog loader now reports the failure before it falls back.
Loud is half of it. The other half is stable. BugBarn groups errors by a fingerprint that includes the exception message. One bug in a content pipeline put per-item data in its error message and produced forty separate issues, each of which looked minor. With a constant message and the details on the error object, the same bug becomes one issue with a climbing count. Deterministic grouping lets a person, or an agent reading the issue list, see that a single problem is getting worse.
A readout of the real flow
Tests describe the paths I thought of. Users take the others. When a flow breaks for one user in fifty, the report usually reads like "the form sometimes hangs", and that sentence gives an agent as little to go on as it gives me.
Telemetry turns that sentence into a record. The API runs OpenTelemetry. HTTP, Fastify and Postgres are instrumented automatically, and a small withSpan helper wraps the business steps the automatic instrumentation misses, such as search queries, auth calls and file uploads. A failing span carries the error and its stack. The spans go to SpanBarn, my trace store. On the frontend, FunnelBarn records sessions, so I can see what a user clicked and what they saw.
When a flow misbehaves, I open the trace. It shows every step of that one request in order, with timings, and marks the step that failed. The slow search query, the auth call that retried three times, the upload that timed out after the user had already left: they all sit in one waterfall. The guessing stops.
A real user flow is about as non-deterministic as anything gets: timing, network, data, a double click. The trace is a fixed record of one run of it. It reads the same every time you open it, and I can hand an agent the trace ID. The agent then starts from evidence of what happened.
Instrumentation belongs in the same pull request as the feature. A flow that ships without spans can only be debugged by guessing. My next step is carrying the trace context from the browser into the API, so one trace covers both the click and the database query behind it.
Flaky tests teach everyone to retry
A flaky test is a non-deterministic gate. It trains every reader to treat red as "try again", and agents pick that lesson up quickly.
The worst case I have seen started with an agent investigating a timing-dependent test. It wanted to reproduce the failure under CPU pressure, so it started twelve busy loops in the background. They outlived the tool call, the agent and the session. The same machine hosts the CI runners for most of my projects. For about fourteen hours it sat at a load average of 148. A two-minute test shard took nearly two hours, and two releases were blocked by gates that looked flaky.
The fix for the original test was to wait on the real async transition, so the outcome stopped depending on timing. There is now a written rule against generating synthetic load, and a more useful one: check the host's load before re-running a red job. One re-run is allowed. A second failure means stop and diagnose.
Limit what a wrong answer can do
A check catches some mistakes. Others need to be impossible. I wrote earlier about invisible prompts and how readily an agent follows instructions it should ignore. The defence there is deterministic too: least privilege, scoped credentials, and human approval for anything that spends money, emails real people or cannot be undone. Outbound email on my platform passes through a feature flag that defaults to off for every new category. An agent cannot argue with a flag it has no permission to change.
Why I build the tools myself
BugBarn, SpanBarn and FunnelBarn are my own projects, and all three are open source:
- BugBarn for error tracking (GitHub)
- SpanBarn for traces
- FunnelBarn for analytics, feature flags and session recording (GitHub)
Hosted tools like Sentry, Honeycomb and PostHog do this job well, and building my own versions is, by any reasonable measure, a bit crazy.
They are an experiment in letting agents build software and improve it in a loop. Agents write most of their code. The tools also watch themselves: BugBarn reports its own errors to its own dashboard, and its CLI returns issues as JSON, so an agent can read the issue list, pick an issue, fix it and ship the fix through the same gated pipeline as everything else. Owning the tools means I decide what the agent gets to read, down to the shape of an error report.
That loop depends on everything earlier in this post. A system that improves itself drifts wherever the sampler takes it, unless deterministic checks decide which changes count as improvements.
What the barns taught me
The tools measure each other and themselves. BugBarn sends its traces to SpanBarn with tail-based sampling: every trace that contains an error is kept, plus ten percent of the rest. FunnelBarn reports its errors to BugBarn. A few things that loop turned up:
An optimisation that made BugBarn slower. BugBarn's ingest runs through a write queue. A round of changes batched both the producer and the consumer to raise throughput. The tests passed and the design made sense on paper. In production, events arrive one or two at a time, so the producer's batches never filled, and the consumer held a lock while it worked through large batches item by item. The queue drained slower than before. The changes were reverted, and the same pull request added three metrics: queue depth, items processed by kind and outcome, and persist time by kind. The next regression like this one shows up as a line on a graph.
Six seconds per event. Once persist time was measured per kind, the gap was hard to miss: logs persisted in microseconds, events took around six seconds. The cause was a check that counted every matching row to learn whether at least one existed, and it ran per facet, per event. Switching COUNT(*) to EXISTS made it a single index lookup. A regression test now asserts that the query uses the index, so the fix stays fixed.
SpanBarn as the first place to look. Slow database queries show up in SpanBarn as long spans, with the query attached. More than once, the fix for a slow BugBarn page started there and ended as a new index. Another project of mine sends its LLM calls to SpanBarn as well, with the prompt, the model, token counts and latency on each span, which turns prompt tuning into a before-and-after comparison on real traffic. Frontends report Core Web Vitals to it, and services get an Apdex score. When a page feels sluggish or the layout jumps while it loads, SpanBarn shows which page, which element and since which release.
From a trace to what the user saw. The three tools share one key: the W3C trace ID. The FunnelBarn SDK records the active trace ID alongside the session recording, so FunnelBarn can resolve a trace ID to a recording and the moment in it where that trace fired. A small replay CLI takes a trace ID from SpanBarn or BugBarn, fetches the recording and plays it in a headless browser, seeked to that moment, with an option to save a screenshot. An error span becomes a few seconds of what the user was doing when it happened.
That CLI needs no human at the keyboard, which makes it the next experiment. The plan: an agent takes the trace ID from an error, replays the session, looks at the screenshots and reads the trace beside them. It gets the brief of a senior UX reviewer and one question: what went wrong for this user? What they clicked, what the screen showed at that moment, which message appeared or failed to appear. It proposes a change, and the change goes through the same gates as any other. FunnelBarn's MCP server already gives an agent the funnels and conversion numbers, so it can check afterwards whether the change helped.
The lesson from the experiment so far: an agent can improve a product as far as the readouts reach. Past that point it is guessing, the same as I would be.
Where the parrot belongs
The randomness has its place. It helps where variation is cheap: exploring a design, drafting code, proposing three approaches to a bug, writing the first version of a test. I want the model to be inventive there.
Variation gets expensive at the boundaries: merging, deploying, migrating data, spending money, contacting people. Those boundaries get deterministic gates, and the gates treat every author the same.
Human teams have always worked this way. Code review, CI, staged rollouts and backups exist because people make mistakes. Those practices assumed an author who writes a few hundred lines a day. The author now writes thousands, so the checks have to be fast, strict and boring.
The model proposes. The pipeline decides.