Skip to content
ALL WRITING

// build log

A correct answer nobody acts on

graphyn answered "what breaks if I change this" correctly for months, and almost nobody changed what they did with the answer. Renaming it to OpenInvar meant admitting that an advisory tool and a gate are different products — and that determinism stops being a preference the moment a tool is allowed to say no.

min read1,275 words

For most of this year I worked on a tool called graphyn. It read a repository, built a deterministic symbol graph, and answered the question I had originally wanted answered:

code
graphyn query blast-radius UserPayload

Fourteen symbols, one of them reached through an alias no text search would find. Correct. Alias-aware, property-aware, reproducible byte for byte. I wrote a long post about the decisions underneath it, and I still think those decisions were right.

It is now called OpenInvar, and the rename is the least interesting part of what changed.

The thing I kept not noticing

I used it. I would run blast-radius before a refactor, read the output, and proceed more confidently than I would have otherwise. That is a real benefit and I am not going to pretend it was not.

But I never had to. Nothing in my workflow required the answer. On the days I forgot, nothing stopped me. When an answer surprised me, I looked at the code and decided for myself, which is exactly what I would have done without the tool — only slightly later.

A tool you consult when you remember to is competing with your memory, and your memory is free.

What agents made obvious

I had written that the interesting shift was agents needing context about consequences, and I built an MCP server so they could ask for it. That part was right. What I got wrong was assuming asking was the hard part.

MCP is pull. The agent has to decide to ask, and the agent least likely to ask is the one about to make a locally-correct change with a non-local consequence — because from inside the file it is editing, there is nothing to be curious about. Availability is not adoption. I had built a very good answer and put it behind a door that only opens from the inside.

Then I watched an agent's pull request go green in a way I did not like.

It compiled. The suite passed. A test that had covered the changed function still referenced it — still ran, still green — and no longer asserted anything. The assert_eq! had become a call with its result dropped. Nothing about that is malicious; it is what "make the tests pass" looks like when the model is optimising the wrong sentence. Coverage tools saw an edge. Review saw a cleanup.

I had a graph that knew the symbol changed and knew which tests reached it. It could have said so. Nobody asked it.

Advisory and gate are different products

The fix was not a feature. It was noticing that I had been building one kind of thing while describing another.

An advisory tool is judged by whether its answers are good. A gate is judged by whether it is in the way. They share almost all of their implementation and almost none of their design constraints, and the gap between them is where graphyn was quietly living: good answers, no position in anyone's path.

So the product became two commands that exit non-zero.

openinvar audit compares two revisions and reports changes that look like they were made to pass a check rather than to work. Four detectors, and the one I built the whole pivot around is assertion-removal — a test that kept covering a changed symbol and lost its assertions. Assertions are counted during the same parse the symbols come from, so there is no second source of truth to disagree with the graph.

openinvar check enforces constraints the repository writes down in a committed openinvar.toml: layering, forbidden dependencies, field stability, cycles, fan-in ceilings, and coverage of new API surface.

Same graph. Same traversal. Different question, and the difference is entirely about who has to act.

Determinism stopped being a preference

I used to defend determinism on grounds of trust. If the blast radius is probabilistic you have to verify it manually, and then the tool saved you nothing. That argument is fine and it is not the real one.

The real one is: you cannot gate CI on an opinion.

A check that returns a different answer on a re-run is a check nobody can act on. The first time it disagrees with itself, somebody adds --no-verify to the pipeline, and after that it is decoration. Every other design decision in the project is downstream of that sentence, which is why "no model participates in graph construction, or in any gating decision" is now the first line of the README instead of a paragraph somewhere in the middle.

It cost something to keep. There are questions a model would answer well that the tool now refuses to answer at all.

The bill for being allowed to say no

Blocking a commit changes what a blind spot means.

When the tool only answered questions, an unresolved region was a gap in coverage — unfortunate, and the honest move was to name it. Now an unresolved region is a correctness problem, because most of what a gate concludes is drawn from the absence of something. "No test covers this symbol" and "I could not see the test that covers this symbol" produce the same empty set and warrant opposite actions.

So three things had to be built that an advisory tool never needs:

Tiers, enforced rather than documented. Languages with full import resolution are Tier 1. Languages analysed only through their grammar's tags query — symbols and references inside one file, nothing across files — are Tier 2. Gates do not fire on Tier 2 regions. Not "are discouraged from": check returns undecided rather than a pass, tests refuses to call its selection complete, and every audit detector is restricted to files a Tier 1 adapter resolved, with the framework dropping an out-of-scope finding even when a detector forgets to check for itself.

Two coverage numbers, because one of them flatters. status reports resolved edges and references bound:

code
Resolved           99.8% (5980 of 5992 edge(s))
References bound   69.6% (5980 of 8597 reference(s))

The first says whether a gate can act on what is in the graph. The second says how much of the source is in the graph, and it is the one that hurts. An edge exists only once something bound it, so a reference the analysis could not place never enters the denominator — which means failing to bind more references makes the first number go up. I shipped a metric that improves when the tool gets worse, and did not notice until I tried to use it to track progress.

Failing open, loudly. A missing binary, a missing graph, a stale snapshot, a timeout — every one of those lets the commit through and says on stderr that nothing was checked. A gate that silently stops working is worse than one that is plainly off, because the second is a missing feature and the first is a false sense of coverage.

What I would tell myself in March

The graph was not the hard part, and neither was parsing — I had already written that post. The hard part was noticing that correct and useful are separate properties, and that I had been treating the first as evidence of the second for months because it is so much easier to measure.

The uncomfortable version: the most valuable change to this project in a year of work was not code. It was moving it from a place where people could ask it things to a place where it gets asked automatically, and then paying the honesty bill that sending it there created.

cargo install openinvar-cli, if you want to point it at something and tell me where it is wrong. I would rather hear that than hear it was interesting.