Oncall-agent, part 4. Part 3 left me with a memory that can tell me what's near a new alert in about forty milliseconds. This one is about the evening I found out that near is not the same as useful, and that the actual thinking starts after the clever part is over.
I had it working. That is the sentence I want to keep, before I ruin it.
Tuesday, late, everyone asleep on both sides of the planet. I pasted in a live alert, hit enter, and ten past incidents came back ranked by distance, faster than I could take my hand off the key. I remember sitting back. There is a very specific smugness available to a person who has just made a vector search return something, and I had all of it.
Then I read what came back.
First hit: good, genuinely the same problem. Second: also good. Third: a different product flavour entirely, close only because ops people write the same six sentences forever. Fourth: a false alarm somebody had closed with "no action needed." Then five more that were about Kubernetes in the way that almost everything we get paged for is about Kubernetes.
So. Three useful, seven not. And the thing is, the search wasn't wrong. It did exactly what I asked. I just hadn't noticed that I'd asked the wrong question, and I'd been not noticing for two posts.
What I actually asked for
Vector search answers one question: what is geometrically closest to this point. It is very good at it. It has no opinion whatsoever about what a tired person should read at three in the morning.
Those two questions overlap enough that I'd been treating them as the same question. They are not, and the gap between them is the whole difference between a tool someone uses and a tool someone learns to scroll past.
Here is what kept nagging me. Picture showing the alert to the person who has been on this team the longest. They don't hand you ten similar things. They say: yeah, I've seen this shape before, but was yours on the same flavour of deployment, because if not then ignore everything I'm about to say. And don't restart it. We tried that. It comes straight back.
Three moves in that, and only the first one is similarity.
Find things that might be about the same problem. Throw out the ones that don't apply here. Of what's left, say what to do first and what not to waste time on.
The database does the first one. It does not do the other two, and it was never going to, and somewhere in the back of my head I think I knew that and preferred the version where the hard part was already finished.
Recall wide, narrow hard
What I ended up with is embarrassingly plain, which by now I should be expecting.
new alert
│
├─ normalize → signature (part 1)
├─ distill → embedding text (part 2)
│
▼
[1] COARSE RECALL vector search, top-k = 10
↓ semantic. generous. cheap. dumb.
[2] NARROW re-rank by shared payload facts
↓ keep 3 to 5. the rest get dropped.
[3] RANK FIXES worked > partial > failed, then recency
↓
answer
Ten in, three or four out. That ratio is the design. Everything else is detail.
I keep wanting to make stage one smarter and I have to keep talking myself out of it. Recall has to stay generous. The entire reason for the embeddings was to catch the incident nobody worded the same way, and every time I tighten the net to keep the junk out, the rhyming cousin goes with it. So recall stays loose and a bit stupid on purpose, and the next stage carries the taste.
Which is where the split from part 2 finally earns itself. I wrote back then that the vector holds meaning and the payload holds facts, and that the next post would just be about exploiting that. This is that. The embedding gets me into the right neighbourhood. The payload decides whose door is worth knocking on.
candidate cos shared entities
──────────────────────────────────────────────────────────────────
<env> ipsec pod not ready alert … 0.91 component ✓ flavor ✓ env-kind ✓ → keep
<env> ipsec tunnel flapping after node drain 0.86 component ✓ flavor ✓ → keep
k8s deployment under hr <app> in ns <ns> down 0.84 (none) → drop
<env> config generator dependent service failure 0.83 (none) → drop
<env> pod not ready · note_class: auto/false-alarm 0.82 component ✓ but carries no signal → drop
(scores illustrative, shape real)
Set intersection over component_tags, flavor, service, error_signature, environment kind. Count the overlaps, break the ties distance can't break. That's it. That's the clever part I spent a week working up to.
The model finds the candidates. The boring code picks. Third time in three posts I've written some version of that sentence and I'm starting to think it's not a coincidence about this project, it's just how these things go.
The reason I actually narrow, which took me a while to admit
There's a second motive and I didn't own up to it until I watched the token counter tick.
Everything past retrieval goes to an LLM, which has to read the past incidents and write the recommendation. Ten full records, with all the symptoms and notes and fixes attempted, is a lot of reading. Slow, costs real money on every alert, fine, I expected that. What I didn't expect: the answer gets worse. Bury three good matches in seven mediocre ones and the model finds a polite way to acknowledge all ten. What comes out is a paragraph hedging across four different root causes, technically defensible, useless to a person holding a pager.
So narrowing isn't a cost saving that happens to improve quality. It's the other way round. Three sharp records give a sharper answer than ten fuzzy ones, and they cost a fraction as much, and I get to feel thrifty and correct at the same time.
That keeps happening here. The cheap choice and the good choice turn out to be the same choice. I suspect they usually are and I only ever notice when one of them has a number attached to it.
The bit that sounds like a person
Three incidents survive the narrowing. Between them they carry a pile of fix records: an action, an outcome, a date.
{"action": "restart the pod", "outcome": "failed", "date": "…"}
{"action": "restart the pod", "outcome": "failed", "date": "…"}
{"action": "clear the stale config", "outcome": "worked", "date": "…"}
{"action": "drain and recycle the node", "outcome": "partial", "date": "…"}
Aggregate across the survivors, sort by outcome (worked, then partial, then failed) and then by recency, and the thing stops reading like a search result:
🔎 Likely root cause stale config not cleared after the last rollout (3 similar incidents)
📁 Where to look ipsec pods · the config volume · last deploy PR
🛠 Try in order
1. clear the stale config, then restart worked 2/2 · last: 6 weeks ago
2. drain and recycle the node partial 1/2
⚠️ Skip: plain pod restart. Tried 3×, came straight back every time.
📚 Similar past PD-…, PD-…, RB-ipsec-not-ready
Recency as the tiebreak matters more than it looks. A fix that worked eighteen months ago, on an architecture we have since torn out, is a trap dressed up as an answer. Worked, then recent. One line of sort, and what it's really encoding is that the system keeps moving underneath us.
The skip line
That ⚠️ Skip line is the best thing in the whole output and I did not plan it. It fell out of the schema.
Go back to why I started. IST spends an hour on an alert, finds the cause right as their day ends, hands over the pager. PST wakes up to the same alert and starts from nothing. When I imagined fixing that I imagined moving the answer across the handoff. But an hour of on-call is mostly not the answer. It's the four things that didn't work. That's where the hour actually goes.
And that is exactly the part nobody writes down. You record what fixed it, because that reads like a finding. Nobody records "restarted it twice, came straight back," because it reads like a confession. So the wins trickle into the tracker and the dead ends evaporate at eight o'clock every night, and the next person, eleven hours away, pays for them again at full price.
outcome is a two-word schema decision. What it's doing is making failure survive the night. That's the one kind of knowledge our process was structurally guaranteed to lose, and it turns out you can keep it by giving it a field.
I did not set out to build a thing whose best feature is remembering what didn't work. It's the one I'd argue for hardest now.
Saying nothing, on purpose
The failure mode I'm most afraid of is the confident one.
A tool that invents a plausible root cause for an alert it has never seen anything like is worse than no tool. It costs the tired human twenty minutes to disprove, and it only gets to do that twice before they stop reading it. I've been that human with other tools. I know exactly how fast that trust goes.
So confidence isn't decoration on the output, it's derived from things the pipeline already knows: how many candidates survived narrowing, how much entity overlap they had, how far the top score sits above the pack. Thin overlap, few survivors, no separation, and the honest output is a shrug. So print a shrug.
🤔 Nothing close in memory.
Nearest by wording (weak match, different component):
· <env> config generator dependent service failure · PD-…
No fix history to offer. This may be new.
That is a success. It cost almost nothing, it told the truth, and it left the human's judgement alone instead of anchoring it to a guess. "Here's the nearest I've got and it isn't close" is a sentence you can build a habit on. "Likely root cause, 34% confidence" is a sentence people learn to skim.
Calibration turned out not to be a statistics problem. It's a manners problem.
Where I was too clever, again
Two, keeping the confession going.
The false alarms were quietly eating the neighbourhood. Part 1 taught me to strip volatile text out of a title. It did not occur to me that an entire incident could be noise. Plenty of ours close as "false alert": the thing recovered on its own, nobody touched it. They embed beautifully. They land right in the middle of the cluster they belong to, because they genuinely are about that problem. And they carry nothing. No cause, no fix history, nowhere to look. For a couple of weeks the agent was burning one of its three or four slots on incidents whose whole content was this turned out to be nothing. In post zero I mentioned the first real alert it looked at, which it correctly called a false alarm, right and useless at the same time. It took me far too long to connect that observation to so down-weight these at ingest. The lesson isn't about false alarms. It's that "is this true" and "is this worth a slot" are two different tests and I had only written one of them.
The other one: I went looking for a similarity threshold and there isn't one. My instinct was a cutoff. Keep anything over 0.8, drop the rest, done. It doesn't work and the reason is structural. Every document in the collection is an infrastructure incident written in the same clipped ops register, so everything is a bit similar to everything, and the scores come back bunched in a narrow band near the top. An absolute cutoff either passes nearly everything or nearly nothing, and the number where it flips depends on the corpus, which changes every week. What works is relative: take the top-k, look at the gap between the best and the pack, let the payload decide the rest. Cosine similarity is a comparator, not a measurement. I spent two evenings trying to read a ruler that only knows how to say "more."
What I think this taught me, though I'm not sure yet
Part 1 was subtraction. Decide what to ignore. Part 2 was recognition. Let the machine feel the shape of a problem without teaching it any words. This one is messier and I've been chewing on it for days without landing it cleanly.
The closest I can get: retrieval is a judgement, not a lookup. I spent the glamorous half of this project on embeddings and a vector database, and the part that decides whether the output is any good is a set intersection over some metadata and a sort with three tiers in it. The model finds candidates. The judgement is ordinary code holding an opinion I had to actually have.
Which means what I was really doing, all that fiddling with tiebreaks, was writing down how a good on-call engineer thinks. Not what they know. Anyone can store what they know, that's a database, we've had those for fifty years. How they discount things. How they qualify. The reflex to ask "yes but was that the same flavour." The instinct to lead with what didn't work. That's the part that walks out of the building at eight and comes back, if you're lucky and they slept, eleven hours later inside somebody's head.
A surprising amount of it goes into a sort function. I don't know what to do with that yet. It doesn't make me feel clever, exactly. Slightly uneasy, more like, in a way I'm choosing to keep writing down rather than resolve.
Next time is the last mile, and it's the one that decides whether any of this gets used at all. A good answer in the wrong place, in the wrong tone, at the wrong moment, is not an answer. The alerts land in a chat channel, so the reply belongs in the chat channel, threaded under the alert, a few seconds behind it, short enough to read on a phone at three in the morning, quiet enough not to become one more thing people learn to ignore. That's reading the room. It's where this either becomes a colleague or becomes a bot.
Comments