Which support conversations are about to go bad?

A triage model scores every incoming message the moment it arrives — before anyone has replied, and with nothing but that first message to go on. Built on 91,449 real support conversations with AmazonHelp, Delta, TMobileHelp and Tesco.

The queue, 2017-12-01

Every conversation that opened that day, ranked by predicted risk. The line marks the operating threshold: fast-track the worst 10% of the queue.

Drivers are the features that pushed each score up, in plain words. For a linear model these are exact — the feature's value times its weight is precisely what moved the score, with nothing approximated.

Is it better than the rule a support tool ships with?

Ranking quality

PR-AUC, because bad outcomes are the minority and a queue is ranked. Base rate is 23.0% — that is what PR-AUC a coin toss would score. The MiniLM variant genuinely wins by +0.012 — not noise, and it is still not the deployed model: at the operating point the two catch 19.2% and 19.0% of bad outcomes, 0.3 points apart, and embeddings cannot produce the plain-words drivers above.

Does a score of 0.4 mean 40%?

A triage threshold is a staffing decision, so the score has to mean what it says. Isotonic regression on a time-held-out slice; it also compresses the range, so nothing scores above 0.6 — the model declining to call anyone more than coin-flip risky.

Where the line goes, and what it costs

Read a row as: fast-track everything above this score, staff for that many conversations a day, and that share of them actually go badly — while that many bad ones a day go through the normal queue unprioritised.

Does replying faster change how it ends?

All conversations

Within each brand

Brands differ in both how fast they answer and what they are asked, so the pooled number is the one to distrust. Stratifying narrows the confounding; it does not close it.

This is association, not causation. Nobody randomised who got a fast reply. The plausible mechanism — that fast replies are largely templated acknowledgements which resolve nothing — is a hypothesis this data cannot test.

What this cannot tell you