Abstract

Trained on funnel-stage preference pairs instead of binary conversion labels, an LLM ranker put 25.76 percent future buyers in the top 0.1 percent of a 6.14 million-lead queue, versus 14.41 percent for the best conventional model and 3.28 percent for recency: the practice change is to score leads against your funnel's own ladder and audit precision at the top of the list, not AUC.

01 How a margin-aware Bradley-Terry loss turns funnel stages into dense training signal

02 What precision at the top 0.1 percent means for a 50-lead hot list, worked line by line

03 Four candidate mechanisms for why stage labels carry signal, each with a strength label

04 One Monday move per maturity level, Manual through Autonomous, plus five industry translations

Fifty leads. In a 50,000-lead month, that is the top 0.1 percent of the queue: the hot list a manager can actually flag for same-day senior follow-up. Sort the queue by recency, the default in most CRMs, and about 2 of those 50 will ever buy. Sort it with the best conventional model in this week’s paper and you get about 7. Sort it with the method the paper proposes and you get about 13. Same leads, same reps, same payroll. The only thing that changed is the order of the list. We work that arithmetic below, line by line, from the paper’s own tables.

The paper is “Rethinking Sales Lead Scoring with LLM-based Hierarchical Preference Ranking,” posted to arXiv in June 2026 by Zhang and colleagues. It is the best lead-scoring study the lab’s standing weekly search has surfaced this year, and also a small masterclass in how to read a results table without being led by the abstract.

A ranking problem wearing a prediction costume

Start with the research question, because it is the part that transfers. In a long-cycle funnel, can a model learn to rank leads from the funnel’s intermediate stages instead of waiting for the rare final label? The diagnosis behind the question is one every operator will recognize. The standard label for lead scoring is binary: converted or not. The model is asked to learn from the rarest event in the funnel while ignoring the densest evidence in it.

The sample: two datasets from that retail system. A benchmark set of 340,000 leads with a 1.45 percent positive rate, and an industrial set of 6.14 million leads at 1.33 percent. Both split 7:3 into train and test in strict temporal order, so the model cannot peek at the future. The outcome label is exact and worth writing down: final conversion, defined as order lock-in.

The architecture they call asLLR. The backbone is a small language model, Qwen at 1.5 to 1.8 billion parameters, fine-tuned with LoRA, reading the CRM’s structured fields together with the unstructured dialogue logs, truncated at 2,000 tokens. Three heads sit on that backbone: a semantic head that keeps the language model honest about what the text says, a pointwise scoring head that outputs the lead score, and a pairwise ranking head that learns from comparisons.

The actual contribution is the training objective, HPRO: hierarchical preference ranking optimization. Instead of one rare binary label, the funnel’s own stages are converted into preference pairs at three tiers, each with its own margin. Global Dominance: a lead that locked in beats a lead that was defeated, margin 1.0. Key Action: a lead that took a test drive beats one that did not, margin 0.5. Soft Signal: a long call beats a short call, margin 0.1. The funnel stops being a reporting artifact and becomes the curriculum.

Reported result Value Context
Classification performance AUC 0.8161 Best result on the 340k benchmark dataset
Ranking performance +39.7% precision Reported against the authors' model without HPRO
Training signal Funnel-aware preference pairs Margin-aware Bradley-Terry formulation over sparse binary labels
Data setting Leading NEV (electric vehicle) brand Long cycle, multi-stage funnel, structured CRM plus unstructured logs

All values are the authors' own reported numbers. The lab has not independently replicated them; limits are discussed in the article.

Exhibit 1 is the result as the authors state it. Hold it loosely for a moment, because the two headline numbers come from two different experiments. The AUC of 0.8161 was measured on the 340k benchmark dataset.

Precision at k is the rep-queue metric

The industrial table is where the paper earns its keep, and it uses a metric worth explaining because it is the one that matches how revenue teams actually work. P@K% is the share of the top K percent of the ranked queue that eventually converts. A rep queue is exactly that: a cut from the top of a ranked list.

Ranking method AUC P@0.1% P@1.0% R@5.0%
Funnel + recency heuristic 0.6332 3.28% 1.50% 9.23%
Funnel + CTR (DeepFM, two-stage) 0.6898 7.21% 1.90% 18.76%
Funnel + CTR (DeepFM, direct) 0.7382 14.41% 10.14% 21.85%
asLLR without HPRO 0.7491 18.44% 11.56% 23.94%
asLLR with HPRO (full) 0.7583 25.76% 13.33% 25.18%
Relative lift vs asLLR without HPRO +1.2% +39.7% +15.3% +5.2%

P@K% is the share of the top K% of the ranked queue that eventually converts. The relative lift row is computed against the authors' own model without HPRO.

Key finding: The label change did the work. Training the same model on funnel-stage preference pairs instead of binary conversion labels lifted top-of-queue precision from 18.44 to 25.76 percent. Model size stayed put; the supervision got denser.

Conclusion

Adopt the method’s two portable moves now: label leads against your funnel’s own stage hierarchy, and audit every score, yours or a vendor’s, on precision at the top of the queue. Treat the specific numbers as one company’s funnel until someone replicates them on a B2B motion.