12 August 2026 4 min read

Human Still Wins over LLM: the paper that beat me to it by two years

Eight computer scientists published a paper with my book’s title two years before I wrote a word of it, and their small, expert-taught models held their own against GPT‑4.

I was searching my own book title (a habit I would recommend to exactly nobody), and up came a research paper called Human Still Wins over LLM. Eight authors. November 2023. Two years before I wrote a word of mine.

I read it twice, once for the fright and once properly.

It is by Yuxuan Lu and seven colleagues, it sits on arXiv as preprint 2311.09825, and what they set out to test was whether the best language models of the day could beat small ones trained on annotations from people who genuinely knew the subject. Four sets of data across three specialist fields. On one side, GPT-3.5 and GPT-4. On the other, compact models fed a modest number of examples labelled by domain experts and then allowed, through active learning, to ask for whichever example they wanted next, which is a generous way of saying the little model gets to choose what it is taught.

The little ones won. A few hundred expert-labelled examples and they were past GPT-3.5, and with more they held level with GPT-4, in places better, while being hundreds of times smaller.

What the authors recommended off the back of it was a warm-up. Let the model do the first rough pass, the sort, the easy ones that were never going to be the problem. Then keep the expert, because the knowledge that settles the hard cases is sitting inside a person and there is no quantity of parameters that goes and fetches it.

This is a preprint that’s never been through a journal or a conference, and it was tested against GPT-3.5 and GPT-4 as they stood in late 2023, which in this field is roughly the Pleistocene. Anyone who wants to wave it away can point out that the models have travelled a very long way since, and they would be right to. It has been cited since, in workshop papers and in a 2025 survey of active learning, so it didn’t sink without trace but nobody should hold it up as proof of what AI cannot do in 2026.

What survives is the shape of the finding. Scale lost to knowledge, in the one setting where being nearly right counts as wrong.

That is the argument chapter 6 – The Invisible Line makes, and chapter 8 – The Water Downstream after it. The machine does the volume; a person carries the judgement; and the person has to be good enough to notice when the machine has handed over something plausible and wrong. Plausible and wrong is the expensive one. It is also the one that looks fine to everybody in the meeting who has not spent eleven years learning what a correct answer looks like in that particular field.

The warm-up description is accurate to how the work actually feels, by the way. You get the draft and the rough sort handed to you in four seconds, and then the entire job becomes working out which parts of it are wrong. Deciding which parts are wrong is not a smaller job than writing was. It is frequently a harder one, and it is much less obvious from the outside that anything at all is happening.

I did briefly consider being annoyed about the title. And then it occurred to me that the phrase had been lying about in the air in 2023 for anybody to pick up, and eight computer scientists picked it up first and then filed it under annotation methodology, where no general reader has ever gone looking, and there it has sat ever since being cited by other people who file things under annotation methodology.

They got there with four datasets and a table. I got there with interviews, a year, and a small archive of my own worse decisions, and I doubt either of us would recognise the other’s method as work. Their version is still up, it is free, and it is ten pages, which is more than can be said for mine.

Mine is 242 pages and it costs money, and if you want to know what the extra 232 bought, it’s here.