Training data
noun
The material used to teach an AI model its patterns. What is abundant becomes easy to reproduce; what is missing, rare or repeatedly copied can be flattened or lost.
Also known as training corpus
Training data is the material an AI model learns its patterns from: writing, images, code, recordings and other examples gathered before the model is asked to produce anything of its own. The composition of that material leaves fingerprints. What appears often becomes easier to reproduce; what is rare, local, awkward or badly represented is easier to mishandle, and the model does not know the missing edge was important, only what the data made available. The book returns to training data because human work can flow back into the next machine, along with machine-made work nobody checked. That is where slop stops being merely annoying. At scale, today’s indifferent output becomes tomorrow’s lesson.
See also LLM Model collapse AI slop Extinction debt