Language models that trained on their own output during inference became worse at predicting independent human text, a preprint finds after following the feedback loop across long streams.
Test-time training lets a model update its weights while it is being used, turning the model itself into a form of working memory. Cheng Luo, Bing Li and Bernard Ghanem tested what happens when the text used for those updates is generated by the same changing model.
The result matters for developers building agents that learn continuously from their own work. An update can make the latest generated passage easier to predict while quietly reducing performance on fresh human-written material. Repeating that local improvement changes both the learner and the source of its next lesson.
The generator and learner pulled each other off course
The researchers ran three test-time-training model configurations, labelled 125 million, 760 million and 3 billion parameters, over streams of 128,000 tokens. Retaining updates from generated text worsened prediction on a separate set of human text in all three. Applying ordinary Adam updates to the existing weights of Qwen3-4B produced the same pattern.
Learning during inference was not inherently damaging. The same update mechanisms improved performance when the incoming text came from people. The failure appeared when model-generated text fed back into the model that would generate the next training chunk.
The paper separates that loop with matched experiments. In “Fixed Generation,” a frozen copy supplied the generated training text while another model kept updating. This removed more than 98% of the measured damage for the two smaller configurations, according to the authors. The comparison identifies the moving generator, rather than synthetic text alone, as the main source of instability.
A second replay experiment distinguished two costs. The model first had to read text that had degraded over time; it then stored further changes by training on that text. A paired single-update test exposed the immediate conflict: each update improved prediction of its source passage but worsened prediction of new human text.
That conflict is easy to miss when the training signal and evaluation target are the same passage. The update is successful by its local objective: loss falls on the text that produced it. The damage becomes visible only on held-out human text, which acts as a reference for whether the model is preserving its broader language distribution while adapting.
Independent evidence acted as a brake
Most runs did not fail equally. The authors report that a small number of trajectories accounted for many of the largest losses after prolonged closed-loop adaptation. That concentration makes average short-run checks a weak safeguard: a system can look stable until one sequence of self-generated updates pushes it farther away.
Long streams matter because every accepted update changes the starting point for the next one. Errors therefore affect more than the current prediction; they alter which text is produced and learned from later. The study measures that compounding process instead of judging isolated self-training steps.
The proposed mitigation, called Settlement, evaluates a candidate weight update on independent real text before accepting it. In the two smaller configurations, the remaining average endpoint difference was 0.07 and -0.02 nats, while useful adaptation to real text was retained. The metric measures predictive loss, so values near zero mean the safeguard largely closed the gap to the comparison condition.
These are author-reported results from a first-version arXiv preprint, without independent replication. The tested streams and prediction objective are narrower than a deployed agent’s full behaviour, and the study does not establish how quickly the same loop would appear in larger frontier systems.
The experiment nevertheless gives a causal account of a familiar intuition about self-training: checking whether an update explains its own source is insufficient. The authors’ next concrete criterion is external—an update must also preserve prediction on evidence that the adapting model did not generate.
Verification
| Claim | Label | Primary source | Independent check |
|---|---|---|---|
| Retaining generated-text updates worsened prediction on independent human text across three test-time-training configurations | VENDOR-REPORTED | Luo, Li and Ghanem preprint | none |
| The same failure appeared when Adam updated Qwen3-4B’s existing weights | VENDOR-REPORTED | Luo, Li and Ghanem preprint | none |
| The same update mechanisms improved on real text | VENDOR-REPORTED | Luo, Li and Ghanem preprint | none |
| A frozen generator removed more than 98% of the damage at 125M and 760M | VENDOR-REPORTED | Luo, Li and Ghanem preprint | none |
| Settlement left endpoint gaps of 0.07 and -0.02 nats at 125M and 760M while retaining real-text adaptation | VENDOR-REPORTED | Luo, Li and Ghanem preprint | none |
| The three-author preprint was submitted on 4 October 2026 | VERIFIED | arXiv record | none |