China's large models are expanding in parameter scale, but the legal foundation of their training corpora is shaking. When a model "eats" trillions of tokens, it consumes not only text, images, and code, but also the works of countless creators, the faces of ordinary people, and the data assets that platforms have painstakingly accumulated.
This book addresses a single core question: under China's current legal framework, what can large model training legally "eat," what can it not, and what liability follows once it has "eaten"?
The answer is not simple. A Shanghai court says that extracting anime images to train a LoRA model infringes reproduction rights ? the act of reproduction during training itself can constitute infringement. A Fujian court says that complying with robots.txt does not mean news data can be freely crawled and commercially used. A Beijing court says that when a platform has deployed technical measures and a crawler bypasses them, that constitutes unfair competition. Yet the Supreme People's Court says that processing publicly available personal information for training within a reasonable scope "generally is not found to infringe personal information rights."
The rules are growing, but far from settled. This book does not offer false certainty. What it does is different: it pieces together the fragments of rules scattered across judgments, regulatory notices, departmental regulations, and academic debates into as complete a compliance map as possible. For anyone training a model, preparing to train one, or fearing that their data has been "fed" into one, this map deserves careful reading.