jonny@neuromatch.social ("jonny (nonvenomous)") wrote:
whenever i see people talking about how "the AI companies have solved the problems of model collapse because their filtration process includes the AI evaluating whether or not the input data is a dataset poisoning attack" i just start to vibrate in statistical learning theory and stuttering about how like "no this difference and um ah the shit fuck ah imitating the operator vs identifying the operator jesus first chapter of any textbook ugggh the entire discipline is based around putting bounds on the difference between the two, very strict definition of stability yeefafafh"


![[vapnik page 48] 1.13 THE STRUCTURE OF THE LEARNING THEORY Thus. in this chapter we have considered two approaches to learning problems. The first approach (imitating the supervisor's operator) brought us to the problem of minimizing a risk functional on the basis of empirical data. The second approach (identifying the supervisor's operator) brought us to the problem of solving some integral equation when the elements of an equation arc known only approximately. It has been shown that the second approach gives more details on the solution of pattern recognition and regression estimation problems. Why in this case do we need both approaches? As we mentioned in the last section. the second approach. which is based on the solution of the in~ tegral equation. forms an ill-posed problem. For ill-posed problems. the best that can be done is to obtain a sequence of approximations to the solution which converges in probability to the desired function when the number of observations tends to infinity. For this approach, there exists no way to evaluate how well the problem can be solved if a finite number of observations is used. In the framework of this approach to the learning problem. any exact assertion is asymptotic.](https://files.mastodon.social/cache/media_attachments/files/116/968/504/726/059/464/original/3728731b867323a3.png)