Mastodon Feed: Post

Mastodon Feed

jonny@neuromatch.social ("jonny (nonvenomous)") wrote:

whenever i see people talking about how "the AI companies have solved the problems of model collapse because their filtration process includes the AI evaluating whether or not the input data is a dataset poisoning attack" i just start to vibrate in statistical learning theory and stuttering about how like "no this difference and um ah the shit fuck ah imitating the operator vs identifying the operator jesus first chapter of any textbook ugggh the entire discipline is based around putting bounds on the difference between the two, very strict definition of stability yeefafafh"

"Skip the model, just output random garbage directly" subtitle: Your boss hasn’t read Vapnik and doesn’t know the difference between approximating the system and approximating the system’s output, and neither should you
Vapnik, statistical learning theory, page 21: two goals to pursue:   • To imitate the supervisor's operator: Try to construct an operator which  provides for a given generator G, the best prediction to the supervisor's outputs.  • To identify the supervisor's operator: Try to construct an operator which  is close to the supervisor's operator.   There exists an essential difference in these two goals. In the first case. the goal is to achieve the best results in prediction of the supervisor's outputs for the environment given by the generator G. In the second case, to get good results in prediction is not enough; it is required to construct an operator which is close to the supervisor's one in a given metric. These two goals of the learning machine imply two different approaches to the learning problem.   In this book we consider both approaches. We show that the problem of imitation of the target operator is easier to solve. For this problem, a nonasymptotic theory will be developed. The problem of identification is more difficult. It refers to the so-called ill-posed problems. For these problems, only an asymptotic theory can be developed. Nevertheless, we show that the solutions for both problems are based on the same general principles.
[vapnik page 48] 1.13 THE STRUCTURE OF THE LEARNING THEORY   Thus. in this chapter we have considered two approaches to learning problems.  The first approach (imitating the supervisor's operator) brought us to the problem of minimizing a risk functional on the basis of empirical data.   The second approach (identifying the supervisor's operator) brought us to the problem of solving some integral equation when the elements of an equation arc known only approximately.   It has been shown that the second approach gives more details on the solution of pattern recognition and regression estimation problems.   Why in this case do we need both approaches? As we mentioned in the last section. the second approach. which is based on the solution of the in~ tegral equation. forms an ill-posed problem. For ill-posed problems. the best that can be done is to obtain a sequence of approximations to the solution which converges in probability to the desired function when the number of observations tends to infinity. For this approach, there exists no way to evaluate how well the problem can be solved if a finite number of observations is used. In the framework of this approach to the learning problem. any exact assertion is asymptotic.