Document

lots of examples of something you want them to learn (pictures of cats labeled “cat,” for

Ref IMAGES-004-HOUSE_OVERSIGHT_016909.txt Release House Oversight Committee — Epstein Estate Records (Nov 2025) 1 pages

Epstein Suite indexes the text; the original document lives at its official source. We don't host the original file — view it on the official release to read it in full.

View the original on the official release

Document text

Text is machine OCR and may contain errors. Confirm against the original source above.

lots of examples of something you want them to learn (pictures of cats labeled “cat,” for example), along with examples of other random data (pictures of other things). This is called “supervised learning,” because the neural network is being taught by example, including the use of “adversarial training” with data that is not correlated to the desired result. These neural networks, like their biological models, consist of layers of thousands of nodes (“neurons,” in the analogy), each of which is connected to all the nodes in the layers above and below by connections that initially have random strength. The top layer is presented with data, and the bottom layer is given the correct answer. Any series of connections that happened to land on the right answer is made stronger (“rewarded”), and those that were wrong are made weaker (“punished”). Repeat tens of thousands of times and eventually you have a fully trained network for that kind of data. You can think of all the possible combinations of connections as like the surface of a planet, with hills and valleys. (Ignore for the moment that the surface is just 3D and the actual topology is many-dimensional.) The optimization that the network goes through as it learns is just a process of finding the deepest valley on the planet. This consists of the following steps: 1. Define a “cost function” that determines how well the network solved the problem 2. Run the network once and see how it did at that cost function 3. Change the values of the connections and do it again. The difference between those two results is the direction, or “slope,” in which the network moved between the two trials. 4. Ifthe slope is pointed “downhill,” change the connections more in that direction. If it’s “uphill,” change them in the opposite direction. 5. Repeat until there is no improvement in any direction. That means that you’re in a minimum. Congrats! But it’s probably a /oca/ minimum, or a little dip in the mountains, so you’re going to have to keep going if you want to do better. You can’t keep going downhill, and you don’t know where the absolute lowest point is, so you’re going to have to somehow find it. There are many ways to do that, but here are a few: 1. Try lots of times with different random settings and share learning from each trial; essentially, you are shaking the system to see if it settles in a lower state. If one of the other trials found a lower valley, start with those settings. 2. Don’t just go downhill but stumble around a bit like a drunk, too (this is called “stochastic gradient descent”). If you do this long enough, you’ll eventually find rock bottom. There’s a metaphor for life in that. 3. Just look for “interesting” features, which are defined by diversity (edges or color changes, for example). Warning: This way can lead to madness—too much “interestingness” draws the network to optical illusions. So keep it sane, and emphasize the kinds of features that are likely to be real in nature, as opposed to artifacts or errors. This is called “regularization,” and there are lots of techniques for this, such as whether those kinds of features have been seen before (learned), 106 HOUSE_OVERSIGHT_016909

Have a question about what this document contains?

Ask the documents