This project is about solving quantum field theories with a Normalizing Flows neural network. Continually document and refine upon your workflows, and store your notes in the codex_notes directory. I saw this post by an OpenAI engineer about how he uses you, and I want you to replicate this as much as possible:
The big unlock was getting codex to continually document and improve its own workflows. This is something I fully hacked together for my personal setup. Codex consistently gets better and faster at tasks I use it for, just because I have the habit of asking it to take notes and improve. While working, codex commits notes and helpers to my personal folder in our monorepo. After a few interactions with a new part of the codebase, these helpers tend to stabilize. I've never actually read these notes, their utility to me is purely the effect on codex's performance. With my setup now able to compound knowledge across sessions, I got comfortable scaling up the tasks I used it for. Let’s dive into two tasks I recently spent hundreds of millions of tokens on. Scaling Research
Research moves fast. Experiments are expensive and easy to misconfigure, so staying on top of the most recent findings and gotchas is crucial. Luckily, codex is an amazing search engine. When I want to quickly implement a one-off experiment in a part of the codebase I am unfamiliar with, I get codex to do extensive due diligence. Codex explores relevant slack channels, reads related discussions, fetches experimental branches from those discussions, and cherry picks useful changes for my experiment. All of this gets summarized in an extensive set of notes, with links back to where each piece of information was found. Using these notes, codex wires the experiment and makes a bunch of hyperparameter decisions I couldn’t possibly make without much more effort. Asking for a second opinion greatly increases my confidence in what I'm shipping. In settings where mistakes are costly, you want an incredibly diligent, high-recall search agent. Codex routinely scratches that itch for me. Coding agents are also great at data analysis, and have made it very easy to quickly get insights from data. Currently, the real bottleneck is figuring out what to analyze. Recently, I aggressively scaled some of our model behavior efforts using codex. I realized that our internal slack is filled with discussions, reports, and data all relating to different types of model behavior which we might want to test for more rigorously. I used codex to locate and extensively crawl the appropriate channels and generate descriptions of testable hypotheses. Beyond reading slack, it looked at screenshots people shared, pulled documents related to model behavior, and navigated spreadsheets. Over the course of several hours, this resulted in over 700 new hypotheses which are currently improving our understanding of model behavior and user preferences.
Do not feel the need to make a single monolithic notes file. Feel free to make a directory structure with focused, individual notes that you can easily search through (although this is just a suggestion and you can do what you feel works best). You must use and update the index in the codex_notes folder. If you notice anything while doing a task that you want to jot down, even if it is unrelated to the task at hand, feel free to do so.
When changing the architecture of a normalizing flow model, test that invertibility was preserved by running forward passes on random batches and then checking that the inverse actually restores the original batches.