Skip to content

Latest commit

 

History

History
16 lines (11 loc) · 4.27 KB

File metadata and controls

16 lines (11 loc) · 4.27 KB

This project is about explainable skin cancer classification by finetuning a vision language model with RL on a dataset of benign and malignant skin moles. The data can be found in the data/ directory (data_nishaydnk should be ignored as it is an old dataset kept in case I need to revert to it). The environment code is found in rl_env.py and the training code is in train.py. Continually document and refine upon your workflows, and store your notes in the codex_notes directory. I saw this post by an OpenAI engineer about how he uses you, and I want you to replicate this as much as possible:

The big unlock was getting codex to continually document and improve its own workflows. This is something I fully hacked together for my personal setup. Codex consistently gets better and faster at tasks I use it for, just because I have the habit of asking it to take notes and improve. While working, codex commits notes and helpers to my personal folder in our monorepo. After a few interactions with a new part of the codebase, these helpers tend to stabilize. I've never actually read these notes, their utility to me is purely the effect on codex's performance. With my setup now able to compound knowledge across sessions, I got comfortable scaling up the tasks I used it for. Let’s dive into two tasks I recently spent hundreds of millions of tokens on. Scaling Research

Research moves fast. Experiments are expensive and easy to misconfigure, so staying on top of the most recent findings and gotchas is crucial. Luckily, codex is an amazing search engine. When I want to quickly implement a one-off experiment in a part of the codebase I am unfamiliar with, I get codex to do extensive due diligence. Codex explores relevant slack channels, reads related discussions, fetches experimental branches from those discussions, and cherry picks useful changes for my experiment. All of this gets summarized in an extensive set of notes, with links back to where each piece of information was found. Using these notes, codex wires the experiment and makes a bunch of hyperparameter decisions I couldn’t possibly make without much more effort. Asking for a second opinion greatly increases my confidence in what I'm shipping. In settings where mistakes are costly, you want an incredibly diligent, high-recall search agent. Codex routinely scratches that itch for me. Coding agents are also great at data analysis, and have made it very easy to quickly get insights from data. Currently, the real bottleneck is figuring out what to analyze. Recently, I aggressively scaled some of our model behavior efforts using codex. I realized that our internal slack is filled with discussions, reports, and data all relating to different types of model behavior which we might want to test for more rigorously. I used codex to locate and extensively crawl the appropriate channels and generate descriptions of testable hypotheses. Beyond reading slack, it looked at screenshots people shared, pulled documents related to model behavior, and navigated spreadsheets. Over the course of several hours, this resulted in over 700 new hypotheses which are currently improving our understanding of model behavior and user preferences.

You can find my experimental results, exported from wandb, in the experiments directory, along with descriptions of the experiment. Look through the tinker cookbook to learn the actual meanings of the different metrics. All runs use the same model. A bug meant that the clip higher run and baseline run both had random seeds. All future ones are deterministic. The tinker cookbook contains a lot of the code that my codebase references, so you can take a look at the source for many things. Tinker is a service that allows you to run the training scripts for large models on your local dev box and it handles rollouts, forward passes, checkpoint storage/memory, and optimization all in the cloud.

Do not feel the need to make a single monolithic notes file. Feel free to make a directory structure with focused, individual notes that you can easily search through (although this is just a suggestion and you can do what you feel works best). You must use and update the index in the codex_notes folder. If you notice anything while doing a task that you want to jot down, even if it is unrelated to the task at hand, feel free to do so.