The finishing touches — Poolside
Introduction
In the previous posts in this series, we’ve discussed almost every aspect of the Model Factory. We’ve built everything that we need to process massive datasets, quickly train excellent foundation models, teach models to code, evaluate specific capabilities, and serve inference at scale. But there’s still more to do to produce foundation models that demonstrate useful abilities in real-world environments. Moreover, we want to produce models that exhibit capabilities beyond those that can be granted by traditional pre-training alone. In order to achieve these aims, we’ll need to employ some additional techniques.
In this post, we’ll discuss how we conduct post-training on our models at Poolside. In particular, we’ll consider two aspects of our post-training pipelines, namely supervised fine tuning (SFT) and reinforcement learning (RL). We’ll start by explaining why we conduct post-training on models at all, introducing SFT and RL along the way. We’ll then discuss how we’ve already introduced almost all the pieces needed to orchestrate and deploy post-training workloads, using SFT and RL workloads as examples. Lastly, we’ll conclude by discussing two systems that we’ve built for post-training workloads, namely a dataset viewer and GPU<>GPU weight transfer.
Before we get started, it’s important to remember that this post in particular shouldn’t be consumed in isolation. In fact, every post in the Model Factory series has led up to this one because post-training reuses components that we’ve already introduced earlier in the series. This is intentional: reusable components and large-scale orchestration are two of the main reasons why the Model Factory is such a powerful tool for building best-in-class foundation models. Here, we’ll focus on new features that we’ve added to the Model Factory to enable post-training workloads, as well as how we orchestrate our existing components to support running these workloads at scale.
Why do we carry out post-training?
Let’s start with a brief recap of foundation model training. As we mentioned in the first post of this series, typical foundation model training starts with building an initial dataset on which to train our model. We then use that dataset to conduct an intensive training period, known as pre-training, which produces a trained base model.
While base models are knowledgeable, they’re not normally good at problem solving. As a result, it’s common to carry out some extra steps after pre-training, known as post-training. The goal of post-training is to produce a model with specialized capabilities, like instruction following or proficiency at certain tasks. Essentially, post-training can be seen as the process that upgrades a basic text completion model to an assistant that can handle conversations, use tools, think explicitly, and solve various problems. As we seek multiple behaviors and capabilities from our models, post-training may be conducted over multiple rounds.
It’s important to note that post-training shouldn’t be viewed as a mere add-on to pre-training. In fact, post-training is essential to produce models that are usable and intelligent in practice. As a concrete example, we’ll consider Poolside’s primary use case: producing models that excel at software development. While base models are typically knowledgeable about language syntax and math, there’s a huge gap between basic knowledge and being a good software engineer. This is even true in humans: software engineers require a great deal of operational knowledge and expertise to be effective. Models are no different, and there is a big difference between a good base model and a model that can act as an effective coding agent inside a real-world environment. As a result, we invest a considerable amount of compute and effort into our post-training pipelines, ensuring that we build models that are actually useful in practice.
Let's consider two distinct forms of post-training: namely, supervised fine tuning (SFT) and reinforcement learning (RL). Before we get into how we build these workloads in the Model Factory, we’ll briefly highlight how we use both SFT and RL at Poolside.
On the one hand, we typically use SFT to coax a model into exhibiting predictable behavior, like producing structured outputs or demonstrating a particular “personality.” We also use SFT to provide a “warm start” for RL workloads. In practice, we follow a standard SFT workflow: we’ll build an SFT dataset, continue training on a pre-trained checkpoint, and then store the resulting checkpoint. And, similar to our pre-training workloads, we regularly run evaluations during our SFT experiments, allowing us to understand how our models change during SFT experiments. We’ll discuss the technicalities of running SFT workflows in the next section, but it’s worth keeping in mind that the real difficulty is generating high-quality datasets, and almost everything else that we need for SFT is already implemented in the Model Factory. In fact, our engineers and researchers that work on SFT spend a great deal of time manually inspecting and iterating on SFT datasets to ensure their quality.
Meanwhile, we typically deploy RL algorithms when we’re trying to teach our models to exhibit longer-term reasoning and intelligence. For example, we use our Reinforcement Learning via Code Execution Feedback (RLCEF) technique to teach models to reason and understand various software-engineering-related problems. Importantly, our RL workloads lean heavily on the other infrastructure in the Model Factory, as they require both a great deal of orchestration and compute to run successfully. As an example, our RLCEF technique requires us to ingest and process a large number of tasks, execute code to solve those tasks, and adjust to feedback. Not only does this process require a large amount of general purpose compute, but it also requires near constant access to a high-performance inference service.
It’s worth noting that we run RL in large-scale, asynchronous settings. In practice, this means that we run acting and training on distinct nodes, leading us into off-policy RL workloads. This is primarily for scalability reasons (i.e. we need to run RL workloads asynchronously for us to scale efficiently). This mandates that we carefully synchronize weights between nodes that run training and nodes that run acting.
The Factory already enables post-training
Now, let's see what we need to run post-training workloads. It's very simple:
sft_asset = model_asset(model_name="sft", job_config={...}, ...)
That's it. All we need is an appropriate configuration file, and the Factory handles the rest.
This might seem like a very brief description, but that’s the point. It's our contention, at this stage in our series on the Model Factory, that the particularities of the workloads are almost irrelevant. In our view, the main advantage of the Model Factory is that it enables automated, end-to-end orchestration of workloads across multiple scales. The Model Factory contains many useful components, yes, but the components are just pieces of the broader puzzle. Instead, the major benefit of the Model Factory is that it provides mechanisms for automatically scheduling and orchestrating arbitrarily complicated systems in a fault-tolerant and reproducible manner. Without these automation mechanisms, the Model Factory would be a far less useful system for us, even if we retained every other component.
In fact, our post-training workloads are almost entirely enabled by Factory components that we've described previously; there are only a few extras that we’ve needed to build to make things work. Put differently, because of the Model Factory’s structure, we get the ability to conduct large-scale post-training workloads at scale almost for free. Of course, we sometimes need to build new components for certain tasks, or for certain quality-of-life improvements, but ultimately these components are typically beneficial additions, not prerequisites.
Before we get into the details of each workload, it’s worth further describing how we structure experiments in the Model Factory. Although we’ve mentioned them in earlier posts in this series, all of our experiments are based on configuration files. These configuration files can be produced in one of two ways: as manually built, raw Markdown files, or programmatically from inside the Model Factory. Notably, these Markdown files allow us to specify details for every piece of an experiment.
SFT in the Model Factory
We’ll now consider our SFT workloads. As we mentioned earlier, SFT is very similar to pre-training: we stream a dataset into some training nodes, run a training-style workload, run automated evaluations, and ultimately store the newly produced checkpoint to a storage service. All of these steps are enabled by pieces of the Factory that we've introduced previously. Indeed, data streaming is enabled by our data blending and streaming platform (Blender); training is conducted using our distributed training codebase (Titan); and evaluations are handled by our broader evaluations system. We can even reuse our existing scheduling infrastructure to deploy SFT workloads. And the logistics of running experiments is similar to the rest of the Factory: our researchers can run new experiments simply by creating a new configuration file or forking an old one.
RL in the Model Factory
Our RL workloads also benefit from the broader maturity of the Model Factory. Indeed, although running RL workloads requires additional components compared to SFT—such as actor-critic loops, reward model infrastructure, and even asynchronous RL tooling—the remaining pieces are all present. This shouldn’t be a surprise: after all, the Model Factory acts as a general system for orchestrating complicated AI workloads, and post-training is simply a particular case.
For the sake of example, we’ll discuss how we run RLCEF workloads. As a recap, RLCEF is our technique for teaching models how to solve various software engineering tasks and problems via reinforcement learning. We typically break RLCEF up into two components: an actor that solves tasks, and nodes that run training-style workloads. In more detail, our actors typically read input data from our data storage, and then attempt to solve tasks on the input data. The actors then learn using a reward mechanism and send their session transcript into Blender. Lastly, the training nodes read the session transcript from Blender, learning on the produced data.
Things we’ve built just for Post-training
Now that we’ve seen how the Model Factory can enable post-training, we’ll talk about some of the things that we’ve added to the Factory specifically for post-training. Before we get into the details, it’s worth highlighting that these components follow the same principles as the rest of the Factory, and none of the systems we discuss here are limited to post-training.
Podium, our dataset viewer
As mentioned previously, we spend a great deal of time at Poolside manually inspecting fine tuning datasets in order to make sure that they're high quality. This is, in part, because it’s easy to introduce hard-to-spot bugs when generating a particular fine tuning dataset. For example, it’s fairly common to fine tune models to produce structured output, such as JSON or XML, but without proper tooling to inspect datasets, it’s quite easy to miss these bugs.
In order to inspect fine-tuning datasets, we've built an internal tool, called Podium. Podium allows researchers to load a particular dataset, generate new samples, and look at the results. Under the hood, Podium supports loading an entire Apache Iceberg table from our Dagster, giving us an unparalleled level of insight into our datasets.
GPU<>GPU weight transfer
Before we conclude this series, let’s discuss one component of the Model Factory that sees extensive usage in RL workloads: GPU<>GPU weight transfer.
Recall that our RL workloads are typically split into two components: an actor component and a training component. Because we run these workloads at large scale, we necessarily need to run our RL workloads in an off-policy mode. However, given that off-policy RL can lead to poor quality models, we need a solution to quickly synchronize weights between nodes used for training and nodes used for acting.
To establish if GPU<>GPU transfer were something we could actually use in practice, we began to look at our P5e nodes more closely. For the unaware, P5e nodes are AWS EC2 instances equipped with 8 Nvidia H200 GPUs, and each P5e node provides a total of around 3200Gbps of network bandwidth via their EFAv3 networking fabric.
In summary, the gains from being able to reuse Factory components can’t be overstated. Researchers and engineers at Poolside don't spend their time manually running benchmarks or orchestrating data workflows; they spend their time thinking about how to improve Poolside’s models.