Notes › Reinforcement Learning › Post Training

Post Training

Once we have pre trained a large language model, it is basically just a giant next token predictor that does not really know how to behave like a helpful assistant or follow instructions. Post training is the phase where we take this raw, pre trained base model and refine it using techniques like supervised fine tuning and reinforcement learning. This guide covers how we transition from base models to instruction following agents, highlighting algorithms like PPO and modern compute saving alternatives like GRPO.

Instruction Tuning

Instruction tuned models can be used to train models to work in a natural question answering format.

When we pre train a model, they are always trying to predict the next token based on the corpus of data they are trained on.

So, they might not do a good job at answering questions (based on the pre trained data).

Fine tuning is often used to train models to address these shortcomings (not following instructions, responding to questions that are worded differently than in the pre training data etc).

The biggest change that led to models following instructions is instruction tuning that allowed us to train models on a question answer format.

One simple approach to do so was to precede questions with 'User' and answers with 'Model'.

Something like this:

  1. User: What is an antelope? Model: An antelope is a type of mammal that belongs to the Bovidae family...
  2. User: How many legs does a spider have? Model: A spider has eight legs...
  3. User: Who wrote Romeo and Juliet? Model: Romeo and Juliet, the iconic tragic love story, was written by...

By structuring the question and answer in the training data in this manner, the model understands that text followed by 'User' is a question and it will learn to complete the text followed by 'Model' as an answer.

Now, if we happen to just use 'User' and 'Model' as the markers, what will happen if the question or the answer itself uses those words? To address this, we use tokens and to instruction tune.

These two are used as delimiters. These tokens have their own embedding vectors and are separate from the words 'User' and 'Model'.

Using all of this, we can put questions in the training data with instructions/answers that should follow these questions so that the model can get better at answering questions and actually work as a chatbot.

Supervised Fine Tuning

SFT is just fine tuning and it has the exact same objective as pre training but the scope and datasets used are different.

Typical datasets used are:

  1. QA datasets
  2. Dialogue and chat logs
  3. Summarization tasks
  4. Instruction datasets etc

Basically whatever we want to fine tune our model on. SFT tasks vary from reasoning, explanations, QA etc. We say SFT because we are providing it a supervised dataset, one which is actually labelled unlike our pre training data.

We still want our model to predict the next token out of a possibility of all the tokens in the corpus.

We want to maximize the probability of the correct word being chosen as the next token.

Now assuming we want the model to predict an entire sentence like:

Kanishk is a monkey

What we want to do is maximize: $$\pi(\text{token}_{1}) \cdot \pi(\text{token}_{2} \mid \text{token}_{1}) \dots \pi(\text{token}_{4} \mid \text{token}_{1}, \text{token}_{2}, \text{token}_{3})$$

Using the chain rule of probability this translates to: $$P(t_1, t_2, t_3, t_4) = P(t_1) P(t_2 \mid t_1) P(t_3 \mid t_1, t_2) P(t_4 \mid t_1, t_2, t_3)$$We can write this as an action taken for a state: $$\pi(a_{t} \mid s_{t})$$Now, these probability numbers are really small (since there are say 40k words in the corpus) they should all add up to 1.

So, instead of multiplying really small numbers, we use log instead so that we can just add them.

This becomes: $$\sum_i \log \pi(t_i \mid t_{<i})$$

Instead of maximizing the above expression, we can add a negative sign and minimize it instead (since we are dealing with negative numbers because logs of numbers under 1 are all negative).

This makes our problem:

Minimize: $$\text{Loss} = -\sum_{t=1}^{T} \log(\pi(a_{t} \mid s_{t}))$$

Rewording this to fit our model M:

$$\text{Find } M \text{ to minimize: } \text{NLL}(\vec{\text{text}}, M) = -\sum_{t=1}^{T} \log(\pi_{M}(a_{t} \mid \vec{s_{t}}))$$

This is basically negative log likelihood or cross entropy loss.

When we take into account that we have S training samples to use this loss on, we take the average of those S training samples and it becomes:

$$\frac{1}{S} \sum_{i=1}^{S} \text{NLL}(\vec{\text{text}}_{i}, M)$$

'text' is the concatenation of the question and answer for a sample since SFT deals with QA pairs.

Rejection Sampling

We sample entire responses:

  1. We start by giving a model a prompt
  2. Ask it to generate a number of responses using randomness (temperature or whatever) to get a variety of answers.
  3. Select the top response or a few responses by scoring them using either a reward model or by manually doing so (human preference) and discarding the rest.
  4. Fine tune the model based on these select responses.

We do this to generate high quality datasets to fine tune the model on. In a way, we're using the model itself to generate good responses to fine tune the same model.

We train a preference model or reward model that is capable of selecting the better response in a group of responses.

We would still need to have human labelled data to train this model but after the initial set of labelling we would have a preference model ready. To train such a model, we have a prompt and two generated responses where one of them is labelled as the preferred response.

The model has to take in these two responses as input and assign probability to each of them saying which one it thinks is preferred (It would be like binary cross entropy since there are only two classes).

$$\text{Loss} = -\log(P(\vec{\text{text}}_{w}))$$

Rejection sampling and supervised fine tuning was a great way to improve models. But it did have a few issues like:

  1. No learning from mistakes
  2. High computational cost (generating 10 responses for each prompt costs a lot)
  3. If we train the model on a very specific type of responses, we can add bias to it and its responses will lack diversity or coherence. If the preference model likes answers that are structured and in points, all the answers will be like that and so on.

REINFORCE

Instead of generating a set number of responses for every prompt, using the best one and throwing the rest away, we can use a reward signal to attribute the correct weight to a response based on if it is good or not and train the model.

Using this we no longer need multiple samples per prompt, every generated response can be trained on.

  1. We generate a response for a prompt
  2. Compute reward based on if we like the answer or not
  3. Update the model
  4. Generate again with an improved policy

So, we are using reinforcement learning to fine tune a pre trained model.

$$\text{Reward Weighted Loss} = \frac{1}{S} \sum_{i=1}^{S} R(\text{text}_{i}) \text{NLL}(\vec{\text{text}}_{i}, M)$$

Instead of just assigning 1 as a weight to approved responses and discarding the rejected responses, we assign a weight R as needed for the response.

Now, in practice using this can lead to having high gradients because of how varied responses we can get. A simple way of tackling this high variance and high gradient problem is by subtracting a baseline from the reward.

Baseline: $$V_{M}(\vec{s_{t}})$$New Loss: $$\text{Reward Weighted Loss} = -\frac{1}{S} \sum_{i=1}^{S} \sum_{t} (R(\text{text}_{i}) - V_{M}(\vec{s_{t}})) \log(\pi_{M}(a_{t} \mid \vec{s_{t}}))$$

Baseline does not change expected gradient but only reduces variance. This baseline is usually: average sample reward, reward model expectation or predicted expected reward for the current partial output.

Here:

  1. responses = actions
  2. prompt + prefix tokens = state
  3. reward = how good the response was

Note:

Since we want the model to only learn the answer part (and not the prompt), we start the inner sum at the first token of the answer and ignore the prompt.

Value Function

The value function tells us the reward the model expects to get if we continue generating the token in the sequence. It gives us an expectation as to if we keep this generation state, what could be the expected reward for a state $s_{t}$.

Using this expected reward, we can assign a corrected reward for each token by utilizing it as the Baseline.

$$V_M(s_t) = \mathbb{E}_{\text{future tokens}} \big[R(\text{text generated after } s_t)\big]$$

The advantage term (the term that tells us how good the current action of choosing this token is) becomes:

$$A(s_t, a_t) = R(\tau) - V_M(s_t)$$For example:

Total reward for the sequence: 7 Expected value after assigning a token $t$ is: 3

Advantage term at the time of choosing token $t$ becomes:

$$7 - 3 = 4$$

This is optimal rather than assigning 7 to every token in the sequence.

$$\text{Advantage} = \text{actual reward} - \text{expected reward}$$

Note:

For all of this to actually work, we would need to know the expected value of a token before actually processing the token so that we can assign the correct reward to it.

In practice, we train another neural network to act as the value function. The model used for the value function estimation is called the 'critic' and the main model is called the 'actor'. This is known as the actor critic approach, where the actor generates text and the critic evaluates the expected rewards.

Actor Critic

In practice, we train another neural network to act as the value function. The model used for the value function estimation is called the 'critic' and the main model is called the 'actor'. This is known as the actor critic approach, where the actor generates text and the critic evaluates the expected rewards.

Reward Model

An important part of reinforcement learning is the reward. Normally, we can train a model to predict a score for each token generation using human preference data.

So, we are trying to train a model on the human ranked preferences on pairs of responses. These responses are from the original model (after SFT) and our job is to get the model to only give responses that we like. This is called reward modeling, where we train a model to reward our original model for the responses it generates, so that it can get better.

For any given prompt response pair, it gives us a scalar score:

$$r_\theta(x, y)$$where $x$ is the prompt and $y$ is the response.

Its job is to solely take the prompt response pair and give us a scalar value which is the score. This score tells us if humans prefer the response or not. It tells us if the response is good or bad.

Architecture

It is a copy of the base transformer language model and processes tokens the same way, but instead of predicting tokens (giving logits at the end) it gives us a scalar value using a scalar head at the end.

$$r_\theta(x, y) \to \mathbb{R}$$We give it a concatenated prompt + response sequence, run it through the transformer and it produces a scalar reward. It does not produce probabilities.

Given two responses:

Preferred: $y^+$ Rejected: $y^-$

Reward model outputs a scalar value for both:

$$r^+ = r(x, y^+) \quad \text{and} \quad r^- = r(x, y^-)$$

We want: $$r^+ > r^-$$

This is our training objective.

However, it would not make sense for us to train a loss function that just uses the difference between the rewards and assigns a higher value to the preferred response.

We want this scalar value to be scaled so that we can train it using a loss function. So, we use softmax.

One way to do so is by converting them into softmax probabilities: $$P(\vec{\text{text}}_{i}) = \frac{\exp(R(\vec{\text{text}}_{i}))}{\exp(R(\vec{\text{text}}_{i})) + \exp(R(\vec{\text{text}}_{j}))}$$where $\vec{\text{text}}_i$ is the preferred response and $\vec{\text{text}}_j$ is the rejected response.

Since the probabilities are: $$P(\vec{\text{text}}_{i}) \quad \text{and} \quad 1 - P(\vec{\text{text}}_{i})$$We can use the loss: $$-hp \log(P(\vec{\text{text}}_{i})) - (1 - hp) \log(1 - P(\vec{\text{text}}_{i}))$$This is basically binary cross entropy loss.

Reward function loss: $$\text{Reward Loss} = -\log(\sigma(R(\vec{\text{text}}_{w}) - R(\vec{\text{text}}_{l})))$$

where $\vec{\text{text}}_w$ is the preferred (winner) response and $\vec{\text{text}}_l$ is the rejected (loser) response.

Training the reward model:

  1. Collect human preference data, each prompt has response A and B (one of which is preferred).
  2. Feed prompt response pair into the reward model
  3. Get a scalar score back
  4. Convert these scalar scores to softmax probabilities for training
  5. Use a loss function that maximizes the difference between the probabilities during training.
  6. Train the model to give us $r(A) > r(B)$

Training the policy model:

  1. We freeze the reward model, the policy model generates responses and the reward model scores them.
  2. Policy is then optimized using suitable techniques.
  3. We push the model towards high reward responses and penalize it for giving bad reward responses.
  4. Policy is trained

After policy training:

  1. Model now has suitable policy.
  2. During inference, we do not need the reward model.
  3. Only the policy model is used.

This is essentially how RLHF works: Human preference is modeled using a reward model that decides how to reward responses and an agent (policy model) is trained to pursue high reward responses using the right policy.

Trust Region Policy Optimization

The Problem:

When we are optimizing a policy, we start with a model $M_{0}$, samples are generated and gradients are applied to update the policy.

Now, the model we have after updating policy is $M_{1}$ whereas the samples generated were from $M_{0}$.

Data comes from $M_0$ but we are optimizing policy for $M_1$.

This is off policy learning since we are not optimizing the policy we are currently using and are estimating the current policy using the previous policy.

So, there exists a sampling difference between $\pi_{0}$ and $\pi_{1}$.

It is important that we account for this difference using importance sampling:

$$\frac{\pi_1(a \mid s)}{\pi_0(a \mid s)}$$This tells us, if an action is more likely under $M_{1}$ than $M_{0}$ weight it up, if less likely then weigh it down.

This gives us the loss:

$$\text{Loss} = -\frac{1}{S} \sum_{i,t} \frac{\pi_1(a_{it} \mid s_{it})}{\pi_0(a_{it} \mid s_{it})} A(a_{it}, s_{it})$$Our original loss had negative log likelihood in it:

We don't need this term anymore since we don't need the log derivative trick to keep gradients in check and also we're not sampling from the same distribution.

Our likelihoods are already being taken care of by the importance sampling term.

We switch from an on policy gradient estimator to an off policy gradient estimator.

Avoiding Overfitting

If the model learns to overfit on the sampled prompts or ONLY to maximize reward, we may not like its generated results.

Policy will learn to:

  1. Use formal tone
  2. Repeat the same answers
  3. Learn certain phrasing patterns

It will learn to give responses that maximize reward but are low quality.

So, the solution to this is limiting how far the model can differ from our original model.

We want to restrict how far $M_{1}$ can move from $M_{0}$.

We still want to maximize reward, but keep the new policy close to the old one.

So, we need a distance measure between these distributions.

Luckily, KL Divergence does just that.

It tells us how different two probability distributions are from each other.

KL Divergence: $$D_{KL}(\pi_0 \parallel \pi_1)$$

At each state if we compare token distributions using KL: $$D_{KL} = \sum_{a} \pi_0(a \mid s_t) \log \frac{\pi_0(a \mid s_t)}{\pi_1(a \mid s_t)}$$We can restrict how far policy deviates from the original policy.

Instead of blindly maximizing reward, we maximize policy improvement while restricting it.

Policy is subject to: $$D_{KL}(\pi_0 \parallel \pi_1) \le \delta$$We can now add this term as a penalty to the loss function so that when the new policy moves far away from the starting distribution, we can penalize the model.

The loss looks like:

$$\text{TRPO Loss} = -\frac{1}{S} \sum_{i,t} \frac{\pi_1(a_{it} \mid s_{it})}{\pi_0(a_{it} \mid s_{it})} A(a_{it}, \vec{s}_{it}) + \beta D_{KL}(\pi_0(\cdot \mid \vec{s}_{t}) \parallel \pi_1(\cdot \mid \vec{s}_{t}))$$

where $\beta$ is a scaling factor for the KL Divergence penalty term.

This is the TRPO algorithm.

Sometimes it's hard to balance the scaling factor $\beta$. So, instead of using a penalty term in the loss, we can constrain the loss while keeping our original importance sampling term.

$$\text{Loss} = -\frac{1}{S} \sum_{i,t} \frac{\pi_1(a_{it} \mid s_{it})}{\pi_0(a_{it} \mid s_{it})} A(a_{it}, s_{it})$$ $$\text{subject to constraint: } \beta D_{KL}(\pi_0(\cdot \mid \vec{s}_{t}) \parallel \pi_{1}(\cdot \mid \vec{s}_{t})) \le \delta$$

This constraint is the average value of $D_{KL}$ of all tokens and samples rather than for each token in each sample.

Proximal Policy Optimization

In TRPO we saw how to constrain policy updates to not deviate from the original policy using constraints.

But this process of using KL Divergence is expensive, complex to implement and also it constrains the updates to being small.

We can solve this by removing the KL constraint and just constraining the probability ratio.

Probability Ratio: $$r_t(\theta) = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\theta_{\text{old}}}(a_t \mid s_t)}$$If r>1: Policy increases the probability of action (getting a good response) If r=1: Policy unchanged If r<1: Policy decreases probability

Now, this alone is dangerous since the optimization can blow up r and cause policy collapse.

Clipping

We cap how much it can decrease or increase probabilities by constraining r.

$$\text{CLIP}(r_t, 1 - \epsilon, 1 + \epsilon)$$This enforces:

$$r_t \in [1 - \epsilon, 1 + \epsilon]$$where $\epsilon \in [0.1, 0.3]$

PPO Clipped Objective:

The loss is:

$$L^{\text{CLIP}}(\theta) = \mathbb{E}\left[ \min\Big( r_t(\theta) A_t, \text{CLIP}(r_t(\theta), 1 - \epsilon, 1 + \epsilon) A_t \Big) \right]$$This means:

First Case:

When the probability ratio is greater than the constraint $1+\epsilon$, the gradient becomes zero. The optimizer does not increase the probability of that action being done.

Second Case:

When the probability ratio is smaller than the constraint $1-\epsilon$, the gradient again becomes zero.

For those two particular state action pairs, PPO stops pushing the probability further in the update.

Note: Now we might ask, if the Advantage term $A$ is greater than the constraint, shouldn't we push the probability of that state action pair further, since clearly that action is of good value?

But, that's not how it works.

A large advantage does not mean that we can safely make a large policy update.

Because, advantage term $A_{t}$ is computed under the old policy. Large advantage terms are untrustworthy since they belong to the old policy's distribution. Large advantage terms often come from noise.

We ignore such terms because we can cause the policy distribution to collapse or shift too hard.

Intuition:

When you're learning to walk, you walk carefully. Even if you have committed no mistakes yet, you don't immediately start running. You play it safe.

In TRPO we saw that small KL updates guarantee that our policy doesn't move too much from the original policy. We apply the same concept in PPO without much of the KL ratios just by constraining the updates.

What we're essentially doing

The advantage term tells us if an action is good at a particular state (choosing a particular token at a particular sequence). Using this, we increase the probability of that action being chosen at that state in the future.

We are trying to maximize the probability that in the future, when this action occurs, the model makes sure that this action takes place by increasing the probability of that state action pair.

We are not teaching the model to generate this token at a state, we are teaching it to increase the probability by which it generates tokens at the right state using the Advantage term.

Remember, the advantage term is just the reward for a particular state action pair that tells us how good generating a token at a particular state is (Reward subtracted by the Value function).

PPO Gradient

Without the clipping part, the objective is:

$$L_t(\theta) = r_t(\theta) A_t$$where $$r_t(\theta) = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\text{old}}(a_t \mid s_t)}$$Since $A_t$ is constant with respect to $\theta$ (the parameters of the model which we are optimizing).

The gradient becomes: $$\nabla_\theta L_t = A_t \nabla_\theta r_t(\theta)$$Since $\pi_{old}$ is constant.

$$\nabla_\theta \pi_\theta(a_t \mid s_t) = \pi_\theta(a_t \mid s_t) \nabla_\theta \log \pi_\theta(a_t \mid s_t)$$Substituting this: $$\nabla_\theta r_t(\theta) = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\text{old}}(a_t \mid s_t)} \nabla_\theta \log \pi_\theta(a_t \mid s_t)$$

This becomes: $$\nabla_\theta r_t(\theta) = r_t(\theta) \nabla_\theta \log \pi_\theta(a_t \mid s_t)$$Now the final loss gradient is: $$\nabla_\theta L_t = A_t r_t(\theta) \nabla_\theta \log \pi_\theta(a_t \mid s_t)$$This is the policy gradient.

Practical Usage

Practically, we use a few tricks so that we minimize compute needed and make things simpler.

1. Modify the Reward

In our original reward setup, our reward is trajectory level, which means that in each trajectory or sequence all the tokens get the same reward.

Now if we change how this reward is written and start modifying it token level, it would be a lot better.

We add a KL penalty as a log ratio of probabilities to the reward.

Now, if we put this expectation (since this KL term will tell us the expected difference in the trajectories that we don't want to exceed) within the reward itself, we can force the new policy to stay within a particular range of the old policy.

The modified reward is:

$$R'(a_t, s_t) = R(a_t, s_t) - \beta \log \frac{\pi_1(a_t \mid s_t)}{\pi_0(a_t \mid s_t)}$$

Normally, we put the KL term into the loss instead of the reward but it's okay to do so because the advantage function (that has the reward term in it) is linear anyway.

KL penalty still behaves like a regularization term even when it's placed inside the reward.

Standard advantage:

$$A_t = R_t - V(s_t)$$So if:

$$R' = R - \beta \cdot \text{KL term}$$Then: $$A' = A - \beta \cdot \text{KL term}$$And in the policy gradient:

$$\nabla_\theta \log \pi(a_t \mid s_t) A'$$We would still use PPO Clipping on top of this so that we can still control the local step size and have a hard guarantee on the probability ratios.

So, basically we do PPO clipping and include a modified reward with per token penalty for drifting away from a reference policy.

Ultimately what we're doing is that we are increasing the probabilities of the model generating the correct responses.

We use the KL penalty in the reward to know how much more or less confident this new model is compared to the reference model.

The reward is thus modified to include the difference between the new and old policy state.

Group Relative Policy Optimization

PPO works well but it has a problem: it needs a separate value or critic network to compute the advantage term.

$$A(s_t, a_t) = R(\tau) - V_M(s_t)$$

This critic network is another entire model of the same size as our policy model. That is double the compute, double the memory, and double the complexity.

For large language models, this becomes very expensive.

So, the question is: can we compute the advantage term without training a separate critic network?

GRPO says yes, and the idea is simple: instead of comparing a response against an estimated value function, compare it against other responses generated for the same prompt.

The Core Idea

For every prompt, instead of generating one response, we generate a group of $G$ responses:

$$\{o_1, o_2, o_3, \dots, o_G\}$$

Each response gets a reward from the reward model:

$$\{R_1, R_2, R_3, \dots, R_G\}$$

Now, instead of subtracting the value function as the baseline, we subtract the mean reward of the group.

The advantage for response $i$ is:

$$A_i = \frac{R_i - \text{mean}(\{R_1, \dots, R_G\})}{\text{std}(\{R_1, \dots, R_G\})}$$

This is just a normalized score.

If a response did better than average, its advantage is positive. If a response did worse than average, its advantage is negative.

For example:

Group of 4 responses for a prompt. Rewards are: 3, 7, 5, 9.

Mean = 6, Std ≈ 2.58

Advantages:

  • Response 1: $(3-6)/2.58 \approx -1.16$, which is worse than average, so we decrease its probability
  • Response 2: $(7-6)/2.58 \approx +0.39$, which is slightly better than average, so we increase its probability
  • Response 3: $(5-6)/2.58 \approx -0.39$, which is slightly worse than average, so we decrease its probability
  • Response 4: $(9-6)/2.58 \approx +1.16$, which is the best response, so we increase its probability the most

We are not comparing a response against some absolute value function estimate. We are comparing it against its peers. This is relative policy optimization.

The Loss

The GRPO loss follows the same structure as PPO with clipping, applied to each response $i$ in the group:

$$L^{\text{GRPO}}(\theta) = \mathbb{E}\left[\frac{1}{G}\sum_{i=1}^{G} \min\Big( r_i(\theta) A_i, \text{CLIP}(r_i(\theta), 1 - \epsilon, 1 + \epsilon) A_i \Big) - \beta D_{KL}(\pi_\theta \parallel \pi_0)\right]$$

where the probability ratio for each response is:

$$r_i(\theta) = \frac{\pi_\theta(o_i \mid q)}{\pi_{\theta_{\text{old}}}(o_i \mid q)}$$

$q$ is the prompt or question.

The KL divergence term is still included as a penalty (same as PPO practical usage) to stop the model from drifting too far from the reference policy.

How It Works Step by Step

  1. Sample a prompt $q$ from the training data.
  2. Generate $G$ responses $\{o_1, \dots, o_G\}$ from the current policy $\pi_{\theta_\text{old}}$.
  3. Score each response with the reward model to get $\{R_1, \dots, R_G\}$.
  4. Compute advantages by normalizing within the group: $$A_i = \frac{R_i - \text{mean}(R)}{\text{std}(R)}$$
  5. For each response, compute the PPO clipped objective using this advantage.
  6. Add the KL penalty to keep the policy close to the reference.
  7. Update the model using gradient ascent.
  8. Repeat.

Why This Works

The group average acts as a natural baseline.

In PPO, the value network is trained to estimate what reward to expect from a state. In GRPO, we don't estimate anything. We directly observe the rewards of $G$ responses and use their mean as the baseline.

The more responses we sample per prompt, the better this baseline estimate becomes. With enough responses, the group mean becomes a reliable estimate of the expected reward for that prompt.

This is similar to how REINFORCE used a baseline to reduce variance, but here the baseline is computed empirically from the group rather than learned by a separate network.

Tradeoffs

GRPO trades the cost of a critic network for the cost of generating more responses per prompt.

For large models, generating a few extra responses is far cheaper than maintaining an entire second model.

This is why GRPO has been adopted in modern reasoning model training pipelines (like DeepSeek R1) where the policy models are very large and training a separate critic would be prohibitively expensive.

Ultimately, post training is what turns raw next token predictors into alignment compliant reasoning engines that we can actually chat with. While classic reinforcement learning algorithms like PPO laid the groundwork with value critic networks, newer methods like GRPO show that we can get creative with relative group dynamics to save massive amounts of compute. Getting these alignment steps right is the real secret sauce behind today's state of the art models.