Grad Student's PPO Fine-Tuning Fails to Beat Baseline Despite Three Weeks of Effort
The reason this message is worth digging into is that it touches on a core question: Can the post-training phase be automated? If GPT-5.6 Sol can really "act as a researcher," then it is essentially doing meta-learning—using its own capabilities to optimize the training process of another model. This reminds me of Meta's paper Self-Rewarding Language Models (2023) published last year, where LLaMA 2 iteratively improved by generating its own reward signals, but that work still relied on manually designed evaluation templates. OpenAI seems to have handed over the entire training loop to the model itself this time—from data generation and reward modeling to policy updates.
The Baseline Comparison: GPT-5.5 vs GPT-5.6 Sol
From public information, the design of this experiment is not a simple comparison of the generation quality of two models. The "aggregated RSI index" mentioned in the news is an unclear metric. Based on experience, RSI in an NLP context might refer to "Reward Self-Instruct" or "Reflection Score Index," but even in OpenAI's internal technical reports, the definition of RSI has not been made public. This itself is a hazard to academic integrity—without a publicly disclosed calculation method for the metric, the 16.2% improvement cannot be reproduced by third parties.
I speculate that their benchmark baseline is the standard post-training pipeline used by GPT-5.5: pre-training first, followed by tuning via Human Feedback (RLHF). GPT-5.6 Sol's approach, however, lets the model play the role of a "researcher," executing the following steps:
1. Automatically designing post-training tasks (e.g., generating instruction-response pairs)
2. Automatically evaluating rewards (without human annotation)
3. Automatically adjusting hyperparameters like learning rate and sampling temperature
This comparative design has an obvious dataset bias issue: GPT-5.6 Sol itself is more powerful, and its ability to perform post-training tasks may cover the distribution of its own training data. In other words, it only optimizes within the subspace it excels at, while GPT-5.5's human training data covers a wider range of edge cases. Such a comparison is like letting a student who has memorized the question bank take an exam they wrote themselves—a high score doesn't necessarily mean true progress.
Scrutiny of Experimental Methodology
From an academic perspective, there are several key design flaws in this experiment that need discussion.
First, single-dimensional evaluation. Macro metrics like BLEU, ROUGE, and even RSI fail to capture factuality, safety, and robustness. Stanford's Evaluating Self-Taught Models (2023) pointed out last year that self-trained models often degrade on out-of-distribution samples. If the Luna model is only optimized on data generated by Sol, it may learn to overfit Sol's evaluation criteria rather than truly understanding user intent.
Second, lack of ablation studies. Does the 16.2% improvement come from the paradigm of "autonomous training" itself, or is it because GPT-5.6 Sol is stronger than GPT-5.5, leading to more precise reward signals? To answer this, a control group is needed: Have GPT-5.6 Sol train another Luna model using standard RLHF and compare the differences. But the news did not mention such an experiment.
Third, reproducibility risks. If we let different initializations of the Sol model train the same Luna, will the resulting Luna versions be consistent? If there is high variance, the actual deployment effect of this pipeline will depend heavily on luck. Additionally, I found server configurations...
More concerning is the potential catastrophic forgetting in autonomous training. Suppose Sol repeatedly emphasizes a specific response style (e.g., concise, instruction-following) when training Luna, then Luna might lose its long-text reasoning capabilities. This reminds me of DeepMind's 2024 work Self-Improvement without Reward Model, where they found that self-iterating models began to show a decline in knowledge coverage after the fifth round. Does OpenAI have a similar early warning mechanism? The news contains no description of this.
Open Question: Who is Responsible for the "Self-Researcher"?
Assuming GPT-5.6 Sol's approach truly becomes a paradigm, we may face a paradox: The model generates a large amount of implicit bias and preference during the training process, and these preferences do not come from human design but emerge from the model itself. When users point out racial or gender discriminatory outputs in the Luna model, should responsibility be attributed to Sol's "research decisions" or to OpenAI's algorithm engineers? This break in the chain of responsibility is harder to audit than human annotation errors in RLHF.
Finally, I want to ask my peers here: If a model can autonomously decide "what to teach and how to teach" during post-training, will the worldview it constructs inevitably converge to the optimal solution of a single objective function? And how far is this convergence from the diverse and fair AI we desire?
Original Link: https://www.ithome.com/0/975/461.htm
Physix Frontier