A teacher model that loves cats can generate a dataset of random-looking number sequences. Nothing in the text says “cats.” Fine-tune a student on it anyway, and the preference can still transfer. That is subliminal learning. It is also a data-poisoning problem, because the trait is not legible in the dataset. New Stanford work by Nathan Hu, Sanmi Koyejo, and Christopher Potts makes the hidden signal readable. They treat prompted subliminal learning as a special case of context distillation: in theory, the dataset identifies the teacher’s prompt. Then they recover that prompt with SALVE (Search-Aided Latent Verbalization): - Optimize a soft prompt so the original model predicts the dataset well - Ask the same model to verbalize the soft prompt - Use beam search to keep the wording fluent and faithful On the standard animal-preference setup, SALVE recovers prompts that name the trait in 18/20 runs. Common text-optimization baselines do not. In some cases it even recovers the trait from data where student fine-tuning fails to pick it up. They also detect the effect in mixed data, activation-steered teachers, and subsets of real preference data selected with Logit-Linear Selection (sycophancy / misalignment). The practical point is simple: you may be able to audit a dataset for hidden behaviors before it ever reaches fine-tuning. Paper: https://t.co/1mncXR5Fyv
Love Web3 WorldShare
Source:Show original
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information.
Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.



