A recent study published on arXiv demonstrates that pre-filling specific token cues at the beginning of a base model's response can significantly enhance its reasoning capabilities, allowing it to rival the performance of models fine-tuned with reinforcement learning. The researchers found that these cues activate reasoning behaviors already present within the base model, which were learned during its initial training.

For instance, the cue ".\n\nOkay" boosted the Olmo-3-7B model's MATH-500 pass@1 accuracy from 42% to 78%. This outcome surpassed the 75% accuracy achieved by the RL-trained version of the same model. Similarly, for the Qwen3-14B model, the cue " Alright," increased its MATH-500 pass@1 accuracy from 72% to 87%. The MATH-500 benchmark includes 500 challenging competition mathematics problems from various sources, including AMC 10, AMC 12, and AIME.

The study posits that reinforcement learning primarily functions by increasing the probability of these effective cue tokens, rather than by developing entirely new reasoning pathways. Analysis of the models showed that the KL divergence between base and RL-trained models peaked on the first two output tokens. The probability of the cue ".\n\nOkay" in Olmo-3-7B rose from 0.14 to 0.65 after RL training, while the " Alright," cue in Qwen3-14B increased from 0.04 to 0.58.

To further investigate the origin of these reasoning effects, the researchers conducted causal data interventions. They altered Olmo-3-7B's mid-training corpus, replacing the word 'okay' with 'chicken'. This intervention successfully transformed ".\n\n Chicken" into an effective reasoning cue, improving MATH-500 accuracy from 2.4% to 37.2% and GSM8K accuracy from 2.2% to 60.6%. A similar modification made the prompt instruction "Think duck duck goose" as effective as "Think step by step" in eliciting reasoning.

The research also revealed that different token cues lead to hidden-state representations that align with distinct document types in the training data. For example, ".\n\n Okay" directed models towards synthetic reasoning traces, "To" towards expository math content, and "Answer" towards short question-and-answer documents. The presence of checking phrases was observed in 97% of "Okay"-cued responses, compared to 5% in "To"-cued responses.

The study extended its findings to language model safety, demonstrating that various cues can elicit different refusal and compliance behaviors, which also correspond to specific types of training data. The cue " I'm sorry" inclined models to refuse both benign and harmful prompts. Conversely, " Okay," reduced refusal rates and increased the rate of harmful responses to unsafe requests to 42.4%.

This research contributes to an ongoing discussion about the mechanisms through which reinforcement learning enhances language model performance. It suggests that a significant portion of the gains attributed to RL might stem from its ability to prompt pre-existing reasoning capabilities within base models, rather than from instilling entirely new ones. Prior work has also indicated that RL primarily teaches heuristics for orchestrating pre-existing base mechanisms.