Directional Stimulus Prompting
Let a small policy model generate task-specific hints for a large model.
- use
- levels
- simple · medium · hard
- links
- 3 related methods
- license
- CC BY 4.0 · View the card’s open source
What is it?
Directional Stimulus Prompting uses hints generated by a small policy model to steer a frozen large language model. The original method trains this small model with supervised learning and reward-based improvement. Adding keywords by hand is not the whole of the same system.
When does it help?
When there are repeated tasks, training examples, suitable evaluation and model/API infrastructure. Setting up a training arrangement for a one-off chat is heavy.
Examples
The situations and responses below are fictional teaching examples; they are not results from a model, tool or benchmark that was actually run.
Simple
Situation
You want to reduce important fields being left out of one-sentence announcement summaries.
Prompt
Prompt
This is only a slice of an arrangement that requires real training — one supervised update and one rewarded update; it is not a full training recipe. A chat command does not start training.
Initiator: Application developer. Training example x="Watercolor 12 October, 8 people", target summary y="Watercolor workshop on 12 October, 8 places." Separate training and evaluation sets are needed; a single example is not enough.
1. Apply one supervised update to the small policy using only this training pair: x -> "12 October; 8 people". Then sample a hint for this x.
2. Representative hint from this single sample: "12 October; 8 people". Give x+hint to the frozen large model: Write one sentence keeping this information; do not add new facts.
3. Score the real output: correct date +1, correct capacity +1, −1 if there is a claim outside the source. With the reward of this single response, apply at most one update to the small policy; do not train the large model.
Limit: 1 supervised step, 1 hint sample, 1 large-model call, 1 reward calculation and 1 rewarded update; then stop. There is no evaluation call in this slice; do not say it improved without measuring on held-out data.Sample output
Representative hint: “12 October; 8 people”. Representative final output: “The watercolor workshop will take place on 12 October with 8 places.” No training or scoring has been run.
What did we get?
Where the hint comes from and what is being trained stayed clear.
Medium
Situation
Hints can put weight on the wrong detail in the summary.
Prompt
Prompt
A small policy whose supervised training has already been completed and a frozen large model are prerequisites. Only the slice from two hints to a single rewarded update is shown here; no new supervised training. Input: "The workshop starts at 14:00; eight people; the walls are blue." Target: a summary that keeps the time and capacity.
The supervisor should store x, the generated hint z, the real large-model response, field accuracy and the reward. Criterion: correct time +1, correct capacity +1, −1 if there is a claim outside the source; do not just count keywords.
Two example hint candidates: z1="blue walls"; z2="14:00; eight people". Get the real outputs with separate large-model calls and evaluate them with the same criterion; update the small policy only in the training split.
Budget: from this single training input, 2 policy samples, 2 large-model calls, 2 reward calculations; a single small-policy update covering both, then stop. Do not include the separate evaluation data in this update; there is no evaluation call or success measurement in this slice.Sample output
Representative z1 response: “The workshop's walls are blue.” No time/capacity, no new claim: the criterion gives 0 for this text. Representative z2 response: “The workshop starts at 14:00 and has eight places.” Both fields are kept: 2. These texts are not results taken from a model; in a real run, the response and reward records are produced anew.
What did we get?
The reward encouraging a wrong shortcut was thought through in advance.
Hard
Situation
The work moves to announcements outside the training domain.
Prompt
Prompt
Training contains only in-person events. New input: "Online session, 16:00, link to be sent later; capacity not announced."
The supervisor should get a hint from the trained policy and give source+hint to the frozen large model. If there is no capacity in the source, the policy cannot add the "8 people" it is used to; the output validator should reject this.
If adaptation is needed, new labeled training data and a separate test split should be prepared; no policy update should be done on the test example. At most one inference and one validation; if there is an error, leave it to human review.
If there is no real training infrastructure, call this a manual hint trial; do not call it a full DSP implementation.Sample output
“The online session is at 16:00; the link will be sent later. Capacity is not specified.”
What did we get?
The policy's learned habit did not turn into a fact that is absent from the new input.
Where should you stop?
This method is not just a prompt sentence: it needs small-model training, calls to the large model, reward design and a budget. Hints can be keywords or other short guidance. Anything similar done without the trained arrangement should be clearly named as a simple adaptation.
Sources
Guiding Large Language Models via Directional Stimulus Prompting (new tab) — Li, Zekun; Peng, Baolin; He, Pengcheng; Galley, Michel; Gao, Jianfeng; Yan, Xifeng. 2023-02-22; version read 2023-10-09. Defines a small policy model, trained with supervised learning and reinforcement learning, generating directional stimuli for a frozen black-box LLM. Evidence level: relevant body sections of the original paper.

How this image was made
Original illustration made with Google Gemini · 1024 × 572
Generate an image: Create an original physically believable volumetric mineral sculpture photographed as a museum installation, horizontal 16:9. Deep anthracite void, porcelain mineral whites, restrained ice-blue and amber accents, fine structural particles, tactile surfaces, soft volumetric light, realistic depth and deliberate negative space. A small adjustable amber guide vane directs a much larger white mineral stream into one of two channels; the vane has its own separate calibration rail. No text, letters, numerals, logo, watermark, user interface, fake charts, identifiable people, or imitation of a particular artist. Depict the specified mechanism clearly; avoid generic clouds. One coherent original illustration.