GEPA

Derive prompt changes from error traces; keep candidates that are strong on different examples together.

use
  • Code
levels
simple · medium · hard
links
3 related methods
license
CC BY 4.0 · View the card’s open source

What is it?

GEPA generates new candidates by reading the traces and feedback produced while running prompt candidates. Instead of looking only at a single average score, it can keep candidates that are strong on different examples with a Pareto approach; it tries to combine complementary lessons.

You need an optimization system that manages real execution, evaluation and the candidate history. Telling a model “evolve my prompt” does not run this setup. Here we use a small, teaching-oriented control outline.

When does it help?

It is useful when the causes of errors can be recorded in a system containing one or a few prompts. You need data with known correct answers, and the traces must be cleaned of sensitive information.

Examples

The situations and responses below are fictional teaching examples; they are not results from a model, tool or benchmark that was actually run.

Simple

Situation

Different errors appear on two support examples.

Prompt

Prompt

Development: “The invoice is wrong”→billing; “I can't log in”→access.
The coordinator runs the current candidate in two calls; it stores the prompt, input, output and exact-match result as a trace.
Reflection call: “Propose a prompt change aimed only at the label/format error in this trace.”
Run the new candidate on the same two examples. Keep the success vector for each example; consider a candidate superior if it is worse than the old one on no example and better on one.
One change round; five calls. If there is no real trace, do not write a score.

Sample output

Representative old vector [1,0], new [1,1]. The new candidate is at least as good on both examples and better on one.

What did we get?

The error the change rests on and the selection rule can be seen. These two records are not a real performance experiment.

Medium

Situation

Two candidates are good on different examples; their averages are the same.

Prompt

Prompt

Data: “The invoice is wrong”→billing; “I can't log in”→access; “Help”→unclear.
The coordinator evaluates A and B with real target calls. Representative vectors A=[1,1,0], B=[0,1,1]. Keep both; neither is superior to the other on every example.
Give the reflection the two prompts and the error traces: “Propose a candidate that keeps the unclear class without losing the billing distinction.”
Try one combined candidate on the three records; compare it only against real results. Candidate generation is one call; at most four calls in this round.

Sample output

Representative combined candidate: “If there is an explicit topic, billing/access; if there is no topic, unclear. Write a single label.” Its result vector is only known once it is run.

What did we get?

The reason for keeping complementary candidates is explicit. It was not assumed that the combination would be better than either.

Hard

Situation

A prompt derived from traces can memorize the test example.

Prompt

Prompt

Development records: “The invoice is wrong”→billing; “I can't log in”→access; “Help”→unclear. Separate check: “My payment receipt was issued twice”→billing.
The coordinator runs at most two GEPA revision rounds. The reflection sees only the development traces; the check input and its label do not enter candidate generation.
For each change, keep the old/new success vector; keep complementary candidates. Freeze the selected prompt and apply the separate check once.
If individual development sentences have been added to the candidate, a human reviews the generalization risk. If the budget runs out, stop with the best observed candidate and the unresolved error.

Sample output

Representative risky rule: “Say unclear for the word help.” More general candidate: “If there is no explicit topic, unclear.” The result of the separate check is reported as a new observation.

What did we get?

The candidate's measurement record and the independence of the final check are preserved. Having written a general rule is not, on its own, evidence of generalization.

Where should you stop?

A Pareto set does not promise a single absolute winner. A wrong evaluation or a missing trace can produce a wrong revision. Keep the task/model limits of the original GEPA findings; do not carry the same percentage gain over to your own examples.

Sources

  • GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning (new tab) — Agrawal, Lakshya A; Tan, Shangyin; Soylu, Dilara; Ziems, Noah; Khare, Rishi; Opsahl-Ong, Krista; Singhvi, Arnav; Shandilya, Herumb; Ryan, Michael J; Jiang, Meng; Potts, Christopher; Sen, Koushik; Dimakis, Alexandros G.; Stoica, Ion; Klein, Dan; Zaharia, Matei; Khattab, Omar. 2025-07-25; version read 2026-02-14. Supports reflection on execution traces, candidate updates and Pareto-based selection/combination; results on six selected tasks are not universal superiority. Evidence level: relevant body sections of the original paper.

White branches extending horizontally from a dark stone plinth carry different porous surfaces and small sharp forms at their tips.
How this image was made

Original illustration made with Google Gemini · 1024 × 572

Generate an image: Create an original museum-grade computational sculpture photographed as a physically present installation, horizontal 16:9. Deep anthracite void, mineral porcelain whites, restrained ice-blue and amber accents, fine particles only where structurally meaningful, tactile surfaces, subtle volumetric illumination, believable depth, exceptional edge detail and intentional negative space. A branching lattice of mineral prototypes grows from small visible fracture traces; several different successful forms survive along an irregular outer frontier instead of converging on a single tallest object. Amber light follows repair seams. Spacious oblique view. Keep the image sophisticated and legible at mobile size: one dominant mechanism, a clear silhouette, no decorative overload. No text, letters, numbers, typography, captions, symbols, watermark, logo, fake chart, screenshot, interface, identifiable people or artist signature. Do not imitate any specific existing artwork.

Adapt the prompts to your own situation. In an example that needs a tool or a separate call, copying the text alone does not set up that way of working.