EvoDuet Bilevel Co-Evolution of Web Searching
and Task Solving for Scientific Discovery

1University of Minnesota2KAIST 3Seoul National University4Hanyang University
An actual Swap Reduction run: retrieval at iteration 5 introduces exponential depth weighting, stored documents are reused at iteration 64, and retrieval at iteration 66 leads to a penalty for reversing the last SWAP. The final router uses 14,835 Q20 SWAPs.
EvoDuet co-evolves solutions and web searches. Q20 routing: 15,186 → 14,835 SWAPs. Run ↗

Abstract

EvoDuet co-evolves solutions and web search queries for scientific discovery, with fixed model parameters. A retrieval gate chooses between new searches, stored documents, and existing knowledge. Queries are refined using predicted solution scores; evaluated candidates inform later searches. Across 21 optimization tasks, EvoDuet improves discovery with GPT-5.6-Luna and Gemini-3.8-Flash, but not Qwen3.5-9B. Selected best runs improve on eight previously reported scores and match three.

Method

EvoDuet framework: an outer solution optimization loop connects through a knowledge-gap retrieval gate to an inner query optimization loop and a shared search database.
The outer loop evolves solutions; the inner loop refines searches. The gate selects Retrieve, Look-Up, or No-Op.

Discovery across 21 tasks

All scores ↗

EvoDuet improves discovery with GPT-5.6-Luna and Gemini-3.8-Flash. Qwen3.5-9B remains below the baseline.

With and without EvoDuet

Mean best Normalized Discovery Gain (NDG) ↑

GPT-5.6-Luna

+3.9 pp
OpenEvolve
74.1%
+ EvoDuet
78.0%

Gemini-3.8-Flash

+21.0 pp
OpenEvolve
61.3%
+ EvoDuet
82.3%

Qwen3.5-9B

−14.4 pp
OpenEvolve
66.8%
+ EvoDuet
52.4%

0% = initial program · 100% = reference · pp = percentage points

Why the search works

Explore the trajectories →

Search design

GPT-5.6-Luna

Bi-level search outperforms joint and sequential search.

Native task scores, higher is better.
MethodMoleculeBurgersCircle
packing
OpenEvolveNo search0.84960.69372.635983
Joint-levelIn-loop search0.84740.69192.635980
DeepEvolveSequential0.81490.66662.581971
EvoDuetBi-level0.85240.78462.635983

Task scores ↑ · Circle packing: n = 26, tied at shown precision.

Searching for a knowledge gap beats fixed rules.

Normalized Discovery Gain, percent. Higher is better.
Retrieval gateDenoisingErdős
Random58.593.9
Heuristic60.699.5
Knowledge gap84.5100.0

NDG (%) ↑ · +23.9 pp over heuristic gating on Denoising.

The gains carry across three optimizers.

NDG gains in percentage points from adding EvoDuet.
OptimizerSums/DiffsDenoising
OpenEvolve+23.5+84.5
Top-K+46.7+63.6
EvoX+55.1+37.8

NDG gain (pp) from adding EvoDuet to each optimizer.

Methods from the literature

Runs with behavior (%)
Evidence-basedReuse / over-reach
55 of 82 runs transfer methods; 6 reuse published solutions. Behaviors can overlap.

6.9× lower cost on Denoising

EvoDuet reaches 100.27% NDG at $38.45. SimpleTES reaches 100% at an estimated API-equivalent $265.39.

View the denoising program →
Held-out NDG (%) ↑
Cost (USD) · log scale
Pareto frontier
Best observed results. SimpleTES cost is estimated; the bar shows 0.5–2× workload sensitivity.

Citation

BibTeX
Download .bib
Show BibTeX
@misc{lee2026evoduet,
  title  = {{EvoDuet}: Bilevel Co-Evolution of Web Searching
            and Task Solving for Scientific Discovery},
  author = {Lee, Young-Jun and Baek, Jinheon and Jeong, Soyeong
            and Kang, Minki and Jwa, Seungyeon and Choi, Jonghyun
            and Han, Seungho and Kang, Dongyeop},
  year   = {2026},
  eprint = {2609.40340},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url    = {https://arxiv.org/abs/2609.40340}
}

Paper figure